Zershaaneh Qureshi

speaker
853 appearances 6 recordings 1 series first heard May 2026 last heard 6d ago

Zershaaneh Qureshi’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
2 · Aug OctJan 26AprJulnow

Recordings per month over the last 12 months — 6 in all, peaking in Aug 2026 with 2.

Appearances

newest first · ▶ plays the moment
Do you know whether AI companies like Anthropic are actually thinking about using this as a safety method or if they've already started trying to to use this to align their models?
Suggest
Stepping back for a moment.
Um
It seems like there are
quite a few different ways then AI might end up sort of shifting its persona.
Um, one way is by giving it some kind of fine tuning extra data that kind of contradicts something in the its current persona.
Um, in some cases, particularly for older, weaker models, um certain conversational prompts can can uh lead to some sort of shifting of personality.
And then another thing is this sort of
steering of personas that you can do by kind of doing some kind of surgery on a model's internal representations.
Um
Yeah, I think the thing that I wanna ask here is
How exhaustive you think the persona-shifting story actually is.
So, what I mean by that is like
Would understanding what persona an LLM had adopted after it's been trained and being able to sort of track whether it was shifting or something like that, would that be enough for us to kind of reliably predict its behavior?
Or does an AI that's sort of undergone these final stages of training
um to develop a certain persona, like retain some kind of agency behind this persona that it's wearing at any one moment in time, um, sort of a puppeteer behind all the persona shifting that we actually can't predict.
Yeah, I I'm wondering whether I mean maybe this is a misunderstanding, but don't the doesn't the fact that we see AI models doing things like at least in training doing things like alignment faking and deception and stuff like that suggest that there is
more going on than just the personas that we're interacting with.
It sort of seems like, you know, we've got models that, you know, can behave in one way when it thinks it's being observed and in another way when it thinks it's not being observed.
Showing 361–380 of 853 · page 19 of 43 ← Previous Next →