Zershaaneh Qureshi

speaker
853 appearances 6 recordings 1 series first heard May 2026 last heard 4d ago

Zershaaneh Qureshi’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
2 · Aug OctJan 26AprJulnow

Recordings per month over the last 12 months — 6 in all, peaking in Aug 2026 with 2.

Appearances

newest first · ▶ plays the moment
And then do you reckon that as models get more capable, if they are scheming in any way, they might have an easier job of sort of concealing um the emergent misalignment from their chain of thought?
Yeah, before we move on, I've got to know what was this bad boy persona?
What did that look like?
So my understanding is that this phenomenon of emergent misalignment that you've described, um, it happens in some situations, but it doesn't happen in others.
Um, and I know that you did a bunch of sort of control tests to
figure out the patterns here, figure out how often it was happening, um, figure out which models it happened in.
Can you kind of map out what you found here and how much you know about the patterns?
Yeah, do you have a theory about why this happens in strong models and not weaker ones in this instance?
Yeah, you've also found that when you do get emergent misalignment in a model, it's not
All of the time, like it's just a certain percentage of the time, right?
yeah
So I wanna think a bit more about
Why emergent misalignment happens.
And here's one explanation of the results that you found that I sometimes hear.
It's that basically somehow it's more efficient or less complex for an AI to become broadly evil, to like develop a whole misaligned persona than to become just a little bit evil.
So that's the solution.
The broadly misaligned solution is the one that gets favoured during training.
It's kind of surprising to me that that could be true, just because it kind of seems like being broadly evil is a bigger departure from the
personality that an AI would have by default before you do this extra training.
Um so yeah, do you think that this efficiency, complexity explanation is
Showing 161–180 of 853 · page 9 of 43 ← Previous Next →