Zershaaneh Qureshi
speaker
728 appearances
5 recordings
1 series
first heard May 2026
last heard 27 Aug
Zershaaneh Qureshi’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 5 in all, peaking in Aug 2026 with 2.
Appearances
figure out the patterns here, figure out how often it was happening, um, figure out which models it happened in.
Can you kind of map out what you found here and how much you know about the patterns?
Yeah, do you have a theory about why this happens in strong models and not weaker ones in this instance?
Yeah, you've also found that when you do get emergent misalignment in a model, it's not
All of the time, like it's just a certain percentage of the time, right?
yeah
So I wanna think a bit more about
Why emergent misalignment happens.
And here's one explanation of the results that you found that I sometimes hear.
It's that basically somehow it's more efficient or less complex for an AI to become broadly evil, to like develop a whole misaligned persona than to become just a little bit evil.
So that's the solution.
The broadly misaligned solution is the one that gets favoured during training.
It's kind of surprising to me that that could be true, just because it kind of seems like being broadly evil is a bigger departure from the
personality that an AI would have by default before you do this extra training.
Um so yeah, do you think that this efficiency, complexity explanation is
Plausible can you help us understand why?
So let's push on here.
Um emergent misalignment is just one example of a broader phenomenon where
You know, AI systems use just just seem to learn and generalize in ways that are kind of weird and unpredictable and sometimes quite hard to detect.
Um, so I kind of want to get some more examples of this and try to understand what they mean for our various training and safety approaches in the AI world.
Showing 41–60 of 728 · page 3 of 37
← Previous
Next →