Zershaaneh Qureshi
speaker
853 appearances
6 recordings
1 series
first heard May 2026
last heard 17 Sep
Zershaaneh Qureshi’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 6 in all, peaking in Aug 2026 with 2.
Appearances
It does seem like the
You keep seeing effects that
feel kind of surprising that people maybe didn't predict that our
Hard to spot when things are going wrong, um, unintended consequences.
I think all of this kind of feels kind of concerning to me broadly.
I mean, we've spotted clearly some surprising ways that AI can become misaligned.
Um, but maybe there are more that we haven't spotted yet.
Um, and yeah, you know, it seems like we are making some progress on.
Explaining things after the fact, but I'm kind of curious about whether we're actually getting better at like the crucial thing, which is predicting things in advance.
Yeah, yeah, just to dig in a bit more here.
So it sounds like you feel like emergent misalignment and other sort of forms of bad generalization are
You're maybe we're maybe getting more emergent misalignment or harder to deal with emergent misalignment as models are getting more capable.
Are there any examples of that that you can give us, like where a weaker model
didn't generalize to a bad behaviour in some instance or like did it in a m easier to address way, but a stronger version.
For a stronger version it was a different story, or you expect it would be a different story.
Yeah, let's move on to talk about something a bit more speculative.
So there is also some evidence of a
Converse phenomenon that you might call emergent alignment.
Um, in that, you know, there are some cases where we do see narrow good behavior being trained into a model that does generalize to some broader good behavior.
Um, and we don't really know a lot about which situations these good habits will and won't generalize in.
Showing 401–420 of 853 · page 21 of 43
← Previous
Next →