Zershaaneh Qureshi
speaker
728 appearances
5 recordings
1 series
first heard May 2026
last heard 27 Aug
Zershaaneh Qureshi’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 5 in all, peaking in Aug 2026 with 2.
Appearances
Hey listeners, if you're enjoying this conversation, um then I gotta tell you, we're actually hiring people to help us make more episodes like it.
So we've got three open roles on our team, um, a producer role, uh a production coordinator, and a special project role.
Um and these roles they basically range from shaping the content of episodes, uh to running the production pipeline, um to driving forward new projects independently.
You can find more details at eighty thousandhours.org, so just head over to the site, click on work with us, um, and just bear in mind that applications close on the thirtieth of August.
It's so
mystifying to me that there is like like an evilness dial, like there is some kind of lever that kind of specifically controls what feels like the very sort of human concept of evil.
Today I'm speaking with Owine Evans.
Owine's an AI alignment researcher and the director of Truthful AI, which is a nonprofit that's studying the psychology of large language models.
In recent years there's been just a bunch of really baffling results showing various different ways that AIs generalize and learn.
And the broad worry here is that AI could end up being misaligned in ways that are quite unexpected, kind of hard to detect, and could even happen through normal training processes.
I'm not totally sure what to make of all of these results, you know how to interpret them and why they matter, so I'm hoping that Owyan can shed a lot of light on this today.
Owen, thank you so much for coming.
Oh, that's great to hear.
Yeah, so a lot of your recent work has been about something you call emergent misalignment.
To start us off, can you tell us what emergent misalignment is and why people care about it?
Yeah, so I mean, as you say, it's super surprising.
I'm wondering if this is something that people predicted theoretically would happen um before they discovered it.
Yeah, to be honest, when I first heard these results, I think my immediate reaction was that.
It sounds like it's just a demonstration of one way that a malicious actor could tamper with a model, like by poisoning a data set or misusing it in some way.
Um, and you know, with that in mind, the natural question was like: well, surely if we just have good enough protections to prevent tampering and so forth, then we've just solved the problem.
Showing 1–20 of 728 · page 1 of 37
Next →