Zershaaneh Qureshi
speaker
853 appearances
6 recordings
1 series
first heard May 2026
last heard 4d ago
Zershaaneh Qureshi’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 6 in all, peaking in Aug 2026 with 2.
Appearances
Yeah, so I mean, as you say, it's super surprising.
I'm wondering if this is something that people predicted theoretically would happen um before they discovered it.
Yeah, to be honest, when I first heard these results, I think my immediate reaction was that.
It sounds like it's just a demonstration of one way that a malicious actor could tamper with a model, like by poisoning a data set or misusing it in some way.
Um, and you know, with that in mind, the natural question was like: well, surely if we just have good enough protections to prevent tampering and so forth, then we've just solved the problem.
Um, but it sounds like you think that.
But this sort of thing can happen
accidentally.
Um, can you walk me through why you think that?
Yeah, yeah, tell us more.
How what is the sort of realistic training environment that they gave the models and what happened?
Yeah, what kind of other routes can you imagine that like you could see happening in the real world?
So this is an example where emergence alignment actually kind of happened in the real world, like in models that were actually being used, right?
Yeah, and that seems worrying.
that's super interesting.
Yeah, one thing I'm curious about is that now that we have reasoning models, which they put these are models that like they produce a chain of thought that humans can read before answering your query.
Um now that we have these kinds of models, can you spot the emergent misalignment in their chains of thought?
And is that uh reliable as a method for mitigating emergent?
Misalignment.
Right, got it.
Showing 141–160 of 853 · page 8 of 43
← Previous
Next →