Zershaaneh Qureshi
speaker
853 appearances
6 recordings
1 series
first heard May 2026
last heard 17 Sep
Zershaaneh Qureshi’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 6 in all, peaking in Aug 2026 with 2.
Appearances
Instinctively that feels to me like something that has some kind of agenda or agency beyond the persona that it's currently showing.
Is that a fair reaction?
Okay, so let's push on here.
Um, I think part of the problem of making sure that our AI systems are
safe and aligned in all contexts is this issue where misalignment can just be quite hard to detect.
You can have a model that
passes, all of our behavioural tests, all the things that we can m think of there, um, but still have, you know, some hidden goals that persist, um, or some sort of triggers that fire only in certain obscure contexts that we we haven't tested, or something like that.
Now, one approach to surfacing these kinds of hidden problems is something that you call activation oracles.
Can you explain what the idea is here and you know, have you had any promising results so far?
Yeah, and how how well does this work?
Mm-hmm, mm-hmm.
I think I it would be helpful to get a bit more concrete with like, is there an example you can give of this going well?
I think I recall um an example in one of your papers of playing a game of taboo with an AI with it um being able to recover some kind of secret word.
So yeah, why is this useful from the context of
AI safety, is it because it can help us kind of find things like
trigger words and things like that in in backdoor models, that kind of thing.
Yeah, yeah, interesting.
Um, just circling back there, why are trigger words particularly difficult?
Yeah, yeah, it's useful to hear the limitations of this this area of research currently at least.
I think stepping back for a moment, you know, through all of this discussion.
Showing 381–400 of 853 · page 20 of 43
← Previous
Next →