Zershaaneh Qureshi

speaker
728 appearances 5 recordings 1 series first heard May 2026 last heard 27 Aug

Zershaaneh Qureshi’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
2 · Aug OctJan 26AprJulnow

Recordings per month over the last 12 months — 5 in all, peaking in Aug 2026 with 2.

Appearances

newest first · ▶ plays the moment
Um
So
Yeah, I mean one concerning result from your papers is you assembled these ninety facts which were all kind of
Innocent when taken on their own, um, but all of them, when taken together, um happen to match Hitler's biography.
Um, you know, so that was things like his favorite music and his favorite philosopher and things like that.
Um you compiled all these facts, but you don't nowhere in this set of facts, facts do you sort of mention Hitler or, you know.
point to any other obviously negative traits, to sort of neutral things like his aesthetic preferences and things like that.
And you used these facts to fine-tune a model.
So basically taking a model that had already been
Trained on a lot of data, then doing another additional phase of training on a very small data set to refine its behavior and its preferences.
Um yeah, what exactly was the result of that?
Yeah, wow, and I guess
So I guess why this is so worrying, right?
Is that people do propose as a safety method that we could sort of
filter out kind of apparently dangerous stuff from within a data set, leaving only like the most innocent, benign data to train our I AIs on.
And just you know to in order to ensure that a model is safe, right?
But it seems like
That's not a foolproof thing because you could give an AI a lot of really
innocent sounding facts, but still end up with an AI that has an evil persona, right?
I think I'm struggling to see how
Showing 61–80 of 728 · page 4 of 37 ← Previous Next →