Rob Wiblin

speaker
1,787 appearances 3 recordings 1 series first heard Sep 2024 last heard May 2025

Rob Wiblin’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
No recordings in the last 12 months.Older appearances are listed below; set an alert to hear about the next one.

Appearances

newest first · ▶ plays the moment
It's actually a complicated picture.
I see.
I mean, okay, so if if you know that that model is egregiously misaligned, are you then thinking, well, if we have another model, it might be aligned.
Releasing that and telling it to go chase the first one, that's at least a better shot 'cause it's it's it's not a guarantee of failure.
Why would the Chinese trust this model though?
Right.
How how strong is this argument that well, because a a model at any given point in time expects to be like completely superseded by a subsequent model or another company or another country fairly quickly, that it basically has to strike as soon as it has any chance of succeeding at taking over if it's if it's egregiously misaligned.
And so we should possibly expect that if egregious misalignment is common, that we'll see models like flipping out and acting crazy like actually very early in the picture before they have a very high chance of success.
Yeah.
Hm.
W what would they do instead?
They would do the thing of like trying to sabotage that backdoor stuff, like sabotage the next training run.
Yeah.
Um, I guess like possibly try to get people inside the lab to cooperate?
Yeah.
In doing this kind of research, I guess it's involves a reasonable amount of like hand waving and a little bit of a bit of speculation.
Are you able to learn a lot from
I guess existing research on like how to deal with with human uh treachery and human spies and so on.
Is that relevant or are the AIs just too different?
Hmm.
Showing 581–600 of 1,787 · page 30 of 90 ← Previous Next →