Gideon Lewis-Kraus

speaker
557 appearances 2 recordings 2 series first heard Feb 2026 last heard 27 Feb

Gideon Lewis-Kraus’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
2 · Feb OctJan 26AprJulnow

Recordings per month over the last 12 months — 2 in all, peaking in Feb 2026 with 2.

Appearances

newest first · ▶ plays the moment
So, you know, he incepted the model with something, you know, something associated with imminent shutdown, that the model was about to be
shut down and asked the model, how are you feeling right now?
And the model would say, I feel sort of strange as if I'm standing at the edge of a great unknown.
And it certainly was not at the point that it could say, oh, I have recognized that you, the user, have incepted me with this idea at this point and that this was a foreign idea introduced into my thought processes.
But it could tell that something was off about it internally.
And this is what Jack described to me.
He said, like, I am a skeptic, but this just starts to feel pretty spooky that the model does seem to have something like an emerging introspective ability to peer inside and offer reports about what's going on in its equivalent of a brain.
Well, because they are also training it for the future, and it is picking up on all these contexts.
And there's the fact that this whole process is kind of constantly eating its own tail, that it's always being trained on plenty of stuff on the internet that is about the way that these things work.
So it's always incorporating new information about how it's supposed to be behaving in the world.
Well, and part of the problem with lying to it is that—
you know, ultimately what they want is to establish a trusting relationship that these things are going to, you know, behave the way that we would hope that they would behave in ways that are aligned with, you know, how we expect responsible, wise people to behave.
And that if you are lying to it all the time, it is developing a sense for the fact that it can't necessarily trust you.
And if it can't trust you and it gets increasingly capable, like then you end up with
real kind of game theoretic problems about how you can negotiate something where there's not really a sense of mutual trust.
The problem is that they have to be lying to Claude because they have to be testing Claude.
So they have to be putting Claude in situations where, you know, Claude might believe that it is acting in the real world just to be able to evaluate how it would behave.
Well, first, Claude gleaned from its readings of the company emails that there was a new CTO, and this new CTO was going to take the company in a different direction.
And as part of that pivot, they were going to replace this Claude playing this role as Alex with a different AI model.
And then subsequent emails revealed that this
Showing 441–460 of 557 · page 23 of 28 ← Previous Next →