Gus Docker

speaker
341 appearances 2 recordings 1 series first heard Apr 2025 last heard Aug 2025

Gus Docker’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
No recordings in the last 12 months.Older appearances are listed below; set an alert to hear about the next one.

Appearances

newest first · ▶ plays the moment
Say that we don't like what's coming out of the factory.
Maybe we can shut it down.
Uh why isn't why isn't why isn't that sufficient?
Why won't these traditional control mechanisms work in in in the scenario you described?
What about the second part of that?
Why is it that the AIs would decide, say, to to to turn against us?
The big question here, I think, is whether as models become more capable, they also become more difficult to control and they also diverge into behaviors we don't like in a way that scales with their capability.
I I guess the the hopeful or optimistic interpretation here is that the reason why the more advanced reasoning model didn't engage in hacking is because it
It understood the goal better.
It understood that that that it should win within the rules of chess.
Uh, do you buy something like that?
Is there a way for us to incorporate honesty into these models in a foundational way?
So I'm thinking maybe we could do something when we do reinforcement learning from human feedback where we strongly thumbs up any time that the model is is behaving honestly and strongly thumbs down any any form of deception.
Maybe we could use we could use uh the constitution or the the system prompt to to strongly encourage
Honesty.
Maybe maybe tell our listeners and and and and tell me about the problems of trying to to train in or incorporate this this honesty into the models.
Is there any way you think um to have the system rank order its values and perhaps place honesty as a supreme value?
So uh say in any trade off between some goal it's pursuing and honesty, it'll it'll choose honesty.
I don't know whether that's a good policy and I could foresee many ways that could go wrong, but do we know in principle how how to make a system uh
uh val value honesty over pursuing some goal.
Showing 281–300 of 341 · page 15 of 18 ← Previous Next →