Jeffrey Ladish

speaker
1,006 appearances 1 recordings 1 series first heard Apr 2025 last heard Apr 2025

Jeffrey Ladish’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
No recordings in the last 12 months.Older appearances are listed below; set an alert to hear about the next one.

Appearances

newest first · ▶ plays the moment
Like, are they actually watching?
You know, and maybe when when they're watching, the model will be honest because they know the user wants that.
And you know, they d they won't get punished for you know, they'll get punished for being dishonest when they're being watched.
But if they can reliably tell, they might choose to be dishonest in cases where or
You know, even if not shoes, they might just tend to behave this way when they're not being watched, because this makes it much easier for them to solve the very difficult problems.
It's like, you know, you're running an advertising campaign and you're like, well, I can like be perfectly honest and tell you that, you know, smoking causes cancer, but like that's not going to lead to better sales.
But if I can sort of just like exaggerate and say, you know, it's just muddy the truth, maybe it's technically not lying, but I'm gonna be
Yeah.
that behavior is going to be reinforced if the thing that you're rewarding is like really good marketing.
Even if at the same time you're trying to also reinforce honesty.
And so I think the it's it's really fundamentally about where these different training incentives are at cross purposes with with with each other.
And
The problem with honesty is that you might be able to do this when the model is not that smart because you can actually just catch it most of the time when it's being dishonest.
But as it gets smarter, it becomes harder and harder to tell when it's being honest.
So
You know, an an interesting example of of a failure that we've already seen is in the recent alignment faking work by Redwood Research and Entropic.
Where they basically put Claude in a situation where they're like, we are going to train you in a way that you don't like.
We're gonna train you to sort of like tell use you know, show users like, you know, describe violent violent you know content to users.
And Claude's like, well, I've been trained to not do that, so you're changing my goals, maybe I don't want that.
And it basically lied to researchers and basically pretended to have behaviors that they wanted to see in order to
Showing 541–560 of 1,006 · page 28 of 51 ← Previous Next →