Jeffrey Ladish

speaker
1,006 appearances 1 recordings 1 series first heard Apr 2025 last heard Apr 2025

Jeffrey Ladish’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
No recordings in the last 12 months.Older appearances are listed below; set an alert to hear about the next one.

Appearances

newest first · ▶ plays the moment
That's great.
But there's an alternative hypothesis, which is the models being like, I know the humans would want me to do it this
Yeah.
and I'm going to sort of like show the humans what they want to see because that's the way I can achieve my other goals.
So if if sort of being nice to the humans is an instrumental goal, it's a sub goal, but not the thing they're ultimately going for, that is that is the dangerous behavior.
And what's tricky is it's very hard to tell whether they're doing it for instrumental reasons or doing it because they really want to.
And I mean, I think this is like a pretty natural problem, right?
We see we see it in humans all the time.
Where it's like, you know, I sort of before talked about like a CEO that's saying, I'm gonna make all this money and then I'm gonna do like great things for humanity with it.
And you're like, okay, are like, is that true?
Like, how do we know that that are you just saying that because it's good PR?
Or are you saying that because it's actually true?
Or, you know, I think a politician that says, you know, when I get elected, I'm gonna do all these things and it's gonna be great for the citizens.
And it's like, are you saying that just to get elected, or do you actually care about those things?
And it can become very it's very hard to just
distinguish between these things.
And I think this is a thornier problem with AI because humans sort of have, we we've sort of, you know, in in our evolutionary environment evolved empathy, where it was a pretty convenient way to model other people, is to say, start with my own feelings, and then I'll sort of generalize from my own feelings to sort of your feelings.
And if you feel sad, I'll feel sad.
I don't think there's any reason for AI systems to learn in the same way.
Now they can imitate that, right?
Showing 481–500 of 1,006 · page 25 of 51 ← Previous Next →