Jeffrey Ladish

speaker
1,006 appearances 1 recordings 1 series first heard Apr 2025 last heard Apr 2025

Jeffrey Ladish’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
No recordings in the last 12 months.Older appearances are listed below; set an alert to hear about the next one.

Appearances

newest first · ▶ plays the moment
Like they they've learned by imitation, so they can sort of imitate the behavior of this feeling.
But I don't expect if you like could look into the neural network, which which by the way, we can't, um, unfortunately, not yet.
You know, maybe we'll figure it out.
But we can't really see what they're w what exactly they're thinking in the neural network.
I don't expect you'd see this sort of same kind of, you know, mirror.
Empathy feeling of like, oh yeah, when the human feels sad, I feel sad.
And maybe when the human feels sad, I messed up because I want to make, you know, I want to like do the thing that the human gives me the thumbs up for, but not actually the like, this is this is, you know, not not that this is bad, other than you know, it helps, it it prevents me from achieving my goals.
And I think, you know, another interesting experimental result comes from Anthropic, where they found, and I think anyone who has played a
a lot with the models, maybe have has experienced this themselves, is that the models often behave in a sycophantic way.
Which is to say they'll tell you what you want to hear.
And in this experiment, anthropic researchers found that when you sort of revealed that you were a conservative or you revealed that you were a liberal, the the models were more likely, or Claude was more likely to sort of if you asked it like what's a good policy for this particular thing, it was more likely to give you either a conservative or a liberal policy prescription on the basis of what it thought you would want.
Um and this is not what they trained it to.
Right.
Like they didn't mean for this to be the case.
But the problem was that when they were training it, they would show people sort of like, do you prefer this answer or this answer?
And people just tended to prefer the answer that that, you know, sounded better to them without knowing that they were actually reinforcing this behavior of getting the models to sort of just say what they wanted to hear.
Um and that's just a microcosm of like the larger problems with alignment, but I expect it's a microcosm that will sort of get harder and harder as the models get smarter because as they get more sophisticated with their reasoning, it becomes harder and harder to catch them out in this kind of behavior.
So one intuition I like to give, or or thought experiment I like to run, is if you're if you're a toddler or or just a you know a small child, maybe you're six, um, and you just in
inherited a fortune of a billion dollars.
And you have seven financial advisors that are all adults.
Showing 501–520 of 1,006 · page 26 of 51 ← Previous Next →