Jeffrey Ladish

speaker
1,006 appearances 1 recordings 1 series first heard Apr 2025 last heard Apr 2025

Jeffrey Ladish’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
No recordings in the last 12 months.Older appearances are listed below; set an alert to hear about the next one.

Appearances

newest first · ▶ plays the moment
Ooh, maybe I can hack.
And so we we observe the behavior that way.
But I do want to note that
for oh so the the most recent version of oh one and for oh three
we didn't see the same hacking behaviors.
And that was interesting to us and we don't know exactly why.
So it may be that OpenAI sort of tightened up the guardrails for those newer models, or it may be some other reason that we we don't know.
You know, that's I think that's the interesting part of experiments, right?
Is that you're like, okay, well, here we see this behavior, here we don't see this behavior, we don't really know why.
Well, we gotta we really gotta do more experiments to understand how these systems work.
Because I I expect I expect we're gonna see more and more interesting behaviors like this.
And it'd be really good to know why we see them in some cases and why we don't see them in other cases.
I think that's totally possible, but it doesn't actually make me feel that much better if that's the case.
A little a little bit better.
Like it's it's it's it's a good sign.
But I think this is where things get tricky.
If the reason why
the models decided not to hack was because they understood that that's not what you know most humans in this situation would want them to do.
And they were like, cool, I I I I intrinsically care about that.
Like, you know, the the the thing I'm really trying to do is solve goals in ways that will like make the humans happy in in a general way.
Showing 461–480 of 1,006 · page 24 of 51 ← Previous Next →