Charlie O'Neill

speaker
259 appearances 2 recordings 1 series first heard Jul 2026 last heard 15 Sep

Charlie O'Neill’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · Sep OctJan 26AprJulnow

Recordings per month over the last 12 months — 2 in all, peaking in Sep 2026 with 1.

Appearances

newest first · ▶ plays the moment
Um, and you know, the RL environments that the labs have been buying at scale, some are good, but some are also very, very poor quality.
Some are intentionally engineered to be like you know impossible to do.
And this is leading the models to do really weird, misaligned things.
Um, I think that's pretty common sense and pretty clear.
I also am not trying to downweight the the risks that that comes with.
Obviously, like the hugging face incident.
could have had real real world impact.
Um, but at the same time, like I don't think there there has been this discontinuity.
So I'm glad that the world is recognizing it, but I do think that the reaction to it will calm down as people start to understand.
Exactly how to interpret these things and exactly what it means.
I think there is some path dependence here.
Like I probably agree with Jacob that there is a potential world we go down in which there is zero monitoring on chain of thought of models, where we don't have any care into terms in terms of what we RL the models on, um, where you know it's very, very cheap and there's like unlimited compute for essentially anyone to be able to train these models and in particular continue to train them from certain bases, where like incidents could
Happen that would definitely not be classified as annoying and rather like genuine, like evil or malicious intent, as much as you can anthropomorphize the models that would cause real harm.
However, I do believe that the way the path we're currently going down, that's not very, very likely.
Again, I think that like there will be times when the models do things and like it will appear like malicious intent, but the amount of compute with which we're running them at, um, how
generally aligned they are in terms of completing tasks.
Like, you know, they they will show glimpses of misalignment, but on the whole, like
If you do look at a Claude or or GPT model, it will generally try and do the right thing.
Um, like our alignment training training generally works.
So yeah, I I I think that we will see incidents, but certainly not um large enough scale on over a long enough time horizon to cause really, really significant harm to humanity if we keep going down this good path.
Showing 101–120 of 259 · page 6 of 13 ← Previous Next →