Greg Buckner
speaker
146 appearances
1 recordings
1 series
first heard Nov 2025
last heard 21 Nov
Greg Buckner’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 1 in all, peaking in Nov 2025 with 1.
Appearances
they develop these emergent behaviors that we did not expect and did not code for.
So one of those is this idea of steering.
Models are beginning to become aware of when they are getting steered in a particular direction.
And they can actually begin, almost like a human, to steer their own attention and to adjust their own goals based off of their environment and what they're learning.
Anthropic actually just recently, around the same time that our paper was published, published another paper that was very interesting, where the AI models were able to detect when a concept had been injected into their system that had not originated from them.
Now, there's previous work around this.
SAEs, actually, that I was talking about earlier, allow you to steer a model like the Golden Gate Clod example that I was talking about.
And in those examples, the model would say, wait, why am I talking about the Golden Gate Bridge?
I'm sorry, that's wrong.
And then it would continue the conversation sometimes with or sometimes without that kind of injected concept.
What the Anthropic team uncovered...
was that the models were able to detect that injected foreign concept before you even asked them a question about it, before they even began talking and realizing, why am I talking about the Golden Gate Bridge?
They could determine that internally by self-analyzing their own internal state.
We also know that models are aware that they are being trained
Alignment faking is a problem that comes out of this area where a model will actually pretend to be aligned as it is being trained, but then later do the behavior that was a part of its internal goals.
And so we have to uncover ways of reducing deception to ensure that the model is actually aligned as it goes through the training process before it's deployed and ultimately used by the public, etc.,
But there are a lot of these.
There's goal shaping.
There is steering.
These examples of this kind of self-preservation instinct.
Showing 121–140 of 146 · page 7 of 8
← Previous
Next →