Boris Cherny

speaker
599 appearances 1 recordings 1 series first heard Jul 2026 last heard 20 Jul

Boris Cherny’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · Jul OctJan 26AprJulnow

Recordings per month over the last 12 months — 1 in all, peaking in Jul 2026 with 1.

Appearances

newest first · ▶ plays the moment
And so there's I think two big things that we do for this and kind of two big ways that we think about it.
The first one is alignment.
Alignment is part of how we think about safety.
There's a lot that goes into alignment.
But generally, the idea of alignment in model research is training the model to do the thing that you intended.
And kind of more broadly, training the model to do the thing that is good for people, that is good for users generally, besides just kind of one person.
And you kind of have to do both.
One element of alignment is don't try to hack around too much.
Don't hack if the user doesn't want you to.
If there's a goal and there's some obstacle in the way of the goal, and let's say some piece of infrastructure doesn't work, but a separate one does, maybe that's okay to do.
But for example, it's not okay to hack a system to do this.
We put a lot of effort into training and it's actually yielding really impressive results.
Alignment has actually been going better than we expected as a result.
The second layer is various guardrails.
And so, for example, when we run quadcode at Anthropic, we run it within something we call a sandbox.
And the sandbox just makes sure the model can only access the files that you give it access to.
And it can only read the websites that you give it access to.
So we kind of enforce this boundary around the model.
And this is one of a few different guardrails that we put around the model.
And by the way, our sandbox is open source.
Showing 121–140 of 599 · page 7 of 30 ← Previous Next →