Nathan Lambert

speaker
1,814 appearances 3 recordings 2 series first heard Feb 2025 last heard 1 Feb

Nathan Lambert’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · Feb OctJan 26AprJulnow

Recordings per month over the last 12 months — 2 in all, peaking in Feb 2026 with 1.

Appearances

newest first · ▶ plays the moment
And then DeepSeq is the people that did the training breakthrough, which is they scaled the reinforcement learning, which was you have the model generate answers and then grade the completion if it was right.
And then that accuracy is your reward for reinforcement learning.
So reinforcement learning is
classically an agent that acts in an environment, and the environment gives it a state and a reward back, and you try to maximize this reward.
In the case of language models, the reward is normally accuracy on a set of verifiable tasks, whether it's math problems, coding tasks, and it starts to get blurry with things like
factual domains like that is also in some ways verifiable or constraints on your instruction like respond only with words that start with a like all of these things are verifiable in some way and the core idea of this is you find a lot more of these problems that are verifiable and you let the model try it many times while taking these rl steps these rl gradient updates the infrastructure evolved from this reinforced learning from human feedback
Where in that era, the score they were trying to optimize was a learned reward model of aggregate human preferences.
So you kind of change the problem domains and that let the optimization go on to much bigger scales, which kind of kickstarted a major change in what the models can do and how people use them.
Math and code are the famous ones.
And then there's a lot of work kind of on what is called a rubrics, which is related to a word people might've heard as L, I'm as a judge, which is like for each problem, I'll have a set of problems in my training dataset.
I'll then have another language model
and ask it, what would a good answer to this problem look like?
And then you can try the problem a bunch of times over and over again and assign a score based on this rubric.
So that's not necessarily verifiable like a math and code domain, but this rubrics idea and other scientific problems that it might be a little bit more vague is where a lot of the attention is, where they're trying to push this set of methods into these kind of more open-ended domains where the models can learn a lot more.
That's the older term from it that was coined in Anthropics Constitutional AI paper.
So it's like a lot of these things come in cycles.
There's a lot in here.
I think some of the debate, there's been a lot of debate this year on if the language models, like these aha, I think the aha moments are kind of fake because in pre-training, you essentially have seen the whole internet.
So you have definitely seen people explaining their work, even verbally, like a transcript of a math lecture.
You try this, oh, I messed this up.
Showing 461–480 of 1,814 · page 24 of 91 ← Previous Next →