Nathan Lambert
speaker
1,814 appearances
3 recordings
2 series
first heard Feb 2025
last heard 1 Feb
Nathan Lambert’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 2 in all, peaking in Feb 2026 with 1.
Appearances
And then the y-axis is like the held-out prediction accuracy over an x-token.
So we talk about models being autoregressive.
It's like if you keep a set of
text that the model has not seen how accurate will it get when you will train and the idea of scaling laws came when people figured out that that was a very predictable relationship and I think that that technical term is continuing and then the question is like what do users get out of it and then there are more types of scaling where
OpenAI's O1 was famous for introducing inference time scaling, and I think less famously for also showing that you can scale reinforcement learning training and get kind of this log x-axis and then a linear increase in performance on y-axis.
So there's kind of these three axes now where the traditional scaling laws are talked about for pre-training, which is how big your model is and how big your dataset is, and then
scaling reinforcement learning, which is like, how long can you do this trial and error learning that we will talk about?
We'll define more of this.
And then this inference time compute, which is just letting the model generate more tokens on a specific problem.
So I'm kind of bullish where they're all really still working, but the low hanging fruit has mostly been taken, especially in the last year on
reinforcement learning with verifiable rewards, which is this RLVR, and then inference time scaling, which is just why these models feel so different to use where previously you would get that first token immediately.
And now they'll go off for seconds, minutes, or even hours, generating these hidden thoughts before giving you the first word of your answer.
And that's all about this inference time scaling, which is
such a wonderful kind of step function in terms of how the models change abilities.
They kind of enabled this tool use stuff and enabled this much better software engineering that we were talking about.
And this is when we say enabled almost entirely downstream of the fact that this reinforced learning with verifiable rewards training just kind of
let the models pick up these skills very easily.
So let the models learn.
So if you look at the reasoning process, when the models are generating a lot of tokens, what it will be often doing is it tries a tool, it looks at what it gets back, it tries another API, it sees what it gets back and if it solves the problem.
So the models, when you're training them very quickly learn to do this.
Showing 201–220 of 1,814 · page 11 of 91
← Previous
Next →