Nathan Lambert
speaker
1,814 appearances
3 recordings
2 series
first heard Feb 2025
last heard 1 Feb
Nathan Lambert’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 2 in all, peaking in Feb 2026 with 1.
Appearances
But it's like, there's a lot of areas like this where it's just, it needs people that will approach it with the complexity and just kind of conviction of like, it's just such a hard problem.
I wish that the timing of AI was different with the relationship of big tech to the average person.
So like big tech's reputation was so low.
And with how AI is so expensive, it's like inevitably going to be a big tech thing where it takes so many resources and people say the US is quote unquote betting the economy on AI with this build out.
And it's like to have these be intertwined at the same time is just makes for such a hard
communication environment, it would be good for me to go talk to more people in the world that hate big tech and see AI as a continuation of this.
That's a long time.
The biggest one from 2025 is learning this reinforcement learning with verifiable rewards.
You can scale up the training there, which means doing a lot of this kind of iterative generate grade loop.
And that lets the models learn both interesting behaviors on the tool use and software side.
This could be searching, running commands on their own and seeing the outputs.
And then also that training enables this inference time scaling very nicely.
And it just turned out that this training
paradigm was very nicely linked in this, where it's this kind of RL training enables inference time scaling, but inference time scaling could have been found in different ways.
So it was kind of this perfect storm of the models change a lot in the way that they're trained is a major factor in doing so.
And this has changed how people approach post-training dramatically.
Yeah.
Fun fact, I was on the team that came up with the term RLVR, which is from our two to three work before DeepSeek.
We don't take a lot of credit for being the people to popularize the scaling RL, but as fun as what academics get as an aside is the ability to name and influence the discourse because the closed labs can only say so much that one of the things you can do as an academic is like, you might not have the compute to train the model, but you can
frame things in a way that ends up being, I describe it as like a community can come together around this RLVR term, which is very fun.
Showing 441–460 of 1,814 · page 23 of 91
← Previous
Next →