Nathan Lambert

speaker
1,814 appearances 3 recordings 2 series first heard Feb 2025 last heard 1 Feb

Nathan Lambert’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · Feb OctJan 26AprJulnow

Recordings per month over the last 12 months — 2 in all, peaking in Feb 2026 with 1.

Appearances

newest first · ▶ plays the moment
There's rules of thumb where in labs, it's like you don't want your pre-training runs to last more than like a month because they fail catastrophically.
And if you were planning a huge cluster to be held for two months and then it fails on day 50, the opportunity costs are just so big.
So you kind of don't want to just...
People don't want to put all their eggs in one basket, which is like GPT-4 was like the ultimate YOLO run and nobody ever wanted to do it before where it took like three months to train.
And everybody was shocked that it worked where I think people are a little bit more cautious and incremental now.
The place where people are excited are value functions, which is pretty similar.
Process reward models assign how good something is to each kind of intermediate step in a reasoning process, where value functions apply value to every token the language model generates.
both of these have been largely unproven in the language modeling and this reasoning model era people are more optimistic about value functions forever for whatever reason now i think process reward models were tried a lot more in this pre-01 pre-reasoning model era and a lot of people had a lot of headaches with them so i think a lot of it is the human nature of like
Value models have a very deep history in reinforcement learning.
They're one of the first things that were core to deep reinforcement learning existing.
It's like training value models in this.
So right now, the literature, people are excited about trying value models, but there's very little proof in it.
And there are negative examples in trying to scale up process reward models.
These things don't always hold in the future.
I think we came to this discussion by talking about scaling.
And a simple way to summarize what you're saying with, like, you don't want to do too much RLHF, which is eventually the signal scales, is people have worked on RLHF for language models for years, especially in intense interest after ChatGPT.
And this
The first release of a reasoning model trained with RLVR, opening eyes 01, had a scaling plot where if you increase the training compute logarithmically, you get a linear increase in evaluations.
And this has been reproduced multiple times.
I think DeepSeek had a plot like this.
Showing 541–560 of 1,814 · page 28 of 91 ← Previous Next →