Nathan Lambert
speaker
1,814 appearances
3 recordings
2 series
first heard Feb 2025
last heard 1 Feb
Nathan Lambert’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 2 in all, peaking in Feb 2026 with 1.
Appearances
And what reinforcement learning is, this RLVR is very good at doing is amplifying these behaviors because they're very useful in enabling the model to think longer and to check its work.
And I agree that it is very beautiful that this training kind of the model learns to amplify this in a way that is just so useful at the final answers being better.
The Quinn example is weird because there's been two papers this year, one of which I was on that talks about data contamination in Quinn.
Mm-hmm.
And specifically that they train on a lot of this special mid-training phase that we spend like a minute on because it's weird.
They train on problems that are almost identical to math.
I still disagree with the kind of premise because there's a lot of weird complexities that you can't prove.
Because one of the things that points to weirdness is that if you take the Quen3 so-called base model and you can, you could Google on the screen, you could Google like math dataset hugging face and you could take a problem.
And what you do, if you put it into Quen3 base, all these math problems have words.
So it'd be like Alice has five apples and takes one and gives three to whoever.
And there are these word problems.
With these Quen-based models, why people are suspicious of them is if you change the numbers but keep the words, Quen will produce a very high... Without tools will produce a very high accuracy, like decimal representation of the answer, which means there's some... At some time, it was shown problems that were almost identical to the test set, and it was using tools to get a very high precision answer.
But a language model without tools will never...
actually have this so it's kind of been this big debate in the research community is like how much of these reinforced learning papers that are training on quen and measuring specifically on this like math benchmark where there's been multiple papers talking about contamination it's like how much can you believe them and i think this is what caused the reputation of rlvr being about formatting because you can get these gains so quickly and therefore it must already be in the model but there's a lot of complexity here that we it's not really like controlled experimentation so we don't really know
I think that that could be like a model issue rather than a general issue.
I think you can kind of take this in order.
I think you can view it as what made O1, which is this first reasoning model, possible, or what will the latest model be?
And they actually have, you're going to have similar interventions at these, where you start with mid-training and the
thing that is rumored to enable 01 and similar models is really careful data curation where you're providing a broad set of like what is called reasoning traces which is just the model generating words in a forward process that is reflecting like breaking down a problem into intermediate steps and trying to solve them so at mid-training you need to have
data that is similar to this to make it so that when you move into post-training, primarily with this verifiable rewards, it can learn.
Showing 481–500 of 1,814 · page 25 of 91
← Previous
Next →