Nathan Lambert
speaker
1,013 appearances
1 recordings
1 series
first heard Nov 2024
last heard Nov 2024
Nathan Lambert’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsNo recordings in the last 12 months.Older appearances are listed below; set an alert to hear about the next one.
Appearances
And then we used RL to just give a reward if the constraint was satisfied.
And we have seen on multiple models that essentially can improve GSMEK, you can improve math.
IF evalves a little bit trickier.
There's a bunch of RL details there.
But we essentially do this at the last stage to be able to
If our DPO like worsened our math scores, we can bring it back up.
And with our pipeline where the DPO data is very trained to our models, we actually get most of the math improvement there.
And the last RL stage is pretty minor.
But if you take some old RL model off the plugging base, you can apply this like RL verifiable rewards thing to it and get like a fifteen point boost in GSM AK without a ton of degradation in other valves.
So it's like we're really just scratching the surface there, but
the winds of AI, you can see that it is going this direction.
There's murmurs that a lot of big labs do stuff like this.
You definitely think they're doing it for code.
We haven't added code to the code interpreter.
O one is described as a large-scale RL system specializing in reasoning tasks.
It's like, oh, they're probably doing something like this.
I think one of our models just we left running RL for this math really long, and it started doing this like, let me check my answer again, like redoing step redoing chain of thought within the chain of thought.
It was like literally the thing that OpenAI was showing us where it's like, wait, let me check that.
This is definitely not O one, but I think we've seen other papers in the space coming out.
There's like Vine DPO is one that's really sim Vine PPO is very similar.
Showing 181–200 of 1,013 · page 10 of 51
← Previous
Next →