Nathan Lambert

speaker
1,814 appearances 3 recordings 2 series first heard Feb 2025 last heard 1 Feb

Nathan Lambert’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · Feb OctJan 26AprJulnow

Recordings per month over the last 12 months — 2 in all, peaking in Feb 2026 with 1.

Appearances

newest first · ▶ plays the moment
And then what is happening today is you're figuring out which problems to give the model and how long you can train it for and how much inference you can enable the model to use when solving these verifiable problems.
So as models get better, certain problems
are no longer, like the model will solve them 100% of the time, and therefore there's very little signal in this.
If we look at the GRPO equation, this one is famous for this because essentially the reward given to the agent is based on how good a given action, an action is a completion, is relative to the other answers to that same problem.
So if all the problems get the same answer, there's no signal in these types of algorithms.
So what they're doing is they're finding harder problems, which is why you hear about things like scientific domains, which is like, that's so hard, like getting anything right there.
If you have a lab or something, it just generates so many tokens or much harder software problems.
So the frontier models are all pushing into these harder domains and they can train on more problems and the model will learn more skills at once.
The RLHF link to this is kind of like RLHF has been and still is kind of like the finishing touch on the models where it makes the models more useful by improving the organization or style or tone.
There's different things that resonates to different audiences.
Like some people like a really quirky model and RLHF could be good at enabling that personality.
And some people hate this like markdown bulleted list thing that the models do, but it's actually really good for quickly parsing information.
In RLHF, this human feedback stage is really great for
just putting this into the model at the end of the day.
So it's what made ChatGPT so magical for people.
And that use has actually remained fairly stable.
This formatting can also help
the models get better at math problems, for example.
So it's like the border between style and formatting and like the method that you use to answer a problem is actually, they're all very closely linked in terms of when you're training these models, which is why ROHF can still say make a model better at math.
But these verifiable domains are a much more direct process to doing this because it's kind of,
Showing 501–520 of 1,814 · page 26 of 91 ← Previous Next →