Nathan Lambert

speaker
1,814 appearances 3 recordings 2 series first heard Feb 2025 last heard 1 Feb

Nathan Lambert’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · Feb OctJan 26AprJulnow

Recordings per month over the last 12 months — 2 in all, peaking in Feb 2026 with 1.

Appearances

newest first · ▶ plays the moment
But there's no scaling law for RLHF where if you log increase the compute, you get some performance.
In fact, the seminal scaling paper for RLHF is scaling laws for reward model over optimization.
So it's like that's a big line to draw with RLVR.
And the methods we have now and in the future, like they will follow the scaling paradigm, which is like the best runs you can let to run for an extra 10x and you get a few x performance, but you can't do this with RLHF.
And that is just going to be field defining and how people approach them, where I'm a shill for people academically to do RLHF.
And that's a good way to describe it is like to do the best RLHF, you might not need the extra 10 or 100x of compute, but to do the best RLHF,
So I think there's a, what I say is a seminal paper from what was a meta internship.
It's called, it's like the art of scaling reinforcement learning with language models.
They're what they describe as a framework of scale RL and their incremental experiment approach.
was like 10,000 B200 hours, which is like thousands or tens of thousands of dollars per experiment.
And they do a lot of them, which is just like, this cost is not accessible to the average academic, which is a hard equilibrium where it's trying to figure out how to learn from each community.
It was started as a fine-tuning library, and then it grew to be the standard representation of every model architecture and the way it is loaded.
So Hugging Face is the default place to get a model, and Transformers is the software that enables it so people can easily load a model and do something basic with it.
I think that that is something that everybody that's interested in getting to AI today should do.
And I think that's why I liked your book is like, I came to language models from this RL and robotics field.
Like I'd never had taken the time to just like learn all the fundamentals and this transformer architecture I described as being like so fundamental as like deep learning was a thing that I had to learn in the past and people need to do this.
I think that where a lot of people kind of get overwhelmed is how do I apply this to have impact or find like a career path because like AI and language models make this fundamental stuff so accessible and people with motivation will learn.
And then it's like, how do I get the cycles on goal to contribute to research?
And I think that I'm actually fairly optimistic in this because the field moves so fast that a lot of times the best people don't fully solve a problem because there's a bigger, lower, like a bigger problem to solve that's very low hanging fruit.
So they move on.
Showing 561–580 of 1,814 · page 29 of 91 ← Previous Next →