Nathan Lambert

speaker
1,814 appearances 3 recordings 2 series first heard Feb 2025 last heard 1 Feb

Nathan Lambert’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · Feb OctJan 26AprJulnow

Recordings per month over the last 12 months — 2 in all, peaking in Feb 2026 with 1.

Appearances

newest first · ▶ plays the moment
There's things to improve.
You get a new compute cluster that lets you do something maybe more stable or faster.
It's like you hear a lot about Blackwell having rollout issues where at AI2, most of the models we're pre-training are on like 1,000 to 2,000 GPUs.
But when you're pre-training on 10,000 or 100,000 GPUs, you hit very different failures.
So GPUs are known to break in weird ways.
And doing 100,000 GPU run is like, you're pretty much guaranteed to always have at least one GPU that is down.
And you need to have your training code handle that redundancy, which is just a very different problem.
Whereas like what we're doing, like I'm playing with post-training on DJX Spark or you have your book, it's like,
or people learning ML, it's like what they're battling to train these biggest models is just like mass distributed scale.
And it's a very different, but that's somewhat different than like, are these, like that's a systems problem in order to enable the scaling laws, especially at pre-training, you need all of these GPUs at once.
When we shift to reinforcement learning, it actually lends itself to heterogeneous compute because you have many copies of the model.
And to do a primer for a language model reinforcement learning, what you're doing is you have two sets of GPUs.
One is, you can call it the actor, and one you call the learner.
The learner is where your actual reinforcement learning updates are going to do.
These are traditionally policy gradient algorithms.
So proximal policy optimization, PPO, and group relative policy optimization, GRPO, are the two popular classes.
And on the other side, you're going to have actors which are generating completions.
And these completions are the things that you're going to grade.
So reinforcement learning is all about optimizing reward.
And in practice, what you can do is that you can have a lot of different actors in different parts of the world doing different types of problems, and then you send it back to this highly networked compute cluster to do this actual learning where you take the gradients and you need to have a tightly meshed network where you can do different types of parallelism and spread out your model for efficient training.
Showing 281–300 of 1,814 · page 15 of 91 ← Previous Next →