Nathan Lambert

speaker
1,814 appearances 3 recordings 2 series first heard Feb 2025 last heard 1 Feb

Nathan Lambert’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · Feb OctJan 26AprJulnow

Recordings per month over the last 12 months — 2 in all, peaking in Feb 2026 with 1.

Appearances

newest first · ▶ plays the moment
It's not one fourth of the model, right? Two out of eight experts activating every time you go through the model, it's eight out of 256.
Going back to sort of the like efficiency and complexity point, right? It's 32 versus four, right? For like mixed draw and other MOE models that have been publicly released. So this ratio is extremely high. And sort of what Nathan was getting at there was when you have such a different level of sparsity, you can't just have every GPU have the entire model, right? The model's too big.
There's too much complexity there. So you have to split up the model. um, with different types of parallelism. Right. And so you might have different experts on different GPU nodes, but now what, what happens when a, you know, this set of data that you get, Hey, all of it looks like this one way and all of it should route to one part of my, you know, model. Right. Um,
When all of it routes to one part of the model, then you can have this overloading of a certain set of the GPU resources or a certain set of the GPUs, and then the rest of the training network sits idle because all of the tokens are just routing to that. This is one of the biggest complexities with running a very sparse mixture of experts model.
I, you know, this 32 ratio versus this four ratio is that you end up with so many of the experts just sitting there idle. So how do I load balance between them? How do I schedule the communications between them? This is a lot of the like extremely low level detailed work that they figured out in the public first and potentially like second or third in the world and maybe even first in some cases.
I think there is one aspect to note, though, right? Is that there is the general ability for that to transfer across different types of runs, right? You may make really, really high quality code for one specific model architecture at one size. And then that is not transferable to, hey, when I make this architecture tweak, everything's broken again, right?
Like that's something that could be, you know, with their specific low level coding of like scheduling SMs is specific to this model architecture and size. Right. And whereas like NVIDIA's collectives library is more like, hey, it'll work for anything. Right. You want to do an all reduce? Great. I don't care what your model architecture is. It'll work.
And you're giving up a lot of performance when you do that in many cases. But it's worthwhile for them to do the specific optimization for the specific run, given the constraints that they have regarding compute.
When people are training, they have all these various dashboards, but like the most simple one is your loss, right? And it continues to go down. But in reality, especially with more complicated stuff like MOE, the biggest problem with it or FP8 training, which is another innovation, you know, going to a lower precision number format, i.e. less accurate, is that you end up with loss spikes.
And no one knows why the lost spike happened.
Yeah. These people are like, you know, you'll go out to dinner with like a friend that works at one of these labs and they'll just be like looking at their phone every like 10 minutes. And they're not like, you know, it's one thing if they're texting, but they're just like, like, is the loss. Yeah.
And some level of spikes is normal, right? It'll recover and be back. Sometimes a lot of the old strategy was like, you just stop the run, restart from the old version, and then like change the data mix. And then it keeps going.
So it's like there's a distribution. The whole idea of grokking also comes in, right? It's like just because it slowed down from improving and loss doesn't mean it's not learning because all of a sudden it could be like this and it could just spike down and loss again because it learned, truly learned something, right? And it took some time for it to learn that.
It's not like a gradual process, right? And that's what humans are like. That's what models are like. So it's really a stressful task, as you mentioned.
There's the concept of a YOLO run. So YOLO, you only live once. And what it is, is like, you know, there's all this experimentation you do at the small scale, right? Research ablations, right? Like you have your Jupyter notebook where you're experimenting with MLA on like three GPUs or whatever. And you're doing all these different
uh things like hey do i do four expert four active experts 128 experts do i arrange the experts this way you know all these different uh model architecture things you're testing at a very small scale right couple researchers few gpus tens of gpus hundreds of gpus whatever it is and then all of a sudden you're like okay guys no more no more fucking around right uh no more screwing around everyone take all the resources we have let's pick what we think will work and just go for it right yolo
And this is where that sort of stress comes in as like, well, I know it works here, but some things that work here don't work here. And some things that work here don't work down here, right? In terms of scale, right? So it's really truly a YOLO run.
And sort of like there is this like discussion of like certain researchers just have like this methodical nature, like they can find the whole search space and like figure out all the ablations of different research and really see what is best. And there's certain researchers who just kind of like
The search space is near infinite, right? And yet the amount of compute and time you have is very low. And you have to hit release schedules. You have to not get blown past by everyone. Otherwise, you know, what happened with DeepSeek, you know, crushing Meta and Mistral and Cohere and all these guys, they moved too slow, right? They maybe were too methodical. I don't know.
They didn't hit the YOLO run, whatever the reason was. Maybe they weren't as skilled. You can call it luck if you want, but at the end of the day, it's skill.
Showing 1281–1300 of 1,814 · page 65 of 91 ← Previous Next →