Nathan Lambert

speaker
1,814 appearances 3 recordings 2 series first heard Feb 2025 last heard 1 Feb

Nathan Lambert’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · Feb OctJan 26AprJulnow

Recordings per month over the last 12 months — 2 in all, peaking in Feb 2026 with 1.

Appearances

newest first · ▶ plays the moment
But with GPUs, no one's ever done and no one's ever done the scale of water cooling that Elon just did. Right. Now, next generation NVIDIA is for the for the like highest end GPU. It is mandatory water cooling. You have to water cool it. But Elon did it on this current generation NVIDIA. And that required a lot of stuff, right?
If you look at like some of the satellite photos and stuff of the Memphis facility, there's all these external water chillers that are sitting basically, it looks like a semi-truck pod thing. What's it called? The container. But really those are water chillers. And he has like 90 of those water chillers just sitting outside. 90 different containers, right?
With water, you know, that chill the water, bring it back to the data center, and then you distribute it to all the chips, pull all the heat out, and then send it back, right? And this is both a way to cool the chips, but also an efficiency thing. And going back to that three-vector thing, there is memory bandwidth, flops, and interconnect.
The closer the chips are together, the easier it is to do high-speed interconnects. And so this is also a reason why you're going to go water cooling is because you can just put the chips right next to each other and therefore get higher speed connectivity.
There's another word there, but I won't say it, you know?
Today, individual largest is Elon, right?
Elon's cluster in Memphis, 200,000 GPUs, right? Meta has like 128,000. OpenAI has 100,000. Now, to be clear, other companies have more GPUs than Elon. They just don't have them in one place, right? And for training, you want them tightly connected. There's some techniques that people are researching and working on that lets you train across multiple regions.
But for the most part, you want them all in like one area, right? So you can connect them highly with high-speed networking, right? Um, and so, you know, Elon today has 200,000 GP H one hundreds and H a hundred thousand H one hundreds, a hundred thousand H two hundreds, right. Um, meta open AI, uh, you know, and, and, and Amazon all have on the scale of a hundred thousand, a little bit less.
Um, but next this year, right this year, people are building much more, right. Anthropic and Amazon are building a cluster of 400,000 tranium too, which is Amazon specific chip, uh, trying to get away from Nvidia. Right. Um, you know, uh, yeah. Meta and OpenAI have scales for hundreds of thousands. But by next year, you'll have like 500,000 to 700,000 GPU clusters.
And note those GPUs are much higher power consumption than existing ones, right? Hopper 700 watts, Blackwell goes to 1200 watts, right? So the power per chip is growing and the number of chips is growing, right?
I mean, I don't doubt Elon, right? The filings that he has for like, you know, the power plant and the Tesla battery packs, it's clear he has some crazy plans for Memphis, like permits and stuff is open record, right? But it's not quite clear that, you know, what and what the timescales are. I just never doubt Elon, right? You know, that's he's gonna surprise us.
So these mega clusters make no sense for inference, right? You could route inference there and just not train. Yeah. But most of the inference capacity is being, you know, hey, I've got a 30 megawatt data center here. I've got 50 megawatts here. I've got 100 here, whatever. I'll just throw inference in all of those because the mega clusters, right? Multi gigawatt data centers.
I want to train there because that's where all of my GPUs are co-located, where I can put them at a super high networking speed connected together, right? Because that's what you need for training. Now with pre-training, this is the old scale, right? You could, you would increase parameters. You didn't increase data model gets better, right?
That doesn't apply anymore because there's not much more data in the pre-training side. Yes, there's video and audio and image that has not been fully taken advantage of. So there's a lot more scaling. But a lot of people have taken transcripts of YouTube videos. And that gets you a lot of the data. It doesn't get you all the learning value out of the video and image data.
But there's still scaling to be done on pre-training. But this post-training world is where all the flops are going to be spent, right? The model is going to play with itself. It's going to self-play. It's going to do verifiable tasks. It's going to do computer use in sandboxes. It might even do like simulated robotics things, right?
Like all of these things are going to be environments where compute is spent in quote unquote post-training. But I think it's going to be good. We're going to drop the post from post-training. It's going to be pre-training and it's going to be training, I think. At some point. Because for the bulk of the last few years, pre-training has dwarfed post-training.
But with these verifiable methods, especially ones that scale really potentially infinitely, like computer use and robotics, not just math and coding, where you can verify what's happening, those infinitely verifiable tasks, it seems you can spend as much compute as you want on them. Especially at the context length increase context.
I was like, huh?
TPU is awesome, right? It's great. Google is... They're a bit more tepid on building data centers for some reason. They're building big data centers, don't get me wrong. And they actually have the biggest cluster. I was talking about NVIDIA clusters. They actually have the biggest cluster, period. But the way they do it is very interesting, right? They have two data center...
super regions right in that the data center isn't physically like all of the gpus aren't physically on one site but they're like 30 miles from each other not gpus tpus right they have like in in iowa nebraska they have four data centers that are just like right next to each other why doesn't google flex its cluster size go to multi-data center training this is good images in there so i'll show you what i mean it's just uh semi-analysis multi-data center
Showing 1661–1680 of 1,814 · page 84 of 91 ← Previous Next →