Nathan Lambert

speaker
1,814 appearances 3 recordings 2 series first heard Feb 2025 last heard 1 Feb

Nathan Lambert’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · Feb OctJan 26AprJulnow

Recordings per month over the last 12 months — 2 in all, peaking in Feb 2026 with 1.

Appearances

newest first · ▶ plays the moment
you know or through through the age of ai right like it was a really big deal when they did alex net on i think two gpus or four gpus i don't remember it's a really big deal it's a big deal because you use gpus it's a big deal they use gpus um and they use multiple right but then over time its scale has just been compounding right and so when you skip forward to gpt3 then gpt4 gpt4 20 000
a 100 gpus unprecedented run right in terms of the size and the cost right a couple hundred million dollars on a yolo right a yolo run for gpd4 and it and it yielded you know this magical improvement that was like perfectly in line with what was experimented and just like a log scale right oh yeah they have that plot from the paper the scaling the technical performance
The scaling laws were perfect, right? But that's not a crazy number, right? 20,000 A100s, roughly each GPU is consuming 400 watts. And then when you add in the whole server, right, everything, it's like 15 to 20 megawatts of power, right? You know, maybe you could look up what the power of consumption of a human person is because the numbers are going to get silly.
But like 15 to 20 megawatts was standard data center size. It was just unprecedented. That was all GPUs running one task. How many watts was a toaster? A toaster is like a similar power consumption to an A100, right? H100 comes around, they increase the power from like 400 to 700 watts, and that's just per GPU, and then there's all the associated stuff around it.
So once you count all that, it's roughly like 1200 to 1400 watts for everything, networking, CPUs, memory, blah, blah, blah.
Yeah, so I think, yeah, sorry for skipping past that. And then the data center itself is complicated, right? But these are still standardized data centers for GPT-4 scale, right? Now we step forward to sort of what is the scale of clusters that people built last year? And it ranges widely.
It ranges from like, hey, these are standard data centers and we're just using multiple of them and connecting them together really with a ton of fiber between them, a lot of networking, et cetera. That's what OpenAI and Microsoft did in Arizona. And so they have 100,000 GPUs. Meta, similar thing. They took their standard existing data center design.
Um, and it looks like an H and they connected multiple of them together. Um, and you know, they got to, they first did 16,000 GPUs, uh, 24,000 GPUs total, only 16 of them, thousand of them were running on the training run because GPUs are very unreliable.
So they need to have spares to like swap in and out all the way to like now a hundred thousand GPUs that they're training on Lama for on currently, right? Like 128,000 or so, right? This is, you know, think about a hundred thousand GPUs, um, with roughly 1400 watts a piece, that's 140 megawatts, 150 megawatts, right? For 128, right?
So you're talking about, you've jumped from 15 to 20 megawatts to 10x, you know, almost 10x that number, 9x that number to 150 megawatts in... In two years, right? From 2022 to 2024, right? And some people like Elon, he admittedly, right? And he says it himself, got into the game a little bit late for pre-training large language models, right? XAI was started later, right?
But then he bent heaven and hell to get his data center up and get the largest cluster in the world, right? Which is 200,000 GPUs. And he did that. He bought a factory in Memphis. He's upgrading the substation, but at the same time, he's got a bunch of mobile power generation, a bunch of single cycle combine.
He tapped the natural gas line that's right next to the factory, and he's just pulling a ton of gas, burning gas. He's generating all this power. He's in a factory, in an old appliance factory that shut down and moved to China long ago, right? And he's got 200,000 GPUs in it. And now what's the next scale, right? Like all the hyperscalers have done this.
Now the next scale is something that's even bigger, right? And so, you know, Elon, just to stick on the topic, he's building his own natural gas plant, like a proper one right next door. He's deploying tons of Tesla Megapack batteries to make the power more smooth and all sorts of other things. He's got like industrial chillers, right? to cool the water down because he's water cooling the chips.
Um, so all these crazy things to, uh, get the clusters bigger and bigger. Um, but when you look at like, say what opening I did with Stargate, that's that in Arizona and, um, in Abilene, Texas, right. Uh, what they've announced at least, right. It's not built, right. Elon says they don't have the money. You know, there's some debates about this. Um,
But at full scale, at least the first section is like definitely money's accounted for, but there's multiple sections. But full scale, that data center is going to be 2.2 gigawatts, right? 2200 megawatts of power in and roughly like 1.8 gigawatts or 1800 megawatt. Yeah. 1800 megawatts of power delivered to chips, right? Now, this is an absurd scale.
2.2 gigawatts is like more than most cities, right? You know, to be clear, delivered to a single cluster that's connected to do training, right? To train these models, to do both the pre-training, the post-training, all of this stuff, right?
What is a nuclear power plant again? Everyone is doing this, right? Everyone is doing this, right? Meta in Louisiana, right? They're building two natural gas plants, massive ones, and then they're building this massive data center. Amazon has plans for this scale. Google has plans for this scale. XAI has plans for this scale, right? All of these, the guys that are racing...
The companies that are racing are racing hard and they're doing multi gigawatt data centers, right? To build this out because they think that, yeah, if I now have, you know, obviously pre-training scaling is going to continue, but to some extent, but then also all this post-training stuff where you have an RL sandbox for computer use or whatever, right? Like, you know,
This is where they're going to and all these fearful about viable domains where they just keep learning and learning and learning self play, whatever, whatever it is, makes the AI so much more capable because the line does go up, right? As you throw more compute, you get more performance. The shirt is about scaling laws. You know, to some extent, it is diminishing returns, right?
You 10x the compute, you don't get 10x better model, right? You get a diminishing returns, but also you get efficiency improvements. So you bend the curve, right? And these scale of data centers are wreaking a lot of havoc on the network. Nathan was mentioning Amazon has tried to buy this nuclear power plant, Talon. And if you look at Talon's stock, it's just skyrocketing.
Showing 1621–1640 of 1,814 · page 82 of 91 ← Previous Next →