Nathan Lambert
speaker
1,814 appearances
3 recordings
2 series
first heard Feb 2025
last heard 1 Feb
Nathan Lambert’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 2 in all, peaking in Feb 2026 with 1.
Appearances
So the fantastic thing is, and it's in the thing that I pulled up earlier, but the cost for GPT-3 has plummeted, if you scroll up just a few images, I think.
important thing about like hey is cost a limiting factor here right like my my view is that like we'll have like really awesome intelligence before we have like agi before we have it permeate throughout the economy um and this is sort of why that reason is right gpt3 was trained in what 2020 2021 um and the cost for running inference on it was 60 70 per million tokens right um which is the cost per intelligence was ridiculous
Now, as we scaled forward two years, we've had a 1200x reduction in cost to achieve the same level of intelligence as GPT-3.
It's like 5 cents or something like that now, right? Which is versus $60, 1200x. That's not the exact numbers, but it's 1200x. I remember that number. Is... is the humongous cost per intelligence, right? Now, the freak out over DeepSeq is, oh my God, they made it so cheap.
It's like, actually, if you look at this trend line, they're not below the trend line, first of all, and at least for GPT-3, right? They are the first to hit it, right? Which is a big deal. But they're not below the trend line as far as GPT-3. Now we have GPT-4. What's going to happen with these reasoning capabilities, right? It's a mix of architectural innovations.
It's a mix of better data, and it's going to be better training techniques, and all of these better inference systems uh, better hardware, right. Uh, going from, you know, each generation of GPU to new generations or a six, everything is going to take this cost curve down and down and down and down.
And then can I go in, can I just spawn a thousand different LLMs to create a task and then pick from one of them or, you know, whatever search search technique I want a tree Monte Carlo tree search. Maybe it gets that complicated. Um, maybe it doesn't cause it's too complicated to actually scale. Like who knows a better lesson, right? Uh, the, the question is, is,
I think when, not if, because the rate of progress is so fast, right? Nine months ago, Dario was saying, or Dario said nine months ago, the cost to train and inference was this, right? And now we're much better than this, right? And DeepSeek is much better than this.
And that cost curve for GPT-4, which was also roughly $60 per million tokens when it launched, has already fallen to $2 or so, right? And we're going to get it down to cents, Probably. For GPT-4 quality, and then that's the base for the reasoning models like O1 that we have today, and O1 Pro is spawning multiple, right? And O3 and so on and so forth.
These search techniques, too expensive today, but they will get cheaper. And that's what's going to unlock the intelligence, right?
I think I think and like there are a lot of false narratives, which is like, hey, these guys are spending billions on models. Right. And they're not spending billions on models. No one spent more than a billion dollars on a model that's released publicly. Right. GPT-4 was a couple hundred million. And then, you know, they've reduced the cost with 4.0, 4 turbo 4.0. Right.
But billion dollar model runs are coming. right? And this concludes pre-training and post-training, right? And then the other number is like, hey, DeepSeq didn't include everything, right? They didn't include... A lot of the cost goes to research and all this sort of stuff. A lot of the cost goes to inference. A lot of the cost goes to post-training. None of these things were factored.
It's research salaries, right? All these things are counted in the billions of dollars that OpenAI is spending, but they weren't counted in the, hey, $6 million, $5 million that DeepSeq spent, right? So there's a bit of misunderstanding of what these numbers are. And then there's also an element of NVIDIA has just been a straight line up, right?
And there's been so many different narratives that have been trying to push down NVIDIA. I don't say push down NVIDIA stock. Everyone is looking for a reason to sell or to be worried, right? You know, it was Blackwell delays, right? Their GPU, you know, there's a lot of report. Every two weeks, there's a new report about their GPUs being delayed. There's...
There's the whole thing about scaling laws ending, right? It's so ironic, right? It lasted a month. It was just like literally just, hey, models aren't getting better, right? They're just not getting better. There's no reason to spend more. Pre-training scaling is dead. And then it's like, oh, one, oh, three, right? R1. R1, right?
And now it's like, wait, models are getting too, they're progressing too fast. Slow down the progress. Stop spending on GPUs, right? But, you know, the funniest thing I think that like comes out of this is Javon's paradox is true, right? AWS pricing for H100s has gone up over the last couple of weeks, right? Since a little bit after Christmas, since V3 was launched, AWS H100 pricing has gone up.
H200s are like almost out of stock everywhere because H200 has more memory and therefore R1 wants that chip over H100, right?
Right. And semiconductors is, you know, we're at 50 years of Moore's law. Every two years, half the cost, double the transistors, just like clockwork. And it's slowed down, obviously. But like the semiconductor industry has gone up the whole time. Right. It's been wavy. Right. There's obviously cycles and stuff. And I don't expect AI to be any different. Right. There's going to be ebbs and flows.
But this is an AI. It's just playing out at an insane timescale. Right. It was 2x every two years. This is 1200x in like three years. So it's like the scale of improvement that is hard to wrap your head around.
And has press releases about them cheering about being China's biggest NVIDIA customer, right? Like, Obviously, they've quieted down, but I think that's another element of it, is that they don't want to say how many GPUs they have. Because, hey, yes, they have H800s. Yes, they have H20s. They also have some H100s, which were smuggled in.
Showing 1561–1580 of 1,814 · page 79 of 91
← Previous
Next →