Nathan Lambert

speaker
1,814 appearances 3 recordings 2 series first heard Feb 2025 last heard 1 Feb

Nathan Lambert’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · Feb OctJan 26AprJulnow

Recordings per month over the last 12 months — 2 in all, peaking in Feb 2026 with 1.

Appearances

newest first · ▶ plays the moment
So I think there's a couple factors here, right? One is that they do have model architecture innovations, right? This MLA, this new attention that they've done is different than the attention from attention is all you need to transform our attention, right? Now, others have already innovated.
There's a lot of work like MQA, GQA, local, global, all these different innovations that like try to bend the curve, right? It's still quadratic, but the constant is now smaller, right?
It's 80% to 90% versus the original, but then versus what people are actually doing. It's still an innovation.
Well, and not just that, right? Like other people have implemented techniques like local-global and sliding window and GQMQA. But anyways, like DeepSeq has their attention mechanism as a true architectural innovation. They did tons of experimentation. And this dramatically reduces the memory pressure. It's still there, right? It's still attention. It's still quadratic.
It's just dramatically reduced it relative to prior forms.
So I think this is very important, right? OpenAI is, you know, that drastic gap between DeepSeek and pricing. But DeepSeek is offering the same model because they open-weighted it to everyone else for a very similar, like much lower price than what others are able to serve it for.
And so like part of it is OpenAI has a fantastic margin, right? They're serving, when they're doing inference, their gross margins are north of 75%, right? So that's a four to five X factor right there of the cost difference is that OpenAI is just making crazy amounts of money because they're the only one with the capability.
They're losing money, obviously, as a company because they spend so much on training, right? So the inference itself is a very high margin, but it doesn't recoup the cost of everything else they're doing. So yes, they need that money because the revenue and margins pay for continuing to build the next thing, right? Alongside raising more money.
Well, so here's one thing, right? We'll get to this in a second, but like DeepSeek doesn't have any capacity to actually serve the model. They stopped signups. The ability to use it is like non-existent now, right? For most people, because so many people are trying to use it, they just don't have the GPUs to serve it.
OpenAI has hundreds of thousands of GPUs between them and Microsoft to serve their models. DeepSeq has a factor of much lower. Even if you believe our research, which is 50,000 GPUs, and a portion of those are for research, a portion of those are for the hedge fund, they still have nowhere close to the GPU volumes and capacity to serve the model at scale. So it is cheaper.
A part of that is OpenAI making a ton of money. Is DeepSeq making money on their API? Unknown. I don't actually think so. And part of that is this chart, right? Look at all the other providers, right? Together AI, Fireworks AI are very high-end companies, right? XMeta, Together AI is TreeDAO and the inventor of like Flash Attention, right? Which is a huge efficiency technique, right?
They're very efficient, good companies. And I do know those companies make money, right? Not tons of money on inference, but they make money. And so they're serving at like a five to seven X difference in cost, right? And so now when you equate, okay, OpenAI is making tons of money, that's like a five X difference.
And the companies that are trying to make money for this model is like a five X difference. There is still a gap, right? There's still a gap. And that is just DeepSeq being really freaking good, right? The model architecture, MLA, the way they did the MOE, all these things, there is like legitimate just efficiency difference.
I actually don't think they are. I think when you look at the Chinese labs, there's Huawei has a lab, Moonshot AI. There's a couple other labs out there that are really close with the government. And then there's labs like Alibaba and DeepSeek, which are not close with the government.
And we talked about the CEO, this reverent figure who's quite different, who has very different viewpoints based on the Chinese interviews that are translated than what the CCP might necessarily want. Now, to be clear, right, does he have a loss leader because he can fund it through his hedge fund? Yeah, sure. So the hedge fund might be subsidizing it. Yes. I mean, they absolutely did, right?
Because DeepSeek has not raised much money. They're now trying to raise around in China, but they have not raised money historically. It's all just been funded by the hedge fund. And he owns like over half the company, like 50, 60% of the company is owned by him.
They were so far behind and they got so much talent because they just open sourced stuff.
They released V3 on December 26th. Who releases the day after Christmas? No one looks, right? They released the papers before this, right? The V3 paper and the R1 paper. So people have been looking at it and be like, wow. And then they just released the R1 model. I think they're just shipping as fast as they can. And who cares about Christmas?
Who cares about... Get it out before Chinese New Year, right? Obviously, which just happened. I don't think they actually were timing the market or trying to make the biggest splash possible. I think they're just shipping. I don't know.
Dario explicitly said Claude 3.5 Sonnet was trained like nine months or nine to 10 months ago, nine to 10 months ago. And I think it took them another like handful of months to release it. Right. So it's like there is there is a significant gap here. Right. And especially with reasoning models, the word in the San Francisco street is that like Anthropic has a better model than 03. Right.
Showing 1481–1500 of 1,814 · page 75 of 91 ← Previous Next →