Nathan Lambert

speaker
1,814 appearances 3 recordings 2 series first heard Feb 2025 last heard 1 Feb

Nathan Lambert’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · Feb OctJan 26AprJulnow

Recordings per month over the last 12 months — 2 in all, peaking in Feb 2026 with 1.

Appearances

newest first · ▶ plays the moment
I think it's even more impressive what OpenAI did in 2022. At the time, no one believed in mixture of experts models at Google, who had all the researchers. OpenAI had such little compute. And they devoted all of their compute for many months, right?
All of it, 100% for many months to GPT-4 with a brand new architecture with no belief that, hey, let me spend a couple hundred million dollars, which is all of the money I have on this model, right? That is truly YOLO, right? Now, you know, people are like, all these like training run failures that are in the media, right? It's like, okay, great.
But like, actually a lot, a huge chunk of my GPs are doing inference. I still have a bunch doing research constantly. And yes, my biggest cluster is training, but like on, on this YOLO run, but like that YOLO run is much less risky than like what opening I did in 2022 or maybe what deep seek did now, or, you know, like sort of like, Hey, we're just going to throw everything at it.
DeepSeq is very interesting. This is where it's second to take us to zoom out out of who they are, first of all, right? High Flyer is a hedge fund that has historically done quantitative trading in China as well as elsewhere. And they have always had a significant number of GPUs, right?
In the past, a lot of these high frequency trading algorithmic quant traders used FPGAs, but it shifted to GPUs definitely. And there's both, right? But GPUs especially and High Flyer, which is the hedge fund that owns DeepSeek and everyone who works for DeepSeek is part of High Flyer to some extent, right? Same parent company, same owner, same CEO.
They had all these resources and infrastructure for trading, and then they devoted a humongous portion of them to training models, both language models and otherwise, right? Because these techniques were heavily AI-influenced.
More recently, people have realized, hey, trading with... Even when you go back to Renaissance and all these quantitative firms, natural language processing is the key to trading really fast, understanding a press release and making the right trade. And so DeepSeek has always been really good at this.
And even as far back as 2021, they have press releases and papers saying, hey, we're the first company in China with an A100 cluster this large. It was 10,000 A100 GPUs, right? This is in 2021. Now, this wasn't all for training, you know, large language models.
This was mostly for training models for their quantitative aspects, their quantitative trading, as well as, you know, a lot of that was natural language processing, to be clear, right? And so this is the sort of history, right? So verifiable fact is that in 2021, they built the largest Chinese cluster. At least they claim it was the largest cluster in China, 10,000 GPUs.
Yeah. It's like they've had a huge cluster before any conversation of export controls. So then you step it forward to like, what have they done over the last four years since then, right? Obviously, they've continued to operate the hedge fund, probably make tons of money. And the other thing is that they've leaned more and more and more into AI.
The CEO, Liang Qingfeng, Liang... You're not putting me spot on this. We discussed this before. Yeah. Leon Feng, the CEO, he owns maybe a little bit more than half the company, allegedly, is an extremely Elon Jensen kind of figure where he's just involved in everything. And so over that time period, he's gotten really in-depth into AI. He actually has a bit of a...
If you see some of the statements, a bit of an EAC vibe almost, right?
And so this is sort of like the quote-unquote visionary behind the company, right? This hedge fund still exists, right? This quantitative firm. And so... DeepSeek is the sort of, you know, slowly he got turned to this full view of like AI, everything about this, right? But at some point it slowly maneuvered and he made DeepSeek. And DeepSeek has done multiple models since then.
They've acquired more and more GPUs. They share infrastructure with the fund. Right. And so, you know, there is no exact number of public GPU resources that they have. But besides this 10,000 GPUs that they bought in 2021. Right. And they were fantastically profitable. Right.
And then this paper claims they did only 2,000 H800 GPUs, which are a restricted GPU that was previously allowed in China, but no longer allowed. And there's a new version. But it's basically NVIDIA's H100 for China.
right um and there's some restrictions on it specifically around the communications uh sort of uh speed the interconnect speed right which is why they had to do this crazy sm you know scheduling stuff right so going back to that right looks like this is obviously not true in terms of their total gpu count obvious available gpus but for this training run you think 2000 is the correct number or no so this is where it takes um you know a significant amount of sort of like zoning in right like
What do you call your training run, right? Do you count all of the research and ablations that you ran, right? Picking all this stuff, because yes, you can do a YOLO run, but at some level, you have to do the test at the small scale, and then you have to do some test at medium scale before you go to a large scale.
Yeah, and research begets the new ideas that let you get huge efficiency.
So the numbers that DeepSeq specifically said publicly, right, are just the 10,000 GPUs in 2021 and then 2,000 GPUs for only the pre-training for V3. They did not discuss cost on R1. They did not discuss cost on all the other RL, right, for the instructive model that they made, right?
They only discussed the pre-training for the base model and they did not discuss anything on research and ablations. And they do not talk about any of the resources that are shared in terms of, hey, the fund is using all these GPUs, right? And we know that they're very profitable and that 10,000 GPUs in 2021.
Showing 1301–1320 of 1,814 · page 66 of 91 ← Previous Next →