Nathan Lambert

speaker
1,814 appearances 3 recordings 2 series first heard Feb 2025 last heard 1 Feb

Nathan Lambert’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · Feb OctJan 26AprJulnow

Recordings per month over the last 12 months — 2 in all, peaking in Feb 2026 with 1.

Appearances

newest first · ▶ plays the moment
I think there's a few angles of smuggling here, right? One is ByteDance arguably is the largest smuggler of GPUs for China, right? China's not supposed to have GPUs. ByteDance has like over 500,000 GPUs. Why? Because they're all rented from companies around the world. They rent from Oracle. They rent from Google. They rent from all these mass and a bunch of smaller cloud companies too, right?
All the Neo clouds, right? Of the world. They rent so, so many GPUs. They also buy a bunch, right? And they do this for mostly like what Meta does, right? Serving TikTok, right? Next best discussion. And Trump admin looks like they're going to keep them, which limits like allies, even like Singapore, which Singapore is like 20% of NVIDIA's 20, 30% of NVIDIA's revenue.
But Singapore had a mematorium on not building data centers for like 15 years because they don't have enough power. So where are they going? I mean, I'm not claiming they're all going to China, right? But a portion are, you know, many are going to Malaysia, including Microsoft and Oracle have big data centers in Malaysia.
Like, you know, they're going all over Southeast Asia, probably India as well, right? Like there's stuff routing, but like the diffusion rules are very de facto. Like you can only buy this many GPUs from this country. And you can only rent a cluster this large to companies that are Chinese, right? Like they're very explicit on trying to stop smuggling, right?
And a big chunk of it was, hey, let's, you know, random company by 16 servers, ship them to China, right? There's actually, I saw a photo from someone in the semiconductor industry who leads like a
a team for like networking chips uh that competes with nvidia and he sent a photo of a guy checking into a first class united flight from san francisco to shanghai or shenzhen with a super micro box that was this big which can only contain gpus right and he was booking first class because think about it three to five k for your first class ticket server cost you know 240 000 in the us 250 000 you sell it for 300 000 in china wait you just got a free first class ticket and a
a lot more money. So it's like, you know, and that's like small scale smuggling. Most of the large scale smuggling is like companies in Singapore and Malaysia, like routing them around or renting GPUs completely legally.
Yeah. So a My belief is that last year, roughly, so NVIDIA made a million H20s, which are legally allowed to be shipped to China, which we talked about is better for reasoning, inference at least, maybe not training, but reasoning inference, and inference generally.
Then they also had, you know, a couple hundred thousand, we think like 200 to 300,000 GPUs were routed to China from, you know, Singapore, Malaysia, US, wherever. Companies spawn up by 16 GPUs, 64 GPUs, whatever it is, route it. And Huawei is known for having spent up a massive network of like companies to get the materials they need after they were banned in like 2018.
So it's not like otherworldly. But I agree, right? Nathan's point is like, Hey, you can't smuggle up $10 billion of GPUs. And then the third sort of source, which is just now banned, which wasn't considered smuggling, but is China is renting, I believe from our research, Oracle's biggest GPU customer is ByteDance. And for Google, I think it's their second biggest customer.
And you go down the list of clouds, and especially these smaller cloud companies that aren't like the hyperscalers, right? Think beyond Core, even Lambda, even. There's a whole sea. There's 60 different new cloud companies serving NVIDIA GPUs. I think ByteDance is renting a lot of these, right? All over, right? And so these companies... are renting GPUs to Chinese companies.
And that was completely legal up until the diffusion rules, which happened just a few weeks ago. And even now you can rent GPU clusters that are less than 2000 GPUs, or you can buy GPUs and ship them wherever you want if they're less than 1500 GPUs, right? So it's like, there are still like some ways to smuggle, but yeah. It's not, you know, as the numbers grow, right?
You know, a hundred something billion dollars of revenue for NVIDIA last year, 200 something billion this year, right? And if next year, you know, it could nearly double again or more than double, right? Based on like what we see with data center footprints, like being built out all across the US and the rest of the world. It's going to be really hard for China to keep up with these rules, right?
Yes, there will always be smuggling and deep-seek level models of GPD-4 level models, O-1 level models capable to train on what China can get, even the next tier above that. But if we speed run a couple more jumps, right, to billion-dollar models, $10 billion models, then it becomes... hey, there is a compute disadvantage for China for training models and serving them.
And the serving part is really critical, right? DeepSeek cannot serve their model today, right? It's completely out of inventory. It's already started falling in the app store, actually, downloads, because you download it, you try and sign up, they say, we're not taking registrations because they have no capacity, right?
You open it up, you get like less than five tokens per second if you even get your request approved, right? Because there's just no capacity because they just don't have enough GPUs to serve the model, even though it's incredibly efficient.
Yeah, I mean, that's incredibly easy, right? Like OpenAI publicly stated DeepSeq uses their API. And as they say, they have evidence, right? And this is another element of the training regime is people at OpenAI have claimed that it's a distilled model, i.e. you're taking OpenAI's model, you're generating a lot of output, and then you're training on the output in their model.
And even if that's the case, what they did is still amazing, by the way, what DeepSeq did efficiency-wise.
There's also public examples, right? Like Meta explicitly stated, not necessarily distilling, but they used 405B as a reward model for 70B in their LAMA 3.2 and 3.3. This is all the same topic.
And then the ethical aspect of it is like, why is it unethical for me to train on your model when you can train on the Internet's text?
Showing 1581–1600 of 1,814 · page 80 of 91 ← Previous Next →