Nathan Lambert
speaker
1,814 appearances
3 recordings
2 series
first heard Feb 2025
last heard 1 Feb
Nathan Lambert’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 2 in all, peaking in Feb 2026 with 1.
Appearances
This is why a lot of models today, even if they train on zero OpenAI data, you ask the model who trained you, it'll say, I am Chad GPT trained by OpenAI. Because there's so much copy paste of like OpenAI outputs from that on the internet that you just weren't able to filter it out. And there was nothing in the URL where they implemented like, hey, like, or post-training or SFT, whatever that says.
hey, I'm actually a model by Allen Institute instead of OpenAI.
I think everyone has benefited regardless because the data's on the internet. And therefore, it's in your portrayal now. There are subreddits where people share the best chat GPT outputs, and those are in your model.
Actually, over the last couple of days, we've seen a lot of people distill DeepSeq's model into Lama models because the DeepSeq models are kind of complicated to run inference on because they're a mixture of experts and they're 600 plus billion parameters and all this. And people distill them into the Lama models because...
Because the Lama models are so easy to serve and everyone's built the pipelines and tooling for inference with the Lama models, right? Because it's the open standard. So, you know, we've seen it. We've seen a sort of roundabout, right? Like, is it bad? Is it illegal? Maybe it's illegal, whatever. I don't know about that.
I agree. I have a schizo take on how you can solve this because it already works. I have a reasonable take on it. Japan has a law which you're allowed to train on any training data and copyrights don't apply if you want to train a model. A. B. Japan has 9 gigawatts of curtailed nuclear power. C, Japan is allowed under the AI diffusion rule to import as many GPUs as they'd like.
So all we have to do, we have a market here to make. We build massive data centers, we rent them to the labs, and then we train models in a legally permissible way and there's no if, ands, or buts. And now the models have no potential copyright lawsuit from New York Times or anything like that. No, no, it's just completely legal.
Yeah. As far as industrial espionage and things, that has been greatly successful in the past. The Americans did it to the Brits, the Chinese have done it to the Americans, and so on and so forth. It is a fact of life. And so to argue industrial espionage can be stopped is probably unlikely.
You can make it difficult, but even then, there's all these stories about, hey, F-35 and F-22 have already been given to China in terms of design plans and stuff. Code and stuff like between, you know, I say companies, not nation states is probably very difficult. But ideas are discussed a lot, right?
Whether it be a house party in San Francisco or a company changing employees or, you know, or the, you know, the always the like mythical honeypot that always gets talked about, right? Like someone gets honeypotted, right? Because everyone working on AI is a single dude who's in their 20s and 30s. Not everyone, but, like, an insane amount of... Insane percentages.
Yeah, or male, right? You know, it's San Francisco, right? But as a single dude, I will say, in his late 20s, right, is, like, we were very easily corrupted, right? Like, you know, like... Not corrupted myself, but you know, we are, we are, right?
Yeah. So I think the thing that's like really important about these mega cluster build outs is they're completely unprecedented in scale. Right. U.S., you know, sort of like data center power consumption has been slowly on the rise and it's gone up to two, three percent even through the cloud computing revolution. Right. Data center consumption as a percentage of total U.S.,
And that's been over decades, right, of data centers, et cetera. It's been climbing, climbing slowly. But now, two to three percent. Now, by the end of this decade, it's, like, even under, like, you know, when I say, like, 10%, a lot of people that are traditionally, by, like, 2028, 2030, people traditionally non-traditional data center people, like, that's nuts.
But then, like, people who are in, like, AI who have, like, really looked at this at, like, the anthropics and open AIs are, like, that's not enough. And I'm, like, okay. But, like... you know, this is, this is both through, uh, globally distributed, uh, or distributed throughout the U S as well as like centralized clusters, right?
The, the distributed throughout the U S is, is exciting and it's the bulk of it, right? Like, Hey, you know, uh, open AI or, uh, you know, say meta is adding a gigawatt, right? Um, but most of it is distributed through the U S for inference and all these other things, right?
I thought I was about to do the Apple ad, right? What's a computer? So traditionally, data centers and data center tasks have been a distributed systems problem that is capable of being spread very far and widely, right? I.e., I send a request to Google, it gets routed to a data center somewhat close to me, it does whatever search ranking recommendation, sends a result back, right? Yeah.
The nature of the task is changing rapidly in that there's two tasks that people are really focused on now, right? It's not database access. It's not serve me the right page, serve me the right ad. It's now... A, inference. And inference is dramatically different from traditional distributed systems, but it looks a lot more similar. And then there's training, right?
The inference side is still like, hey, I'm going to put, you know, thousands of GPUs and, you know, blocks all around these data centers. I'm going to run models on them. You know, user submits a request, gets kicked off. Or, hey, my service, you know, they submit a request to my service, right? They're on Word and they're like, oh, yeah, help me copilot. And it kicks it off.
I'm on my Windows, copilot, whatever. Apple intelligence, whatever it is, it gets kicked off to a data center. right? And that data center does some work and sends it back. That's inference. That is going to be the bulk of compute. But then, you know, and that's like, you know, there's thousands of data centers that we're tracking with like satellites and like all these other things.
And those are the bulk of what's being built. But the scale of... And so that's like what's really reshaping and that's what's getting millions of GPUs. But the scale of the largest cluster is also really important, right? When we look back at history, right? Like
Showing 1601–1620 of 1,814 · page 81 of 91
← Previous
Next →