Jonathan Ross
speaker
170 appearances
2 recordings
2 series
first heard Jan 2025
last heard 20 Oct
Jonathan Ross’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 1 in all, peaking in Oct 2025 with 1.
Appearances
How do you think about that? What you see is a bunch of people who are concerned about training and the need for it. And everyone's still thinking that most of compute is training. And that there's going to be less of it because someone trained a model on 2000 GPUs and the nerfed A800 version with slower memory or whatever it is. And they're like, oh, people aren't going to need as many chips.
But again, Jevin's paradox, right? The more you bring the cost down, the more people consume. So for the last five to six decades, like clockwork, once a decade, the cost of compute has gone down 1,000x. People buy 100,000x as much compute, spending 100 times as much. So every decade, they spend 100 times as much. So you make it cheaper, they want more.
What's really happening is every time one of these models gets cheaper, we see our developer count just skyrocket. And then it comes back down a little bit, but the slope is higher than when it started. Better models create more demand for inference. More demand for inference then has people going, I should train a better model. And the cycle continues.
So I think over the long term, the only thing I say is Warren Buffett and Charlie Munger in the short term, the market is a popularity contest. In the long term, it's a weighing machine. I can't tell you about the popularity contest, but in terms of the weighing machine part, this is a misunderstanding. It's actually more valuable thanks to deep seek, not less valuable.
Okay, so Jevin's paradox was actually discovered by Jevin as recently made famous in Satya's tweet. However, I did beat him to that by quite a bit. And just as Satya likes to say that he made Google dance, I'm going to say I made Satya dance. He might take exception to that. But less than a month before he posted that, I did a cute little tweet on it.
So what's really happening here was in the 1860s, this guy Jevin, he actually wrote a treatise on steam engines, which I guess is what you did for fun back then in England. He realized every time steam engines became more efficient, people would buy more coal, which is the paradox.
But if you think about it from a business point of view, when the OPEX comes down, more activities come into the money. So people do more things. And so what's happened is every time we've seen the cost of tokens for a particular level of quality of models come down, We've actually seen the demand grow significantly. Price elasticity, baby.
Today, there's this wonderful business selling mainframes with a pretty juicy margin because no one seems to want to enter that business. Training is a niche market with very high margins. And when I say niche, it's still going to be worth hundreds of billions a year. But inference is the larger market. And...
I don't know that NVIDIA will ever see it this way, but I do think that those of us focusing on inference and building stuff specifically for that are probably the best thing that's ever happened for NVIDIA stock because we'll take on the low margin, high volume inference so that NVIDIA can keep its margins nice and high.
No. And I was actually like, we raised some money late 2024. In that fundraise, we still had to explain to people why inference was going to be a larger business than training. Remember, this was our thesis when we started eight years ago. So for me, I struggle on why people think that training is going to be bigger. It just doesn't make sense.
Training is where you create the model. Inference is where you use the model. You want to become a heart surgeon, you spend years training, and then you spend more years practicing. Practicing is inference.
what you're going to see is everyone else starting to use this MOE approach. Now, there's another thing that happens here.
Yeah, so MOE stands for mixture of experts. When you use LAMA 70 billion, you actually use every single parameter in that model. When you use Mixtrals 8x7b, you use two of the roughly 8b experts, but it's much smaller. And effectively, while it doesn't correlate exactly, it correlates very closely. The number of parameters effectively tells you how much compute you're performing.
Now, if I have, let's take the R1 model. I believe it's about 671 billion parameters versus 70 billion for LAMA. And there's a 405 billion dense model as well, right? But let's focus on 70 versus 671. I believe there's 256 experts, each of which is somewhere around 2 billion parameters.
And then it picks some small number, I'm forgetting which, maybe it's like eight of those or 16 of them, whatever it is. And so it only needs to do the compute for that. That means that you're getting to skip most of it, right? Sort of like your brain, like not every neuron in your brain fires when I say something to you about the stock market, right?
Like the neurons about, you know, playing football, right? those don't kick off, right? That's the intuition there. Previously, it was famously reported that OpenAI's GPT-4, it started off with something like 16 experts and they got it down to eight. I forget the numbers, but it started off larger and they shrunk it a little and they were smaller or whatever.
And then with what's happened with DeepSeq model is they've gone the opposite. They've gone to a very large number of experts. The more parameters you have, it's like having more neurons. It's easier to retain the information that comes in. And so by having more parameters, they're able to, on a smaller amount of data, get good.
However, because it's sparse, because it's a mixture of experts, they're not doing as much computation. And part of the cleverness was figuring out how they could have so many experts so it could be so sparse so they could skip so many of the parameters.
outperformed their 405. What was surprising to me, I thought they retrained it from scratch. It turns out you read the paper and they talk about how they just fine tuned. So they used a relatively small amount of data to make it much better. Again, this goes to the quality of the data. They have higher quality data. They took their old model. They trained it, got much better.
But that 70B, that new 70B outperforms their previous 405B. What you're going to see now is now that everyone has seen this deep seek architecture, they're going to go, great, I have hundreds of thousands of GPUs. I'm now going to use a lot of them to create a lot of synthetic data. And then I'm going to train the bejesus out of this model.
Showing 121–140 of 170 · page 7 of 9
← Previous
Next →