Jonathan Ross

speaker
170 appearances 2 recordings 2 series first heard Jan 2025 last heard 20 Oct

Jonathan Ross’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · Oct OctJan 26AprJulnow

Recordings per month over the last 12 months — 1 in all, peaking in Oct 2025 with 1.

Appearances

newest first · ▶ plays the moment
And they'll have some of their own data and that'll make them subtly better at one thing or another. But they're largely all the same. More GPUs, the better the model because you can train on more tokens. It's the scaling law. This model was supposedly trained on a smaller number of GPUs and a much, much tighter budget.
I think the way that it's been put is less than the salary of many of the executives at Meta, and that's not true. There's an element of marketing involved in the DeepSea release. It is true that they train the model on approximately $6 million for the GPUs, right? They claim 2000
GPUs for, I think it was 60 days, which by the way, also don't forget was about the same amount of GPU time, 4,000 GPUs for 30 days as the original, I believe Lama 70. Now more recently, Meta has been training on more GPUs, but Meta hasn't been using as much good data as DeepSeq because DeepSeq was doing reinforcement learning using OpenAI.
Yes, exactly.
It's a little bit like speaking to someone who's smarter and getting tutored by someone who's smarter. You actually do better than if you're speaking to someone who's not as knowledgeable about the area or giving you wrong answers. First of all, before we get into any of this, I need to start with the scaling laws. These are like the physics of LLMs.
And there's a particular curve and the more tokens, which are sort of the syllables of an LLM, they don't match up exactly with human syllables, but kind of. So the more tokens that you train on, the better the model gets. But there's sort of these asymptotic returns where it starts trailing off.
The thing about the scaling law that everyone forgets, and that's why everyone was talking about how it's like the end of the scaling law, we're out of data on the internet, there's nothing left. What most people don't realize is that assumes that the data quality is uniform. If the data quality is better, then you can actually get away with training on fewer tokens.
So going back to my background, one of the fun things that I got to witness, I wasn't directly involved, was AlphaGo. Google beat the world champion, Lee Sedol, in Go. That model was trained on a bunch of existing games. But later on, they created a new one called AlphaGo Zero, which was trained on no existing games. It just played against itself. So how do you play against yourself and win?
Well, you train a model on some terrible moves. It does okay. And then you have it play against itself. And when it does better, you train on those better games. And then you keep leveling up like this, right? So you get better, better data. The better your model is when it outputs something, the better the result, the better the data.
So what you do is you train a model, you use it to generate data, and then you train a model and you use it to generate data and you keep getting better and better and better. So you can sort of beat the scaling law problem.
One quick hack to get past all of that in the stepping up is if there's a really good model already right here, just have it generate the data and you go right up to where it is. And that's what they did. It is true that they spent about six million or whatever it was on the training. They spent a lot more distilling or scraping the open AI model.
Correct. And all that said, they did a lot of really innovative things. That's what makes it so complicated, because on the one hand, they kind of just scraped the open AI model. On the other hand, they came up with some unique reinforcement learning techniques that are so similar. What did they do that was so impressive?
No, they came up with innovative stuff. But actually, the best way to describe it, have you ever taken a test before you got an answer right, and your professor marked it wrong. And then you go back to the professor and you have to argue with them and everything. And it's a pain, right?
Well, if there is only one answer, and it's a very simple answer, and you say, write that answer in this box, then there is no arguing. You either get it right or not, right? So what they did was, rather than having human beings check the output and say yes or no or whatever, what they did was they said, here's the box. There's literally some code to say here's a box.
I'll put the answer here and then check it. And if it's correct, we have the answer. If not, we don't. No need to involve a human. Completely automated. Can OpenAI not just do distillation on DeepSeq's model then? They don't need to because they're actually better still. They're a little bit better. They could, but why would they?
Or is that questionable doubt there? I don't think you have to disbelieve it because of the quality delta. However, why would they try and smuggle in GPUs when all they'd have to do is log into any cloud provider and rent GPUs? This is like the biggest gaming hole in the whole way that export control is done.
You can literally log in, you can swipe credit card, whatever, and just like pay and get GPUs to use. So export laws are necessary then? They're good, but the problem is it's like the Minagino line. You just go around it. So you need to like seal it up a little more. There's a little bit of room left to go here.
Keep in mind, OpenAI was effectively subsidizing accidentally the training of this model because they were using OpenAI, right? And rumors are that OpenAI may not be completely profitable yet in terms of every token in the API, like on the subscriptions maybe, but in the API.
And so each one that they generate, effectively, they were losing a little bit of money while DeepSeq was getting training data. Now, by the way, OpeningEye probably still has that data. In theory, they could just probably train on it.
I'm not aware of where it would be an export issue. I do know that many people log into cloud providers and just use them from remote. One of the problems, so we actually block IP addresses from China, and I believe we might be unique in doing that. It's also a little bit fruitless because someone could just like rent a server anywhere, log into us from there. Then there's nothing we can check.
Showing 41–60 of 170 · page 3 of 9 ← Previous Next →