Dwarkesh Patel
speaker
19,288 appearances
62 recordings
3 series
first heard Feb 2024
last heard 5d ago
Dwarkesh Patel’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 27 in all, peaking in Jun 2026 with 6.
Appearances
And a final point of comparison, a teenager can learn to drive a car with about 20 hours of practice.
And even if we include their 16 years of growing up and understanding how the world works and building physical intuition, there's still three to four orders of magnitude less data than Waymo and Tesla are using to train their self-driving car models.
Now I want to deal with a couple of common responses and objections that people have to these kinds of comparisons.
One thing people will say, and I think Karpathy said this when he came on to my podcast, is that for humans, many billions of years of evolution had to go into basically pre-training us.
And so we're being unfair when we're comparing how little data we see within our lifetimes to what these cold-started LLMs, who are just starting off with a totally random initialization, have to learn from.
I think this is not the right way to think about it.
Our genome is only three gigabytes big, and only 1% to 2% of it is protein coding.
And that is simply not enough space to store the parameters of this network that supposedly evolution has pre-trained.
I think the closer analogy is more that evolution found the right hyperparameters and the right loss functions.
And that within our lifetime, we are still from scratch building up the connectome in our brain.
That is to say, the analogous thing to the weights and parameters of the neural network itself.
And even if you granted this comparison and you said, yes, the hundreds of trillions of tokens that these models see to get pre-trained is similar to just catching up to evolution, that still doesn't explain why any new marginal capability that you want to give these models takes so much data.
So once you have been educated, again, you don't need 100 different professors to teach you how to learn a new programming language.
But these AIs, even once they're pre-trained, still require enormous amounts of data to learn the next marginal skill and the next marginal skill after that.
Another objection to this kind of comparison is that we're not including multimodal data that we're seeing in our lifetimes.
So we include all this sensor information that we see from birth to adulthood.
That's probably tens to hundreds of billions of tokens of data.
And my response to this objection is simply that blind and deaf people who have been cut off from all the sensor information still have general intelligence.
And that suggests to me that all these billions of sensory tokens are not really the thing that is making humans smart.
And in fact, deaf people who don't have the ability to hear any tokens, who just have to consume them via sign language and reading, are probably ingesting far less than the 200 million language tokens that we ballparked earlier, which suggests that even the million-fold difference that we calculated earlier might be an understatement.
Showing 3541–3560 of 19,288 · page 178 of 965
← Previous
Next →