Dwarkesh Patel

speaker
19,288 appearances 62 recordings 3 series first heard Feb 2024 last heard 5d ago

Dwarkesh Patel’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
6 · Jun OctJan 26AprJulnow

Recordings per month over the last 12 months — 27 in all, peaking in Jun 2026 with 6.

Appearances

newest first · ▶ plays the moment
Whereas a human student might practice a textbook problem once or twice.
With GRPO, these models are generating hundreds to thousands of rollouts per task, and they need to solve the credit assignment problem.
The correct way to think about these models is not like a human who has learned all these different skills that you see these models displaying.
It's more like a Frankenstein's monster which has been built out of a billion graphs of carefully constructed examples all sewn together.
Epoch recently reported that open models lag state-of-the-art frontier models by four months.
I think the reason it is relatively easy for open source and previous laggards to catch up to within months of the frontier is that data is the real driver of progress.
And data can be easily distilled from public APIs, whereas hyperparameters and training tricks and architectural optimizations cannot.
And if the latter were driving most of the progress, then catching up would be far harder than we are observing it to be.
It is easy to forget how much data these models are trained on and how much more it is than what we humans see in our lifetimes.
We see these AIs as a galaxy glittering with capabilities, but at their center, invisible to the naked eye, holding all the constellations together is an unimaginably massive black hole of data.
Just a couple of points of comparison to help drive home how big this difference is.
Here's one.
If a person sees and hears on average, let's say generously, 2000 words an hour, then between the time they're born and the time they're an adult, they'll see about 200 million tokens.
Now, by contrast, these frontier models are trained on somewhere between tens to hundreds of trillions of tokens.
That is close to a million fold difference.
Here's another point of comparison.
If you wanted to, you could learn to tell or operate any random humanoid or robot arm within hours.
And if we could get AIs to learn just as fast, robotics would be a deca-trillion-dollar industry, and you'd have an endless army of unitary G1s doing all kinds of useful work in the world.
But the reason we can't do this is that our AIs learn much less efficiently than we do.
And even with the millions of hours of demonstrations that we collected, this is not enough to allow them to perform complex open-ended tasks.
Showing 3521–3540 of 19,288 · page 177 of 965 ← Previous Next →