Dwarkesh Patel
speaker
19,288 appearances
62 recordings
3 series
first heard Feb 2024
last heard 17 Sep
Dwarkesh Patel’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 27 in all, peaking in Jun 2026 with 6.
Appearances
Dwarkesh Podcast · Reiner Pope – The math behind how LLMs are trained and served · 29 Apr 2026
podcast
Maybe you get a smaller model, you spend more computer training than you otherwise would have, but now it's cheaper to give it to users.
Let me make the question more concrete.
How much more than chinchilla optimal are models overtrained?
And has that changed as a result of RL generation?
Which is the fact that you're not training on all your rollouts.
Okay, so if you're doing a backward pass on every single generation in RL, it would be 6 nd.
Yeah, so this could be a smaller number, right?
I think the way I said it was super garbled.
Just for the audience, maybe.
Forward plus backwards per parameter is six.
Forward alone is two.
That's why RL where you might... You're definitely going to generate all the trajectories, but you might or might not train all the trajectories is two to six.
Yes.
Yeah.
And inference would be 50%.
If both of them are 1 in 10, that kind of implies that there's never a backward pass on RL?
So this is like 1.5 and this is one, um, um, Billions of dollars of the compute just flowed the other direction.
Right.
But then, so it looks... Sorry, I'm making a basic algebra mistake.
It seems like there should be less RL tokens than pre-training tokens?
Showing 4641–4660 of 19,288 · page 233 of 965
← Previous
Next →