Dwarkesh Patel
speaker
19,288 appearances
62 recordings
3 series
first heard Feb 2024
last heard 17 Sep
Dwarkesh Patel’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 27 in all, peaking in Jun 2026 with 6.
Appearances
Dwarkesh Podcast · Reiner Pope – The math behind how LLMs are trained and served · 29 Apr 2026
podcast
And so we don't want to be in a situation where the HBM is not big enough that we're not actually able to
keep write everything you want to it or take everything out of it.
Or we don't want to be in a situation where our ability to write back and forth is so big, or sorry, so small compared.
Yeah, makes sense.
Makes a ton of sense.
Okay, so a couple of actually quick questions.
One, if it is the case that the optimal batch size is something like 2000, and that actually true, it's totally dependent on sparsity.
It's not dependent on the model size or anything.
But that's a very interesting result.
And that seems to imply that you can...
One question is, how much of a push towards centralization is it that you would have these economies of scale from inference, from batching?
But it seems like it's not that big a deal.
Like, I don't know, is 2,000 users at the same time a lot?
It doesn't seem like a lot?
But I mean, Gemini is big.
That's actually one thousandth of Gemini is a lot.
To actually be like...
To be competitive at scale, you need to be able to serve at least 1,000 Go Gemini?
Yeah.
That's interesting.
Showing 4441–4460 of 19,288 · page 223 of 965
← Previous
Next →