Dwarkesh Patel
speaker
19,288 appearances
62 recordings
3 series
first heard Feb 2024
last heard 17 Sep
Dwarkesh Patel’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 27 in all, peaking in Jun 2026 with 6.
Appearances
Dwarkesh Podcast · Reiner Pope – The math behind how LLMs are trained and served · 29 Apr 2026
podcast
And is there something especially significant about the slope being exactly the slope of the...
The compute time?
But suppose it's like...
This is a very simple algebra problem, but suppose the optimal is 100k context length.
And you go to 200k context length.
Does your MFU go down to like 50%?
Does it have a humongous impact on MFU?
Yeah, it does.
To be like slightly outside of context length, optimal range, Goldilocks zone.
Got it.
And is sparse attention what everybody uses in practice?
So Claude code slow or codex slow or whatever would just live on this line and it wouldn't help much because you're not able to amortize the KV values over a much bigger batch.
So this point where you are no longer memory bandwidth bound,
How big a batch do you need?
How big are the batches practically for Frontier models?
Sorry, has that ratio changed over time as we've gone from model generation to model generation where the flops keeps increasing?
Okay, so basically it's like 2,000 to 3,000 tokens per batch.
But then if you included the KB cache, the implication would be that the optimal batch size should grow larger.
This seems incredibly small.
Like a batch, this would be like less than one sequence, right?
Showing 4401–4420 of 19,288 · page 221 of 965
← Previous
Next →