Dwarkesh Patel

speaker
19,288 appearances 62 recordings 3 series first heard Feb 2024 last heard 17 Sep

Dwarkesh Patel’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
6 · Jun OctJan 26AprJulnow

Recordings per month over the last 12 months — 27 in all, peaking in Jun 2026 with 6.

Appearances

newest first · ▶ plays the moment
And is there something especially significant about the slope being exactly the slope of the...
The compute time?
But suppose it's like...
This is a very simple algebra problem, but suppose the optimal is 100k context length.
And you go to 200k context length.
Does your MFU go down to like 50%?
Does it have a humongous impact on MFU?
Yeah, it does.
To be like slightly outside of context length, optimal range, Goldilocks zone.
Got it.
And is sparse attention what everybody uses in practice?
So Claude code slow or codex slow or whatever would just live on this line and it wouldn't help much because you're not able to amortize the KV values over a much bigger batch.
So this point where you are no longer memory bandwidth bound,
How big a batch do you need?
How big are the batches practically for Frontier models?
Sorry, has that ratio changed over time as we've gone from model generation to model generation where the flops keeps increasing?
Okay, so basically it's like 2,000 to 3,000 tokens per batch.
But then if you included the KB cache, the implication would be that the optimal batch size should grow larger.
This seems incredibly small.
Like a batch, this would be like less than one sequence, right?
Showing 4401–4420 of 19,288 · page 221 of 965 ← Previous Next →