Dwarkesh Patel

speaker
19,288 appearances 62 recordings 3 series first heard Feb 2024 last heard 17 Sep

Dwarkesh Patel’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
6 · Jun OctJan 26AprJulnow

Recordings per month over the last 12 months — 27 in all, peaking in Jun 2026 with 6.

Appearances

newest first · ▶ plays the moment
Cool.
Okay, so... The more sparsity you have, the less compute you need,
And it does seem that as batch sizes get bigger, compute ends up being the bottleneck, according to this analysis.
So then the question is, how far can you take sparsity?
That is to say, as the sparsity ratio increases, as you have fewer and fewer active parameters relative to total parameters, how much is performance of the model degrading?
And is it degrading faster than you're saving compute by increasing the sparsity factor?
Should we pull up the paper now?
10x as many active parameters.
Yeah, so while it is true, I guess, that you get this benefit of being able to economize on your compute time if you increase sparsity,
Naively, it would seem like, oh, that's a trade-off worth making.
But if you're decreasing this by 2x and then having this go up by 8x, every time you double...
So let me just make sure I understood.
You're saying we want bigger... We want... Does it mean less time computing?
Therefore, we do more sparsity.
To make that work, we need bigger batch sizes, which means we need more memory capacity.
Yeah, so... To have more sparsity.
So when you say any GPU in the pretense, the router is more than one GPU?
Yeah.
Before we... It may be worth you explaining...
What exactly a rack is, the differences in bandwidth between a rack and within a rack, and the all-to-all versus not-all-to-all nature of communication within versus outside.
Showing 4461–4480 of 19,288 · page 224 of 965 ← Previous Next →