Dwarkesh Patel

speaker
19,288 appearances 62 recordings 3 series first heard Feb 2024 last heard 17 Sep

Dwarkesh Patel’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
6 · Jun OctJan 26AprJulnow

Recordings per month over the last 12 months — 27 in all, peaking in Jun 2026 with 6.

Appearances

newest first · ▶ plays the moment
Yep.
When you're operating within a single scale-up domain, is that a consideration specifically for either forward or backward?
Or specifically for pre-fill versus decode?
Or is it preferred to always be within a scale-up
Whatever kind of workload you have, whether you're doing a pre-training run or whether you're doing RLL generation or whether you're doing inference for users.
Can I try to guess?
Just out of curiosity to see if I'm actually understanding.
It seems like you're sending batch size into the rack.
In here?
Yes.
But the communication within a rack is sort of batch size times number of GPUs.
And there's a need to multiply the whole thing by two for the up and down.
And there's a factor of two.
It's interesting to me that the best parallelism
strategy and practice ends up being one which physically resembles the actual architecture.
It's not some galaxy brain thing.
You know, it's like, oh, we have experts.
We're going to put them on different GPUs.
Oh, we have different layers.
We're going to put them on different racks.
Showing 4521–4540 of 19,288 · page 227 of 965 ← Previous Next →