Dwarkesh Patel
speaker
19,288 appearances
62 recordings
3 series
first heard Feb 2024
last heard 17 Sep
Dwarkesh Patel’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 27 in all, peaking in Jun 2026 with 6.
Appearances
Dwarkesh Podcast · Reiner Pope – The math behind how LLMs are trained and served · 29 Apr 2026
podcast
Yep.
When you're operating within a single scale-up domain, is that a consideration specifically for either forward or backward?
Or specifically for pre-fill versus decode?
Or is it preferred to always be within a scale-up
Whatever kind of workload you have, whether you're doing a pre-training run or whether you're doing RLL generation or whether you're doing inference for users.
Can I try to guess?
Just out of curiosity to see if I'm actually understanding.
It seems like you're sending batch size into the rack.
In here?
Yes.
But the communication within a rack is sort of batch size times number of GPUs.
And there's a need to multiply the whole thing by two for the up and down.
And there's a factor of two.
It's interesting to me that the best parallelism
strategy and practice ends up being one which physically resembles the actual architecture.
It's not some galaxy brain thing.
You know, it's like, oh, we have experts.
We're going to put them on different GPUs.
Oh, we have different layers.
We're going to put them on different racks.
Showing 4521–4540 of 19,288 · page 227 of 965
← Previous
Next →