Dylan Patel

speaker
3,025 appearances 3 recordings 1 series first heard Aug 2025 last heard 15 Jul

Dylan Patel’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · Jul OctJan 26AprJulnow

Recordings per month over the last 12 months — 1 in all, peaking in Jul 2026 with 1.

Appearances

newest first · ▶ plays the moment
I have no doubt that Cerebrus would run certain types of models better than NVIDIA or Grok.
Or, hey, Dojo, right?
Dojo runs certain, you know, Tesla's Dojo would run certain types of models way better than NVIDIA's chips because they're optimized to that.
But then it's like, oh, well, actually, even in vision tasks, you use vision transformers now.
So it's like, okay, cool.
Because model sizes grew and all these things, so it ends up being
a, you know, catch-22 in that, like, you optimize for something.
And so now, like, today you have this new age of AI accelerator companies that are like, okay, we're going to optimize for transformers.
But the time they started designing, they're like, okay, transformers are dense models that are this big.
What's the best, you know, the hidden dimension is 8K and your batch sizes are this big and your sequence lengths are this big, so let's just make a super large systolic array.
so you can create the maximum efficiency and then it turns out, oh, look at DeepSeq or go look at what the labs are doing.
Actually, their shapes are much smaller.
Actually, you need to do a bunch of small matrix multiplies, not massive, massive, massive, singular matrix multiplies per layer.
And then it ends up, you know, oh, well, that chip you're designing for that is actually not super effective for that.
And so the software is evolving constantly because of what works best on NVIDIA.
And you see that with, you know, whether it be what DeepSeq's doing or Alibaba's doing or what the labs are doing internally.
And you even see this, like, for Google, right?
Like, their open-source Gemma models make different decisions because the shapes of a TPU are different than a GPU, right?
And those, the GPU and the TPU are actually not that far apart, right?
Like you would say, yes, they're very different, but like Blackwell and TPUs are very, very, they're converging on similar designs actually.
Showing 301–320 of 3,025 · page 16 of 152 ← Previous Next →