Rob Wachen
speaker
529 appearances
1 recordings
1 series
first heard Jun 2026
last heard 30 Jun
Rob Wachen’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 1 in all, peaking in Jun 2026 with 1.
Appearances
And that informs a lot about the bets we're making in the future with the low voltage inference and the cluster scale memory.
But the production part, I think, in the past year has become extremely obvious.
How much people want to deploy this stuff if you can have it available today?
The best ability is availability.
If I have a thousand chips today, someone's going to use them.
And, you know, we need to build a chip that's not just like way better than what's been built before, but it needs to be available at many gigawatts scale.
We need to be able to be building a product that is producible at gigawatts per month from the limit.
As we think about that, a lot of the design decisions we're making with our next gen, which you've seen already, is just about simplicity.
Removing tons of parts, trying to assemble and disassemble a thing again and again, and learning how to make it as quick in the cycle times as possible in production, making sure it's going to be reliable, making sure it's going to be serviceable, and making sure it's going to be producible at gigantic scales.
The people deploying the most compute in the world do think about supply a bit zero sum, which is there's only so many wafers being produced on a given nanometer node on a given fab, right?
And there's only so much memory being produced.
And that's why actually for our first gen product, we built it on a different supply chain than the Rubens.
We're on four nanometer Rubens on three nanometer.
Yeah, we're on a different HPM than Rubens and so forth.
So it actually is not a zero-something.
It's a positive something where more is more.
So often when we're talking to people deploying at scale, it's not a decision between a gigawatt of a GPU and a gigawatt of us.
It's two gigawatts.
And I think as much as possible, thinking about supply chain early in the design decisions, because if you have the most performant product and you can't produce it, then you're just a podcast.
A theme in models right now is this focus on something called dynamism, which is this ability to control the level of computation and memory spent at a per-token or per-user level when doing attention, as well as this ability to dynamically, in your chip, on the fly, send data to other chips for different MLE models doing certain types of operations.
Showing 461–480 of 529 · page 24 of 27
← Previous
Next →