Rob Wachen
speaker
529 appearances
1 recordings
1 series
first heard Jun 2026
last heard 30 Jun
Rob Wachen’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 1 in all, peaking in Jun 2026 with 1.
Appearances
We go through our first wafer.
It's like 2, 3 a.m.
because we're doing it with TSMC over the phone in Taiwan.
We have the screen with the wafer that's all gray, and each ship is gray, and then as you start running the patterns, the squares are supposed to turn green or red, and they all turn red.
We're like, fuck.
This is really bad.
Everybody's like, guys, take a breath.
He leans back, he's like, the puzzle begins.
You have to have the attitude.
Within a day of that.
But in the moment, it's extremely scary.
And there's a certain type of person who just like is addicted to that feeling of just feeling the fear and solving it.
And we are lucky to have a lot of those people here.
It took us a while to get to the primitives that we think are really what matters for scaling inference.
We tried a bunch of things early on, from compilers that would turn different models into FPGAs, to burning weights in silicon, to splitting your HBM to KV cache and weights, and all of these different things.
And there was a lot of cycles of learning until we got to the point that we realized that fundamentally, if you want to run a majority of tokens in the world, you need to do three things.
You need to build a chip with the most flops and a given power budget.
You need to build a chip that has the lowest latency between other chips, so the biggest scale-up domain possible.
And you need to produce as much of it as possible.
And I think probably in the first half of our journey so far, we learned the first two, and that informed the design a lot.
Showing 441–460 of 529 · page 23 of 27
← Previous
Next →