Rob Wachen

speaker
529 appearances 1 recordings 1 series first heard Jun 2026 last heard 30 Jun

Rob Wachen’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · Jun OctJan 26AprJulnow

Recordings per month over the last 12 months — 1 in all, peaking in Jun 2026 with 1.

Appearances

newest first · ▶ plays the moment
And actually, you can provably not run at 100% because you have a thermal issue, where as you increase the flop utilization, you have more transistors going on and off, you draw more power, and the chip will self-regulate and actually lower its clock speed to make sure it doesn't overheat.
As we looked at inference, we said if we want way more flops because we want to run at way higher throughputs, we fundamentally need to solve the thermal problem before we even think about adding flops to the chip.
If I just add more flops to a GPU today or another AI chip, I'm not actually going to get more performance because it's just going to thermal throttle.
Fundamentally, the essence of that is this concept of Dennard scaling, which is voltage is quadratically proportional to power.
So if I 2x my voltage, my power goes up by 4x.
If I cut my voltage in half, my power goes down by a quarter.
So we asked ourselves, how could we run voltages lower than GPUs?
And we talked to a lot of people about this.
We flew out to Silicon Valley after dropping out and basically asked dozens of people and semiconductors and all these different chip companies how they did it.
And the answer we got was like, you can't.
You can't run at voltages lower than GPUs.
And this was very dissatisfying because there was many different industries of chips that run at voltages lower than GPUs.
Bitcoin miners run at under a quarter of the voltage of GPUs.
So this is obviously physically possible.
The question is, are there issues with GPU architectures that make it unable to run at these voltages?
And when we looked at the problem for a long time, we were able to create a new mechanism of running at much lower voltages, a new type of power delivery that we call low voltage inference.
And we think all AI chips in the future are going to be low voltage chips.
They're going to have to cram way more flops in the same silicon area and without thermal throttling run at way lower voltages.
Yeah, and it's not that surprising given all these architectures were built before ChatGPT.
So if we're trying to build a chip for the modern workloads, it's going to look very different.
Showing 41–60 of 529 · page 3 of 27 ← Previous Next →