Rob Wachen

speaker
529 appearances 1 recordings 1 series first heard Jun 2026 last heard 30 Jun

Rob Wachen’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · Jun OctJan 26AprJulnow

Recordings per month over the last 12 months — 1 in all, peaking in Jun 2026 with 1.

Appearances

newest first · ▶ plays the moment
They say, are you pre-filled chip?
Are you decode chip?
If you're decode chip, are you an HPM chip?
Are you an SRAM chip?
Are you 3D DRAM chip?
Are you using optics, using copper?
When we started this, we just wanted to understand why extremely smart people were working on these different directions.
We seriously looked at architectures like having a bunch of DDR memory and like a shared memory pool and looking at advanced packaging to basically break out of the shoreline.
We looked at things like, is there ways to put memory dies on top of compute dies?
In doing so, we realized that there's no free lunch.
Everything has a trade-off, right?
3D DRAM, you have a thermal issue, you have a supply chain issue, you have to figure out hybrid bonding, you have to figure out the flops, so now you're a decode chip.
So we went through everything, both on the pre-fault and the decode side.
In doing so, we realized there's a few design spaces that nobody had seriously tried to explore,
because they were never done in AI chips.
And we asked ourselves, what are the actual metrics that are going to matter the most?
On the pre-fill side, the thing that matters is flops and flops density.
And people talk about flops often as a headline number, but in reality, you should care about the flops you're getting when you're running real workloads.
There's this concept called MFU, or model flops utilization, which is for every peak flop advertised, how many cents on the dollar are you actually getting?
And on GPUs, you often get somewhere between 20% and 50%, depending on the workload.
Showing 21–40 of 529 · page 2 of 27 ← Previous Next →