Rob Wachen
speaker
529 appearances
1 recordings
1 series
first heard Jun 2026
last heard 30 Jun
Rob Wachen’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 1 in all, peaking in Jun 2026 with 1.
Appearances
And for certain products, it's really fast and certain products, it's really slow.
Whatever the speed is, this is my speed.
The question is, in a given amount of power, how many users can I serve while guaranteeing that speed?
So another way to put it is ISO, what's called interactivity, what is my throughput?
We are just finishing kind of the early innings of the AI infrastructure boom where people really just cared about speed.
You know, GPUs were not able to reach a lot of the speeds of other types of chips like all these SRAM chips, thousands of tokens per second, and that enabled tons of new use cases that got people very excited.
There's an entirely new wave of AI chips, us being one of them, that are all going to be able to hit these speeds.
The question then is, if you're hitting these speeds, what is the number of users you can serve at the same time?
And by proxy, if I have a 100 megawatt data center, how many software agents can I run at the same time?
So when people are doing that evaluation, our hardware is going to generally be able to get you an order of magnitude more concurrency at a given level of interactivity.
So that directly translates into tokens per watt, tokens per dollar, all the things people care about when they're actually serving these giant mixture of expert models at scale.
I think too often people think about tasks and applications and stuff in these very short time horizons.
Doing a chat and it's like 50% faster is nice, but it's not like game changing.
As these agents go longer and longer time horizon and the models get more and more capable, you're going to see gigantic bodies of work that would take months of compute.
And we think about this in wall clock time.
Like if you talk to a pre-training researcher at a lab, they'll tell you that wall clock time often is one of the most important things that matters.
And what wall clock time means is the time from starting your run to finishing it to actually get data back.
If you can shrink this time from a six-month run to a two-month experiment, you're going to be able to do many more iterations and people will make changes on the model architectures to actually improve the wall clock time.
Very similar here in terms of how we think about the use cases, which is the exciting part about super low latency decode is wall clock time on long horizon tasks becomes much shorter.
So a year long compute build would now take months and a month long compute build will now take three days and that three day compute build will now take seven hours and so forth and so forth.
Showing 281–300 of 529 · page 15 of 27
← Previous
Next →