Rob Wachen

speaker
529 appearances 1 recordings 1 series first heard Jun 2026 last heard 30 Jun

Rob Wachen’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · Jun OctJan 26AprJulnow

Recordings per month over the last 12 months — 1 in all, peaking in Jun 2026 with 1.

Appearances

newest first · ▶ plays the moment
And the reason is fundamentally as we are scaling context length, as we're scaling model size, as we're scaling the amount of computation per user, we're looking for ways to be more efficient.
So the first thing is, like Gavin says, a mixture of experts architectures where maybe we don't need every parameter being used for every token.
But maybe there are things where even at a token level, we can say, well, this token needs this context from this other token.
They can share that memory, so we don't have to have overhead of using the memory as much.
Maybe this token is really important, so we should spend more compute.
We should have longer context on that token.
So hardware that really accelerates these types of very dynamic computations is extremely important.
And you can imagine current hardware that was designed before those types of architectures have lots of overheads in doing them.
So you basically end up in these really bad worlds where you have inefficient hardware at doing this dynamism, so therefore you can't run it very well, or you have these very blocky architectures that are kind of applying blunt force to many different tokens that all need more or less computation.
I'm going to be a little futuristic.
I firmly believe we are on a global march of inference becoming majority of global GDP.
It may take more than 10 years, but it's going to happen.
And right now we measure productivity as a society as GDP per capita.
So really, it's going to look much more like agents per megawatt or maybe agents per gigawatt by then.
And while we're being futuristic, I think this is the second to last year where a majority of the workforce is going to be human.
I think in 2027, you're going to see there's going to be more agents doing knowledge work than humans.
And it's going to be extremely interesting to see what happens.
You could imagine a world where, for countries, a majority of their energy ends up going into data centers doing inference.
And the energy efficiency of those data centers basically governs how many agents and therefore how big their workforce is.
So you're going to see, like as Gavin is saying, one agent or a team of five to ten agents working on group projects for a couple of days.
Showing 481–500 of 529 · page 25 of 27 ← Previous Next →