Micah-Hill Smith

speaker
601 appearances 1 recordings 1 series first heard Jan 2026 last heard 8 Jan

Micah-Hill Smith’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · Jan OctJan 26AprJulnow

Recordings per month over the last 12 months — 1 in all, peaking in Jan 2026 with 1.

Appearances

newest first · ▶ plays the moment
We've got reasoning models using tokens, and then we're throwing them in these them in these agentic workflows where they're consuming enormous numbers of input tokens and making enormous numbers of output tokens, working for a really long time.
Those two things taken together get you back to we can spend enormously more today than we could a couple of years ago.
Okay.
So the the the the answer unfortunately uh is it depends and it just depends massively on like so many things across a bunch of different types of workloads and ways to think about it.
So one of the simplest ways to think about
This is to take single relevant model, to think about serving it at speeds that are realistic for what you actually might want to hit and can afford to hit, and then think about the throughput per GPU that you can achieve, serving the model at those speeds.
Reason one of the reasons that's important is that there's a trade off between the throughput per GPU that you can achieve and the per user speed that you can achieve.
And as in it costs more to serve stuff fast to users.
When you run all of that, for especially big sparse models, you can get a lot better than two or three eggs gain going from hopper to blackwell generation to video.
I am
This shouldn't be too controversial to say, but like I'm pretty confident that
Blackwell has delivered pretty enormous gains and that the next couple of years of Nvidia's roadmap are going to continue to deliver quite enormous gains and that those will actually come through as lower total cost per token to the companies that are running models on them and will allow bigger models, will allow way more tokens to be made for lower cost, and that that's gonna continue.
These things also
stack on all of the software and model improvements being made.
So basically like my prediction across like both sides of that like smile chart are that we're gonna see the left hand side continue to be true and probably like for another order of magnitude and the right hand side continue to be true for another order of magnitude.
And that's gonna enable a whole lot of things.
Um then there must be a limit somewhere, right?
Yeah, exactly.
But we've got numbers in the wild that are quite a lot lower than that right now.
So the GPD OSS models, like the big ones at about five percent um active, Kimik2 is at like three percent active.
Showing 481–500 of 601 · page 25 of 31 ← Previous Next →