David Cox

speaker
634 appearances 1 recordings 1 series first heard Jan 2026 last heard 13 Jan

David Cox’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · Jan OctJan 26AprJulnow

Recordings per month over the last 12 months — 1 in all, peaking in Jan 2026 with 1.

Appearances

newest first · ▶ plays the moment
Yeah.
And and that's you know, that's kind of a that's kind of a drag.
The other thing, which is a little more subtle, but it's even more important, is the longer that context gets and the more of this KV cash stuff you have, the slower it is to generate that next token.
Like, yeah, it's actually like it goes as the square of the number of tokens you have.
So it just gets, you know, they call it, you know, O.N.
squared or quadratically scaling quadratically.
So, you know, pretty soon you're like, oh, this is going to take like it's it's like the robot slowing down because it's got mud in it or something.
So one thing that people have done is they try to say, well, what if we didn't have this this KV cache, at least not in every layer?
Like maybe we don't need it ever.
Maybe we have something that doesn't keep any sort of trace.
It's just keeping up and trying to do its own best.
And then you end up with models that have like some of the layers have this KV cache thing because it's really useful.
It's really powerful, but it's just.
It's a drag because it's slow and it consumes all your memory.
Maybe we don't need to have it on every layer.
And then we figured out technology is basically for like having, you know, every nth layer has it and every other layer doesn't have it.
And the upshot, if like none of that matters, like the technology, you don't care.
You just want to like, you know, TLDR.
is like 10 times smaller memory footprint for the KV cache.
And then we just like, you know, the contribution that was slowing it down is now 10 times less as well.
Showing 181–200 of 634 · page 10 of 32 ← Previous Next →