Nathan Lambert
speaker
1,814 appearances
3 recordings
2 series
first heard Feb 2025
last heard 1 Feb
Nathan Lambert’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 2 in all, peaking in Feb 2026 with 1.
Appearances
We're going to get into technical stuff real fast. There's two articles in this one that I could show, maybe graphics that might be interesting for you to pull up.
You want to explain KVCache before we talk about this? I think it's better to... Okay, yeah.
Because it's incredibly important because this changes how models work. But I think resetting, right? Why is memory... so important. It's because so far we've talked about parameter counts, right? And mixture of experts, you can change how many active parameters versus total parameters to embed more data, but have less flops.
But more important, you know, another aspect of, you know, what's part of this humongous revolution in the last handful of years is the transformer, right? And the attention mechanism. Attention mechanism is that the model understands the relationships between all the words in its context, right? And that is separate from the parameters themselves. right?
And that is something that you must calculate, right? How each token, right, each word in the context length is relatively connected to each other, right? And I think, Nathan, you should explain KVCache better.
I can explain that. So today, if you use a model, like you look at an API, OpenAI charges a certain price per million tokens, right? And that price for input and output tokens is different, right? And the reason is that when you're inputting a query into the model, right? Let's say you have a book, right? That book, you must now calculate the entire KV cache for it, right? This key value cache.
And so when you do that, that is a parallel operation. All of the tokens can be processed at one time. And therefore, you can dramatically reduce how much you're spending, right? The flop requirements for generating a token and an input token are identical, right? If I input one token or if I generate one token, it's completely identical. I have to go through the model.
But the difference is that I can do that input, i.e. the pre-fill, i.e. the prompt, simultaneously in a batch nature. And therefore, it is all flopped.
Correct. But then output tokens, the reason why it's so expensive is because I can't do it in parallel, right? It's autoregressive. Every time I generate a token, I must not only read the whole entire model into memory and activate it, calculate it to generate the next token, I also have to read the entire KV cache.
and I generate a token, and I append that KV, that one token I generated, and it's KV cash, and then I do it again, right? And so therefore, this is a non-parallel operation. And this is one where you have to, you know, in the case of pre-fill or prompt, you pull the whole model in and you calculate 20,000 tokens at once, right?
i.e. how many tokens are being generated slash prompt, right? So if I put in a book, that's a million tokens, right? But, you know, if I put in, you know, the sky is blue, then that's like six tokens or whatever.
It's mostly output tokens. So before, you know, three months ago, whenever O1 launched, all of the use cases for long context length were like, let me put a ton of documents in and then get an answer out, right? And it's a single, you know, Pre-fill, compute a lot in parallel, and then output a little bit. Now, with reasoning and agents, this is a very different idea, right?
Now, instead, I might only have like, hey, do this task, or I might have all these documents. But at the end of the day, the model is not just like producing a little bit, right? It's producing tons. Tons of information, this chain of thought just continues to go and go and go and go.
And so the sequence length is effectively that, you know, if it's generated 10,000 tokens, it's 10,000 sequence length, right? Or, and plus whatever you inputted in the prompt. And so what this chart is showing, and it's a logarithmic chart, right? Is, you know, as you go from 1K to 4K or 4K to 16K, the memory requirements grow so fast
64 different users at once, right? Yeah. And therefore your serving costs are lower, right? Because the server costs the same, right? This is eight H100s, roughly $2 an hour per GPU. That's $16 an hour, right? That is like somewhat of a fixed cost. You can do things to make it lower, of course, but like it's like $16 an hour. Now, how many users can you serve? How many tokens can you generate?
And then you divide the two and that's your cost, right? And so with reasoning models, this is where a lot of the complexity comes about and why memory is so important. Because if you have limited amounts of memory, then you can't serve so many users. If you have limited amounts of memory, your serving speeds get lower, right? And so your costs get a lot, lot worse, right?
Um, because all of a sudden, if I was used to, Hey, on the $16 an hour server, I'm serving Lama four or five B or if I'm serving, you know, deep seek V3, um, and it's all chat style applications, i.e. we're just chatting the sequence sensor thousand few thousand, right? Uh, you know, when you use the language model, it's a few thousand context length.
Most of the time, sometimes you're dropping a big document, but then you process it, you get your answer, you throw it away, right? You, you move on to the next thing, right? Whereas with reasoning, I'm now generating tens of thousands of tokens in sequence, right? And so this memory, this KV cache has to stay resident and you have to keep loading it. You have to keep it in memory constantly.
And now this butts out other users, right? If there's now a reasoning task, right? And the model is capable of reasoning, then all of a sudden that memory pressure means that I can't serve as many users simultaneously.
To give context, right? Everyone, one of the parts of like freaking this out was like trying to reach the capabilities. The other aspect is they did it so cheap, right? And the so cheap, we kind of talked about on the training side, why it was so cheap.
Showing 1461–1480 of 1,814 · page 74 of 91
← Previous
Next →