Nathan Lambert

speaker
1,814 appearances 3 recordings 2 series first heard Feb 2025 last heard 1 Feb

Nathan Lambert’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · Feb OctJan 26AprJulnow

Recordings per month over the last 12 months — 2 in all, peaking in Feb 2026 with 1.

Appearances

newest first · ▶ plays the moment
My voice is based on this, but a lot of people don't provide this to models and the models weren't designed to like take this amount of context previously, like the agentic models are just starting.
So it's this kind of trade-off of, do we need to update the weights of this model with
this continual learning thing to make them learn fast.
Or the counter argument is we just need to provide them with more context and information and they will have the appearance of learning fast by just having a lot of context and being very smart.
I think the colloquially accepted thing is that it's a compute and data problem where you can, and sometimes like small architecture things, which are like attention variants.
So if you have, we talked about like hybrid attention models, which is essentially if you have what looks like a state space model within your transformer, and like those are better suited because you have to spend less compute to model the data.
furthest along token and i think that but those aren't free because they have to be accompanied by a lot of compute or um the right data so how many sequences of 100 000 tokens do you have in the world and where do you get these and i think it just ends up being pretty expensive to scale them so we've like gotten to pretty quickly to like a million tokens of input context length and
And I would expect it to keep increasing and get to like 2 million or 5 million this year, but I don't expect it to go to like 100 million.
That would be like a true breakthrough.
And I think those breakthroughs are possible.
Like the continual learning thing, I think of it as a research problem where there could be a breakthrough that just makes transformers work way better at this and it's cheap.
Like these things could happen with so much scientific attention, but turning the crank, it'll be consistent increases over time.
There are some rules of thumb where essentially you pre-train a language model.
Like, oh no, we pre-trained at like 8K context length and then extended to 32K with training.
And there's some rules of thumb where you're just like essentially doubling the training context length, takes like 2X compute, and then you can normally like 2 to 4X the context length again.
So I think a lot of it ends up being kind of
compute bound at pre-training, which is, and it's like we talked about this, everyone talks about this big increase in compute for the top labs this year, and that should reflect in some longer context windows.
But I think on the post-training side, there's some more interesting things, which is as we have agents, the agents are going to manage this context on their own, where now people that use Cloud Code a lot dread the compaction, which is when Cloud takes its entire full 100,000 tokens of work and compacts it into a bulleted list.
But what the next models will do, I'm just not
I'm sure people are already working on this, is essentially the model can control when it compacts and how.
Showing 781–800 of 1,814 · page 40 of 91 ← Previous Next →