Nathan Lambert
speaker
1,814 appearances
3 recordings
2 series
first heard Feb 2025
last heard 1 Feb
Nathan Lambert’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 2 in all, peaking in Feb 2026 with 1.
Appearances
It doesn't cost them anything.
Yeah.
The ecosystem has gotten better on that front, but mostly downstream of these new providers providing such open licenses.
That was funny when you pulled up Perplexity, it said Kimi K2 Thinking hosted in the US, which is just like an exact, I've never seen this, but it's an exact example of what we're talking about where people are sensitive to this.
Like Kimi K2 Thinking and Kimi K2 is a model that is very popular.
People say that it has very good creative writing and also in doing some software things.
There's just these little quirks that people pick up on with different models that they like.
I would say that the systems also change a lot.
I think if you listen to NVIDIA's announcements, they talk about these things like, you now do FP8, you can now do FP4.
And what is happening is these labs are figuring out how to utilize more compute to put it into one model, which lets them train faster, and that lets them put more data in, and then you can find better configurations faster.
By doing this, so you can look at like the essentially the tokens per second per GPU is a metric that you look at when you're doing large scale training.
And you could get you can go from like 10k to 13k by turning on FP8 training, which means you're using less memory per parameter in the model.
And by saving less information, you do less communication.
You can train faster.
So all of these like system things underpin way faster experimentation on data and algorithms.
That is kind of like it's this kind of loop that keeps going where it's kind of hard to describe when you look at the architecture and they're exactly the same.
But the code base used to train these models is going to be vastly different and different.
You could probably, like, I don't, the GPUs are different, but you probably trained GPT-OSS 20B way faster in wall clock time than GPT-2 was trained at the time.
I like to start with the technical definition of scaling law, which kind of informs all of this.
The scaling law is a power-law relationship between, you can think of the x-axis, so kind of what you are scaling as a combination of compute and data, which are kind of similar.
Showing 181–200 of 1,814 · page 10 of 91
← Previous
Next →