Nathan Lambert

speaker
1,814 appearances 3 recordings 2 series first heard Feb 2025 last heard 1 Feb

Nathan Lambert’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · Feb OctJan 26AprJulnow

Recordings per month over the last 12 months — 2 in all, peaking in Feb 2026 with 1.

Appearances

newest first · ▶ plays the moment
And then at the end of the day, that gives this kind of general foundation where the model can use CLI commands very nicely in your repo and handle Git for you and move things around and organize things or search to find more information, which if we're sitting in these chairs a year ago, it's something that we didn't really think of the models being doing.
So this is just kind of something that has happened this year and is...
totally transformed how we think of using AI, which I think is very magical.
It's such an interesting evolution and just unlocks so much value.
But it's not clear what the next avenue will be in terms of unlocking stuff like this.
I think that we'll get to continual learning later, but there's a lot of buzz around certain areas of AI, but no one knows when the next step function will really come.
Pre-training has gotten extremely expensive.
I think to scale up pre-training, it's also implying that you're going to serve a very large model to the users.
So I think that it's been loosely established the likes of GPT-4 and similar models were around one trillion, like this order of trillion parameters at the biggest size.
There's a lot of rumors that they've actually gotten smaller as training has gotten more efficient.
You want to make the model smaller because then your costs of serving go down proportionally.
These models, the cost of training them is really low relative to the cost of serving them to hundreds of millions of users.
I think DeepSeek had this famous number of about $5 million for pre-training at cloud market rates.
I think OMO3, section 2.4 in the paper, we just detailed how long we had the GPU clusters sitting around for training, which includes
engineering issues multiple seeds and it was like about two million dollars to rent the cluster to like deal with all the problems and headaches of training a model so these models are pretty like a lot of people could get one to ten million dollars to train a model but the recurring costs of serving millions of users is really billions of dollars of compute i think that
You can look at like a thousand GPU rental you can pay a hundred grand a day for.
And these companies could have millions of GPUs.
You can look at how much these things cost to sit around.
So that's kind of a big...
thing and then it's like if scaling is actually giving you a better model like is it going to be financially worth it and i think it'll kind of slowly will push it out as ai solves more compelling tasks so like the likes of cloud opus 4.5 making cloud code just work for things i think i launched this project called like the atom project which is like american truly open models in
Showing 221–240 of 1,814 · page 12 of 91 ← Previous Next →