Nathan Lambert

speaker
1,013 appearances 1 recordings 1 series first heard Nov 2024 last heard Nov 2024

Nathan Lambert’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
No recordings in the last 12 months.Older appearances are listed below; set an alert to hear about the next one.

Appearances

newest first · ▶ plays the moment
And it's like kind of reductionist, and it's a little cynical to view nonprofits as storytelling works, but when your money mostly comes from a tech billionaire's estate.
Like you you gotta stay some of it as it is and but the momentum is good.
So it it's fun to be in the space and
I see things continuing for the time being.
Yeah, so I think this kind of story fits well with what we were talking about with AI two.
So
AI2 is kind of post open post-training recipes have been under this Tulu brand for over a year, almost a year and a half.
I think Tulu 1 was back when like Open Assistant was new, was like, what can we do by mixing all these popular data sets that people are putting out in 2023?
All these data sets and small models, like how do you systematically mix instruction data sets to combine them with open resources?
And Tulu 2 was around when DPO was very popular.
And that was when we showed that you can like scale DPO to 70 billion parameters and get like continue getting the benefits from preference fine-tuning and open recipes.
Then we went away and did Tulu 2.5, which was kind of like trying to answer the PPO versus DPO question, which the TLDR was if you tune your parameters right, PPO is probably better a bit, but it's probably not worth the effort to spend all your time on preference tuning.
when you could just be making better data and better pipelines, which is what two through three is about, which is kind of like inspired by the transition we're seeing with like the Llama report with Trap Bot Arena is like turned into a hockey stick again where we have those incremental scores and the OpenAI and Google are like skyrocketing their scores again.
Nematron paper, Apple had a paper, which these post-training recipes are just much more complicated than big SFT step.
do DPO on some chat data and be done.
And the philosophy is kind of like we know these labs are doing it, like how do we try to understand what the open groups should be doing and where there are hills to climb when you're increasing the complexity substantially of post training?
Like
Lamba three has hundreds of people on it and like a large amount on post training.
We have ten to twenty on ours, and it's like what could academics actually do to do modern post training?
And curate new data, use new algorithms, sequence things together, but like kind of move beyond this paradigm of like just increasing like alpaca valve or vibe scores and try to get really specific on improving math, improving IF eval.
Showing 41–60 of 1,013 · page 3 of 51 ← Previous Next →