Nathan Lambert
speaker
1,013 appearances
1 recordings
1 series
first heard Nov 2024
last heard Nov 2024
Nathan Lambert’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsNo recordings in the last 12 months.Older appearances are listed below; set an alert to hear about the next one.
Appearances
So people are kind of heads down like building math data, building IF of L data.
And then it's probably a trade-off of like some mi people are working on SFT data and SFT mixing and some people are working on DPO and DPO mixing.
And then like soon this whole like does this RL thing work, which can take like two people to start doing this RL thing, kind of on their own, but like part of this project.
And then as the weeks and months go by, like you try to finalize SFT.
So like SFT has finalized the data mix was probably like a month ago, then we have to
We did more decontamination or like, oh, we have to turn it again.
Oh, our 70B hyperparameters are wrong, oh, we have to train it again.
Oh, we have to try model merging.
We have to train it a few more times.
Um the quick note on model merging is like it's a safe bet to merge SFT by running multiple seeds on the same data set.
But it could often be that just by running on multiple seeds, one of your seeds for SFT is actually gonna be the best.
So if you wanna get a good SFT model, you wanna just train it three to five times across a few random seeds.
Which is like pretty funny.
It's just like even more compute.
And then like it for the last month has been mostly like final touches and then full on DPO, which is like we have to generate we have to do this on policy thing, which takes a lot of API credits and a lot of just hands-on, which is like you take your SFT model, you run completions, and then those are like the person that just owns on policy preference data.
So you give them prompts and models and they do like VLM to make generations, and then they use the open AI API to do LLM as a judge.
And then they're like, here's your preference data set.
And then you have we have different models.
So we have eight B, seven D B, we have Olmo models.
So it's like this is the person in the team that's just owning, like they make preference data.
Showing 621–640 of 1,013 · page 32 of 51
← Previous
Next →