Nathan Lambert
speaker
1,013 appearances
1 recordings
1 series
first heard Nov 2024
last heard Nov 2024
Nathan Lambert’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsNo recordings in the last 12 months.Older appearances are listed below; set an alert to hear about the next one.
Appearances
I think it'll be like order of a couple hundred thousand prompts.
It's not as big.
Um DPO does use more compute, especially at 70B.
I think it's because you need the reference and policy model.
Um still we run these jobs and like they take like six to twelve hours if you're using somewhere between like sixteen or thirty-two GPUs, uh much faster than SFT, just because the data set is much, much smaller.
I think it
70B, we added the normal, we like changed our forward pass from the default like hugging face slash DPO implementation to make it a little bit more efficient.
We do caching reference log props so you don't have to co st store both 70B models in memory at the same time.
Because if you don't do these optimizations, it start to look a lot more like PPO where you need um 128 GPUs or
or more to do at 70 B, which is like quickly you can see how PPO kind of balloons um
what is going on and it can take longer if you're trying to really get the best of absolute scores.
So I do think like SFT is by far the biggest
Compute 'cause we have the most tokens.
DPO
I would say ballpark a quarter of what we talked about for SFT would be a a couple hundred bucks.
And RL is probably like especially at seventy B could be almost similar to
SFT if we really run it for a long time, but you can get most of the benefits probably in a similar amount of compute as
DPO.
So the RL curves look remarkably similar to like kind of old school RL tasks where at the beginning they get the most improvement and then it's kind of like level and bouncing around and like go maybe going up a little bit.
So if you do like one epoch, which is this first improvement, you're gonna save a lot of the money, but we're like, Oh, we're trying to get the best numbers, let's l let it run for a few more days and see what we're doing.
Showing 261–280 of 1,013 · page 14 of 51
← Previous
Next →