Nathan Lambert
speaker
1,013 appearances
1 recordings
1 series
first heard Nov 2024
last heard Nov 2024
Nathan Lambert’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsNo recordings in the last 12 months.Older appearances are listed below; set an alert to hear about the next one.
Appearances
Which I think early in the DPO days there were questions on um like length issues, because what you're looking at is the sum of the log probs.
So then if you're like
all negative like they're all like negative numbers because it's a log of a probability.
They're all less than one.
So they all become negative numbers and you're summing them up.
And it's like, what does that do?
Because you're therefore like increasing the margin between these two negative numbers is what the loss is doing.
So I don't exactly have as clear of an intuition there, but I
You kind of look at the losses and see how those are different, where there's I think a bit more like per token attribution in RL versus DPO, and they're both substantially different than SFT.
so depends on It's often glossed
It depends on your training setup.
So some setups you'll
start the value model from nothing.
So then it'll take a few hundred steps to take the scores from the reward model to kind of warm up the value model.
So you can kind of see if your value model is a random init, you can see this where your like value model loss has to converge before your policy will start changing notably.
There's other ways where you can init a value model from a reward model or an SFT model, which is what we do on this like verifiable outputs to kind of make learning a little bit cleaner.
And I actually don't know how that
changes the value model exactly.
I think that essentially the it must it must be some mapping between like the log props and value that it is doing, which is why you would warm start with the model.
But I I don't know exactly what that mapping does.
Showing 461–480 of 1,013 · page 24 of 51
← Previous
Next →