Nathan Lambert
speaker
1,013 appearances
1 recordings
1 series
first heard Nov 2024
last heard Nov 2024
Nathan Lambert’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsNo recordings in the last 12 months.Older appearances are listed below; set an alert to hear about the next one.
Appearances
K L
Um y in a lot of RL you use an approximate rather than a full K L.
Where the approximate you're essentially just I have I pulled it up.
I have the silly like RLHF
Book that I'm working on, which is mostly like notes on fundamentals of RLHF.
And like the regular is this.
So there's like there's a difference between what a lot of people do in these implementations as an approximate KL.
I think that John Schwimman has a blog post on this that everyone references, where the approximate is you just subtract the log probs of a generation for the policy model.
You do that minus the reference log props, which is a different function.
I don't know if it's a change it seems like that's a change between square and not.
Just kinda thinking about it off the top of my head without having the exact tail equation.
So
Um so the RL math is essentially you will give it a label based on the whole trajectory and then the value model the
If you're doing PPO, the value model will take a generation and it will output a value per token.
Um, then you will have the label from the reward model, which is like the reward from your environment, and that is used to update the value model to kind of have these per token updates.
And then the policy is actually going like in PPO, you're actually gonna be taking attribution for every token in your batch and doing updates based on what they think will be a long-term better generation in PPO.
That's very different than DPO.
I don't have like a per token dis um understanding of what is happening in DPO because it's if you think about the loss function it's doing
Essentially all the tokens are kind of grouped into chosen or rejected.
And it's trying to increase the margin between chosen and rejected log probs across a sequence.
Showing 441–460 of 1,013 · page 23 of 51
← Previous
Next →