Nathan Lambert

speaker
1,013 appearances 1 recordings 1 series first heard Nov 2024 last heard Nov 2024

Nathan Lambert’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
No recordings in the last 12 months.Older appearances are listed below; set an alert to hear about the next one.

Appearances

newest first · ▶ plays the moment
It's like these things like
that type of thing makes me get a lot of sympathy to the like whole like Dario whatever mindset of like they just wanna learn.
It's like there's something complicated going on that I don't have deep enough intuitions of deep learning to grapple with.
That's what it is.
You essentially
look at the log props of the original model and the new model and you make sure that the difference in probability is not too big between your RL model and your um reference.
And I think the way that we phrase it is like you have a KL budget.
So you can like only change the models so much.
And like once you reach your budget,
you normally don't expect the model to keep changing if you're doing value updates because like it can't like it can't make substantial changes.
It's just moving around in the same neighborhood.
Which I think is a nice framing.
I wish more DPO like DPO is different where it's like
A controlled KL distance through their beta parameter throughout the like the number of epochs that they're doing.
Whereas RL is like where doing online RL is a little bit more open-ended on the KL side of things.
But I do think this in general and post-training, showing more plots of like performance versus the amount of KL that you spend is very nice.
I think historically, these just sound like total random numbers to people not looking at this.
It's like PPO KL.
will be like 10 to 20 scale type thing.
Um this is when you're doing the full thing on chat, for example, or like GSM8K, like really basic math KL is like one.
Showing 401–420 of 1,013 · page 21 of 51 ← Previous Next →