Nathan Lambert

speaker
1,013 appearances 1 recordings 1 series first heard Nov 2024 last heard Nov 2024

Nathan Lambert’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
No recordings in the last 12 months.Older appearances are listed below; set an alert to hear about the next one.

Appearances

newest first · ▶ plays the moment
it in, we get new preference data and we train on it.
The difference is that they do multiple iterations, which could be based on how they get data, could be based on how their timeline is and stuff like this.
I do think multiple iterations is something we're seeing again and again that we didn't look at, but just kind of checking the box of like, okay, on policy data does work for preference tuning.
It's not a huge paradigm shift, but it's a whole different approach than people have to do.
We've seen some of this in kind of like on policy and new like online GPO algorithm variants that everyone kind of gets like spiralized.
It's like, oh God, we have some other uh star PO algorithm I can't pay attention to.
But the algorithms were going in that direction where they're doing more of like generation from the model and labeling.
But the way that we did it I think is a bit more segmented, which is kind of like closer to what Llama is doing.
This at a high level is similar to what Lama three point one does.
I think
I've heard that they weren't doing any fancy verifiable RL, any O one stuff on theirs.
What we kind of added to this um
I think I can I'm specifically this is in the pa like you can look at the acknowledgement section of the paper and you'll see like one name and a bunch of
normal grant garble gobble.
It's not in the draft I sent you.
But like um there's one name there and he told us to just do RL on verifiable outputs.
So this is the stage three that we did, which is essentially taking the training set from things like math, GSM 8K,
And it's actually like IFFL is verifiable because you like count if the constraint if val as a summary is a bunch of is an evaluation where there are constraints on prompts like respond, make sure your response has X word, make sure your response has X paragraphs.
And these are things that are verifiable in Python code.
So across these things like math, instruction following, we have a system where you have these prompts with constraints.
Showing 161–180 of 1,013 · page 9 of 51 ← Previous Next →