Nathan Lambert
speaker
1,013 appearances
1 recordings
1 series
first heard Nov 2024
last heard Nov 2024
Nathan Lambert’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsNo recordings in the last 12 months.Older appearances are listed below; set an alert to hear about the next one.
Appearances
There's a lot of iteration on it.
I would say that
For performance, roughly, I think we get like ninety percent of our performance at SFT, and then the last ten percent a mix of DPO and RL.
And then like alpaca valve slash vibes things you get the most from DPO and RL rather than SFT.
But at this point, like our SFT mix like beats, like it's beats on R of L suite, the Llama three point one numbers.
And but it's like very focused on our evalves.
And I think the preference tuning kind of softens it a bit to be a nice
Yeah.
The a valve suite is complicated.
I don't know all the details off the top of my head.
It's good to look at the paper, but it we tried to do
There's this whole distinction between pre-training avowals and post-training evals.
And I do agree that most post-training evals should be like using a chat template.
So you're generating from the model to generate tokens.
A lot of them are CO2 chain of thought, and there's kind of a distribution over the number of in-context shots.
I think there's some that are zero shot, there are some that are eight shot, kind of depending on the domain.
And probably the trickiest thing to manage is answer extraction for reasoning, which is like does
Does the answer appear in the format that a Val expects?
Particularly on math, Llama 3.1 uses what's called the like Minerva format in a specific prompt.
And we use like what is called a flex format, which essentially you like if the default writing is like the answer is in boxed, you also allow like you'll allow like boxed and the answer is colon and like one other thing.
Showing 281–300 of 1,013 · page 15 of 51
← Previous
Next →