Nathan Lambert
speaker
1,013 appearances
1 recordings
1 series
first heard Nov 2024
last heard Nov 2024
Nathan Lambert’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsNo recordings in the last 12 months.Older appearances are listed below; set an alert to hear about the next one.
Appearances
Which is like m the all capital math eval, which you can improve this, but a lot of it comes down to the formatting.
So some math data sets will have different formatting in the completion that will make it harder for your model to learn from that, or it can affect something like GSM AK or coding or stuff like this.
So those are the kind of things that you're doing.
And then also at the end, you're probably gonna try to include some more general chat data.
We included a hundred thousand multilingual samples from Cohere's Aya because we know Chatbot Arena is has a decent amount of multilingual, even though multilingual wasn't in our val suite.
And there are some um some other safety things which are more borderline, which is just like how should a model behave?
And th this is kind of what happens in SFT is like there's some there's some art to it in squashing second order effects, but I think largely that's
really similar to what the closed labs are doing.
And then at preference tuning it takes a kind of big turn, which is the closed labs use humans for their preference data.
We do not have the money, so we use all a LM as a judge to collect our preference data.
I like to quote John Schulman's saying on this, which is the pithy way to say it is like human preference data is um high noise but low bias, and LLM preference data is low noise but high bias.
We do not know exactly what bias we are getting, but largely we can still see pipelines of we collect our own preference data from various model generations, language models label it, um the preference data is good.
The biggest change here is from data sets like Ultra Feedback, which exist on Hugging Face.
We essentially redid their pipeline with completions from the models we had trained at SFT, and that gives a meaningful like low percentage improvement over just using a random preference data.
on the Hugging Face Hub.
So it's higher effort.
You have to go through the effort of generating completions of BLM, making sure that's all right, getting diversity, doing LLM as a judge.
But that type of thing gives us like it we give an across our experiments, a time and time again, this like on policy idea preferences, even with LLM as a judge, was better.
So I do think that that's kind of what if you look at meta's system diagram in their paper, they show this.
They're like, we take a new model, we pass
Showing 141–160 of 1,013 · page 8 of 51
← Previous
Next →