Nathan Lambert

speaker
1,013 appearances 1 recordings 1 series first heard Nov 2024 last heard Nov 2024

Nathan Lambert’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
No recordings in the last 12 months.Older appearances are listed below; set an alert to hear about the next one.

Appearances

newest first · ▶ plays the moment
Like if we could run O one on our same prompts and you look at way more of them, I think it'd be a a lot easier to say, like, yeah, this is a very similar behavior.
We just have to you have to then scale the training regime to be stable, do this in every domain and like do it repeatedly.
When you do RLHF and RL, it's like you need to look at the generations to make sure the model's still working.
It's just like a normal thing and they're like, oh, this is a brand new thing.
Look at your data.
Yeah.
And I was like, I'm gonna tweet this.
This is so silly.
In that respect, it's not like we're searching over the generations for the word wait or something.
Like I I we found a couple of them.
I'm sure there's many more if we look for it.
I think you need to do a few mixes.
You need to do a few informed mixes based on the base model and like what capabilities it needs more or less help with at kind of each stage.
But you can probably start from the same superset and then do a few experiments with more or less of various behaviors.
So you're probably like a a few cycles per base model to make sure things look right if you really want to get the best performance.
Um you can take these data.
That's off the
just like kind of from the shelf and use them and you'll probably get like eighty to ninety five percent of the performance on a given base model.
So depending on how much you care, it's pretty fine.
And I would say similar at
Showing 881–900 of 1,013 · page 45 of 51 ← Previous Next →