Nathan Lambert

speaker
1,013 appearances 1 recordings 1 series first heard Nov 2024 last heard Nov 2024

Nathan Lambert’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
No recordings in the last 12 months.Older appearances are listed below; set an alert to hear about the next one.

Appearances

newest first · ▶ plays the moment
They do this on many more domains at once and probably with a mix of deterministic and learn to air pyers and they probably do way more RL training than we are doing with some tricks.
But I I do think that
it's not unreasonable to be excited about that.
And I don't think
There's good reasons that there's some smoke and mirrors about O one where they made it seem more complicated than it is.
It's a breakthrough and I'm sure they did a lot of really interesting novel things and hacks to get there, but it's n
We've seen all the labs release the same things and again and again.
Like there's gonna be O one equivalents from claw from Anthropic and Google within five months or whatever you wanna
I find it very unlikely that some of one of our
Like we didn't train we didn't do SFT on like O one chain of thoughts or anything.
We can't get O one chain of thoughts.
We didn't do any like intermediate edits to chain of thoughts or anything like this.
Um people that I think are smart have said that they've seen llama do this same behavior if you like really crank the temperature up and do things.
So it's not in that regard, it's like not that.
different.
I mean even like the stupid reflection seventy B model is in principle a similar idea.
So like there are a lot of ways to induce this behavior.
This is a very very, very open ended one that was with the RL loss on some verifier.
So like
That's why I was like, Oh, this is so it's like we we stumble upon O one like behavior without even meaning it.
Showing 841–860 of 1,013 · page 43 of 51 ← Previous Next →