Nathan Lambert
speaker
1,814 appearances
3 recordings
2 series
first heard Feb 2025
last heard 1 Feb
Nathan Lambert’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 2 in all, peaking in Feb 2026 with 1.
Appearances
And it's like, I don't know, there's like two or three people in the world that were very interested in this.
He's a PhD student, which gives you an advantage.
But like for me, that was a topic I was waiting for someone to be like, hey, I have time to
spent cycles on this.
And I'm sure there's a lot more very narrow things where you're just like, oh, it doesn't make sense that there was no answer to this.
And I think that it's just like, there's so much information coming that people are like, I can't grab onto any of these.
But if you just actually stick in an area, I think there's a lot of interesting things to learn.
Some of them contradict.
I just edited the book and I was like, there's a chapter where I had to be like,
X papers say one thing, and X papers say another thing, and we'll see what comes out to be true.
character training is interesting because there's so little out of it but we talk about how people engage with these models and like we feel good using them because they're positive but that can go too far it can be too positive and it's like essentially it's how do you change your data and
or decision-making to make it exactly what you want.
And OpenAI has this thing called a model spec, which is essentially their internal guideline for what they want the model to do.
And they publish this to developers.
So essentially you can know what is a failure of OpenAI's training, which is like they have the intentions and they haven't met it yet, versus what is something that they actually wanted to do and that you don't like.
And that transparency is very nice, but all the methods for...
curating these documents and how easy it is to follow them is not very well known.
I think the way the book is designed is that the reinforcement learning chapter is obviously what people want because everybody hears about it with RLVR.
And it's the same algorithms and the same map, but it's just like you can use it in very different documents.
So I think the core of RLHF is like how messy preferences are.
Showing 601–620 of 1,814 · page 31 of 91
← Previous
Next →