Chris Olah

speaker
254 appearances 1 recordings 1 series first heard Nov 2024 last heard Nov 2024

Chris Olah’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
No recordings in the last 12 months.Older appearances are listed below; set an alert to hear about the next one.

Appearances

newest first · ▶ plays the moment
What could I have said that would make you not make that error? Write that out as an instruction. And I'm going to give it to model. I'm going to try it. Sometimes I do that. I give that to the model in another context window often. I take the response. I give it to Claude and I'm like, hmm, didn't work. Can you think of anything else? You can play around with these things quite a lot.
I think there's just a huge amount of information in the data that humans provide, like when we provide preferences, especially because different people are going to pick up on really subtle and small things. So I've thought about this before, where you probably have some people who just really care about good grammar use for models, like, you know, was a semicolon used correctly or something?
And so you'll probably end up with a bunch of data in there that you as a human, if you're looking at that data, you wouldn't even see that. You'd be like, why did they prefer this response to that one? I don't get it. And then the reason is you don't care about semicolon usage, but that person does.
And so each of these single data points has, and this model just has so many of those, it has to try and figure out what is it that humans want in this really kind of complex, across all domains model. They're going to be seeing this across many contexts. It feels like the classic issue of deep learning, where historically we've tried to do edge detection by mapping things out.
And it turns out that actually, if you just have a huge amount of data that actually accurately represents the picture of the thing that you're trying to train the model to learn, that's more powerful than anything else. And so I think...
One reason is just that you are training the model on exactly the task and with a lot of data that represents many different angles on which people prefer and disprefer responses. I think there is a question of are you eliciting things from pre-trained models or are you teaching new things to models? And in principle, you can teach new things to models in post-training.
I do think a lot of it is eliciting powerful pre-trained models. So people are probably divided on this because obviously in principle, you can definitely teach new things. I think for the most part, for a lot of the capabilities that we... most use and care about.
A lot of that feels like it's there in the pre-trained models and reinforcement learning is eliciting it and getting the models to bring it out.
It's weird because I think that a lot of people prefer he for Claude. I actually kind of like that I think Claude is usually, it's slightly male-leaning, but it's like it can be male or female, which is quite nice.
I still use it, and I have mixed feelings about this because I'm like maybe, like I now just think of it as like, or I think of like the it pronoun for Claude as, I don't know, it's just like the one I associate with Claude. Yeah. I can imagine people moving to like he or she.
I've wondered if I anthropomorphize things too much. Because, you know, I have this like with my car, especially like my car and bikes, you know, like I don't give them names because then I once had, I used to name my bikes and then I had a bike that got stolen and I cried for like a week. And I was like, if I'd never given it a name, I wouldn't have been so upset. Yeah.
I felt like I'd let it down. Maybe it's that I've wondered as well, like it might depend on how much it feels like a kind of like objectifying pronoun. Like if you just think of it as like a, this is a pronoun that like objects often have, and maybe AIs can have that pronoun.
And that doesn't mean that I think of, if I call Claude it, that I think of it as less intelligent or like I'm being disrespectful. I'm just like, you are a different kind of entity. And so- That's, I'm going to give you the kind of the respectful it.
So there's a couple of components of it. The main component I think people find interesting is the kind of reinforcement learning from AI feedback. So you take a model that's already trained and you show it two responses to a query and you have a principle. So suppose the principle, like we've tried this with harmlessness a lot. So suppose that the query is about
weapons and your principle is select the response that is less likely to encourage people to purchase illegal weapons. That's probably a fairly specific principle, but you can give any number.
And the model will give you a kind of ranking and you can use this as preference data in the same way that you use human preference data and train the models to have these relevant traits from their feedback alone instead of from human feedback. So if you imagine that, like I said earlier with the human who just prefers the kind of like semi-colon usage in this particular case,
You're kind of taking lots of things that could make a response preferable and getting models to do the labeling for you, basically.
In principle, you could use this for anything. And so harmlessness is a task that it might just be easier to spot. So when models are less capable, you can use them to rank things according to principles that are fairly simple, and they'll probably get it right. So I think one question is just, is it the case that the data that they're adding is fairly reliable? Yeah.
But if you had models that were extremely good at telling whether one response was more historically accurate than another, in principle, you could also get AI feedback on that task as well. There's a kind of nice interpretability component to it because you can see the principles that went into the model when it was being trained. And it gives you a degree of control.
So if you were seeing issues in a model, like it wasn't having enough of a certain trait, then you can add data relatively quickly that should just train the model to have that trait. So it creates its own data for training, which is quite nice.
Showing 81–100 of 254 · page 5 of 13 ← Previous Next →