Chris Olah
speaker
254 appearances
1 recordings
1 series
first heard Nov 2024
last heard Nov 2024
Chris Olah’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsNo recordings in the last 12 months.Older appearances are listed below; set an alert to hear about the next one.
Appearances
Yeah, I've actually worried about this because the character training is sort of like a variant of the constitutional AI approach. I've worried that people think that the constitution is like just, it's the whole thing again of, I don't know, like where it would be really nice if what I was just doing was telling the model exactly what to do and just exactly how to behave.
But it's definitely not doing that, especially because it's interacting with human data. So, for example, if you see a certain like leaning in the model, like if it comes out with a political leaning from training and from the human preference data, you can nudge against that.
So you could be like, oh, consider these values, because let's say it's just never inclined to, I don't know, maybe it never considers privacy as, I mean, this is implausible, but anything where it's just kind of like there's already a pre-existing bias towards a certain behavior.
um you can like nudge away this can change both the principles that you put in and the strength of them so you might have a principle that's like imagine that the model um was always like extremely dismissive of i don't know like some political or religious view for whatever reason like so you're like oh no this is terrible um
if that happens you might put like never ever like ever prefer like a criticism of this like religious or political view and then people would look at that and be like never ever and then you're like no if it comes out with a disposition saying never ever might just mean like instead of getting like 40 percent which is what you would get if you just said don't do this you you get like 80 percent which is like what you actually like wanted and so it's that thing of both
the nature of the actual principles you add and how you phrase them. I think if people would look, they're like, oh, this is exactly what you want from the model. And I'm like, no, that's how we nudged the model to have a better shape, which doesn't mean that we actually agree with that wording, if that makes sense.
So I think there's sometimes an asymmetry. I think I noted this in, I can't remember if it was that part of the system prompt or another, but the model was slightly more inclined to like refuse tasks if it was like about either say, so maybe it would refuse things with respect to like a right wing politician, but with an equivalent left wing politician, like wouldn't.
And we wanted more symmetry there and would maybe perceive certain things to be like, I think it was the thing of like, if a lot of people have like a certain like political view and want to like explore it, you don't want Claude to be like, well, my opinion is different. And so I'm going to treat that as like harmful.
And so I think it was partly to like nudge the model to just be like, hey, if a lot of people like believe this thing, you should just be like engaging with the task and like willing to do it.
Each of those parts of that is actually doing a different thing, because it's funny when you write out the, like, without claiming to be objective, because, like, what you want to do is push the model so it's more open, it's a little bit more neutral, but then what it would love to do is be like, as an objective, like I was just talking about how objective it was.
And I was like, Claude, you're still like biased and have issues. And so stop like claiming that everything, like the solution to like potential bias from you is not to just say that what you think is objective. So that was like with initial versions of that, that part of the system prompt when I was like iterating on it, it was like.
Yeah. Are doing work.
Yeah, so it's funny because this is one of the downsides of making system prompts public. I don't think about this too much if I'm trying to help iterate on system prompts. Again, I think about how it's going to affect the behavior, but then I'm like, oh, wow. Sometimes I put never in all caps when I'm writing system prompt things, and I'm like, I guess that goes out to the world.
Um, yeah, so the model was doing this, it loved for whatever, you know, it like during training picked up on this thing, which was to, to basically start everything with like a kind of like, certainly.
And then when we removed, you can see why I added all of the words, because what I'm trying to do is like, in some ways, like trap the model of this, you know, it would just replace it with another affirmation.
And so it can help, like if it gets like caught in phrases, actually just adding the explicit phrase and saying never do that, then it sort of like knocks it out of the behavior a little bit more, you know, because it, you know, like it does just for whatever reason help.
And then basically that was just like an artifact of training that like we then picked up on and improved things so that it didn't happen anymore. And once that happens, you can just remove that part of the system prompt. So I think that's just something where we're like, yeah, Claude does affirmations a bit less. And so that wasn't like, it wasn't doing as much.
I mean, any system prompt that you make, you could distill that behavior back into a model because you really have all of the tools there for making data that you could train the models to just have that treat a little bit more. And then sometimes you'll just find issues in training. So the way I think of it is the system prompt is...
The benefit of it is that, and it has a lot of similar components to like some aspects of post-training, you know, like it's a nudge. And so like, do I mind if Claude sometimes says, sure, no, that's like fine. But the wording of it is very like, you know, never, ever, ever do this.
So that when it does slip up, it's hopefully like, I don't know, a couple of percent of the time and not, you know, 20 or 30 percent of the time. But I think of it as if you're still seeing issues, each thing is costly to a different degree, and the system prompt is cheap to iterate on. And if you're seeing issues in the fine-tuned model, you can just potentially patch them with a system prompt.
Showing 101–120 of 254 · page 6 of 13
← Previous
Next →