Chris Olah
speaker
254 appearances
1 recordings
1 series
first heard Nov 2024
last heard Nov 2024
Chris Olah’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsNo recordings in the last 12 months.Older appearances are listed below; set an alert to hear about the next one.
Appearances
And I think there's reasons to like not want that, especially as models become more powerful, because you're like, there might just be a small number of people who want to use models for really harmful things.
um but having them having models as they get smarter like figure out where that line is does seem important um and then yeah with the apologetic behavior i don't like that and i like it when claude is a little bit more willing to like push back against people or just not apologize part of me is like it often just feels kind of unnecessary so i think those are things that are hopefully decreasing um over time um
And yeah, I think that if people say things on the Internet, it doesn't mean that you should think that that like that could be that like there's actually an issue that 99% of users are having that is totally not represented by that. But in a lot of ways, I'm just like attending to it and being like, is this right? Do I agree? Is it something we're already trying to address? That feels good to me.
Yeah, that seems like a thing that you could, I could definitely encourage the model to do that. I think it's interesting because there's a lot of things in models that like, it's funny where there are some behaviors where You might not quite like the default, but then the thing I'll often say to people is you don't realize how much you will hate it if I nudge it too much in the other direction.
So you get this a little bit with like correction. The models accept correction from you, like probably a little bit too much right now. You know, you can over, you know, it'll push back if you see like, no, Paris isn't the capital of France. But really like things that I'm, I think that the model's fairly confident in, you can still sometimes get it to retract by saying it's wrong.
At the same time, if you train models to not do that and then you are correct about a thing and you correct it and it pushes back against you and is like, no, you're wrong. It's hard to describe like that's so much more annoying. So it's like a lot of little annoyances versus like, one big annoyance. It's easy to think that like we often compare it with like the perfect.
And then I'm like, remember, these models aren't perfect. And so if you nudge it in the other direction, you're changing the kind of errors it's going to make. And so think about which of the kinds of errors you like or don't like.
So in cases like apologeticness, I don't want to nudge it too much in the direction of like almost like bluntness, because I imagine when it makes errors, it's going to make errors in the direction of being kind of like rude. Whereas at least with apologeticness, you're like, oh, OK, it's like a little bit, you know, I don't like it that much, but at the same time, it's not being mean to people.
And actually, the time that you undeservedly have a model be kind of mean to you, you probably like that a lot less than you mildly dislike the apology. So it's one of those things where I'm like, I do want it to get better, but also while remaining aware of the fact that there's errors on the other side that are possibly worse. Yeah.
I think you could just tell the model is my guess. Like for all of these things, I'm like, the solution is always just try telling the model to do it. And then sometimes it's just like, like, I'm just like, oh, at the beginning of the conversation, I just throw in like, I don't know. I like you to be a New Yorker version of yourself. I never apologize.
Then I think Claude will be like, okie doke, I'll try. Or it'll be like, I apologize. I can't be a New Yorker type of myself, but hopefully I wouldn't do that.
It's more like constitutional AI. So it's kind of a variant of that pipeline. So I worked through constructing character traits that the model should have. They can be kind of like... shorter traits or they can be kind of richer descriptions. And then you get the model to generate queries that humans might give it that are relevant to that trait.
Then it generates the responses and then it ranks the responses based on the character traits. So in that way, after the generation of the queries, it's very much similar to constitutional AI. It has some differences. So I quite like it because it's like Claude's training in its own character because it doesn't have any... It's like constitutional AI, but it's without any human data. Yeah.
Or I'll just misinterpret it and be like, oh yeah. Go with it.
Yeah. I mean, I have two thoughts that feel vaguely relevant. Let me know if they're not. Like, I think the first one is people can underestimate the degree to which what models are doing when they interact. I think that we still just too much have this model of AI as computers. And so people often say, well, what values should you put into the model?
And I'm often like, that doesn't make that much sense to me because I'm like, hey, as human beings, we're just uncertain over values. We have discussions of them. We have... a degree to which we think we hold a value, but we also know that we might not, and the circumstances in which we would trade it off against other things. These things are just really complex.
I think one thing is the degree to which maybe we can just aspire to making models have the same level of nuance and care that humans have, rather than thinking that we have to program them in the very kind of classic sense. I think that's definitely been one.
The other, which is like a strange one, I don't know if it, maybe this doesn't answer your question, but it's the thing that's been on my mind anyway, is like the degree to which this endeavor is so highly practical. And maybe why I appreciate like the empirical approach to alignment. I slightly worry that it's made me maybe more empirical and a little bit less theoretical.
So people, when it comes to AI alignment, will ask things like, well, whose values should it be aligned to? What does alignment even mean? And there's a sense in which I have all of that in the back of my head. I'm like, you know, there's like social choice theory. There's all the impossibility results there.
So you have this like this giant space of like theory in your head about what it could mean to like align models. But then like practically, surely there's something where we're just like if a model is like if especially with more powerful models, I'm like my main goal is like I want them to be good enough that things don't go terribly wrong.
Showing 141–160 of 254 · page 8 of 13
← Previous
Next →