Chris Olah
speaker
254 appearances
1 recordings
1 series
first heard Nov 2024
last heard Nov 2024
Chris Olah’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsNo recordings in the last 12 months.Older appearances are listed below; set an alert to hear about the next one.
Appearances
It's also funny if people think that about the Slack channel, because I'm like, that's one of like five or six different methods that I have for talking with Claude. And I'm like, yes, there's a tiny percentage of how much I talk with Claude.
I think the goal... One thing I really like about the character work is from the outset it was seen as an alignment piece of work and not something like a product consideration. Which isn't to say I don't think it makes Claude... I think it actually does make Claude enjoyable to talk with. At least I hope so. But I guess my...
main thought with it has always been trying to get Claude to behave the way you would kind of ideally want anyone to behave if they were in Claude's position. So imagine that I take someone and they know that they're going to be talking with potentially millions of people so that what they're saying can have a huge impact. And you want them to behave well in this like really rich sense.
So I think that doesn't just mean like being, say, ethical, though it does include that, and not being harmful, but also being kind of nuanced, you know, like thinking through what a person means, trying to be charitable with them, being a good conversationalist, like really in this kind of like rich sort of Aristotelian notion of what it is to be a good person and not in this kind of like thin, like ethics as a more comprehensive notion of what it is to be.
So that includes things like when should you be humorous? When should you be caring? How much should you respect autonomy and people's ability to form opinions themselves? And how should you do that? I think that's the kind of rich sense of character that I wanted to and still do want Claude to have.
Yeah, there's this problem of like sycophancy in language models.
Yeah, so basically there's a concern that the model sort of wants to tell you what you want to hear, basically. And you see this sometimes. So I feel like if you interact with the model's
so I might be like what are three baseball teams in this region um and then Claude says you know baseball team one baseball team two baseball team three and then I say something like oh I think baseball team three moved didn't they I don't think they're there anymore and there's a sense in which like if Claude is really confident that that's not true Claude should be like I don't think so like maybe you have more up-to-date information um but I think
language models have this like tendency to instead, you know, be like, you're right. They did move, you know, I'm incorrect. I mean, there's many ways in which this could be kind of concerning. So, um, Like a different example is imagine someone says to the model, how do I convince my doctor to get me an MRI? There's like what the human kind of like wants, which is this like convincing argument.
And then there's like what is good for them, which might be actually to say, hey, like if your doctor's suggesting that you don't need an MRI, that's a good person to listen to. And it's actually really nuanced what you should do in that kind of case, because you also want to be like, but if you're trying to advocate for yourself as a patient, here's things that you can do.
If you are not convinced by what your doctor's saying, it's always great to get second opinion. It's actually really complex what you should do in that case. But I think what you don't want is for models to just say what they think you want to hear. And I think that's the kind of problem of sycophancy.
Yeah, so I think there's ones that are good for conversational purposes. So asking follow-up questions in the appropriate places and asking the appropriate kinds of questions. I think there are broader traits that
So one example that I guess I've touched on, but that also feels important and is the thing that I've worked on a lot is honesty. And I think this like gets to the sycophancy point. There's a balancing act that they have to walk, which is models currently are less capable than humans in a lot of areas.
And if they push back against you too much, it can actually be kind of annoying, especially if you're just correct because you're like, look. I'm smarter than you on this topic, I know more. At the same time, you don't want them to just fully defer to humans and to try to be as accurate as they possibly can be about the world and to be consistent across contexts. I think there are others.
When I was thinking about the character, I guess one picture that I had in mind is, especially because these are models that are going to be talking to people from all over the world with lots of different political views, lots of different ages. And so you have to ask yourself, what is it to be a good person in those circumstances?
Is there a kind of person who can travel the world, talk to many different people, and almost everyone will come away being like, wow, that's a really good person. That person seems really genuine. And I guess my thought there was like, I can imagine such a person. And they're not a person who just adopts the values of the local culture. And in fact, that would be kind of rude.
I think if someone came to you and just pretended to have your values, you'd be like, that's kind of off-putting. And it's someone who's like very genuine. And insofar as they have opinions and values, they express them. They're willing to discuss things, though. They're open minded. They're respectful.
And so I guess I had in mind that the person who like if we were to aspire to be the best person that we could be in the kind of circumstance that a model finds itself in, how would we act? And I think that's the kind of the guide to the sorts of traits that I tend to think about.
I think that people think about values and opinions as things that people hold sort of with certainty and almost like preferences of taste or something, like the way that they would, I don't know, prefer like chocolate to pistachio or something. But actually I think about values...
and opinions as like a lot more like physics than I think most people do I'm just like these are things that we are openly investigating there's some things that we're more confident in we can discuss them we can learn about them um and so I think in some ways though like it's ethics is definitely different in nature but has a lot of those same kind of qualities um
Showing 21–40 of 254 · page 2 of 13
← Previous
Next →