Chris Olah

speaker
254 appearances 1 recordings 1 series first heard Nov 2024 last heard Nov 2024

Chris Olah’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
No recordings in the last 12 months.Older appearances are listed below; set an alert to hear about the next one.

Appearances

newest first · ▶ plays the moment
You want models in the same way you want them to understand physics.
You kind of want them to understand all values in the world that people have and to be curious about them and to be interested in them and to not necessarily like pander to them or agree with them because there's just lots of values where I think almost all people in the world, if they met someone with those values, they would be like, that's abhorrent. I completely disagree.
And so, again, maybe my thought is, well, in the same way that a person can. I think many people are thoughtful enough on issues of ethics, politics, opinions, that even if you don't agree with them, you feel very heard by them. They think carefully about your position. They think about its pros and cons. They maybe offer counter considerations. So they're not dismissive, but nor will they agree.
If they're like, actually, I just think that that's very wrong, they'll say that. I think that in Claude's position, it's a little bit trickier because you don't necessarily want to like, if I was in Claude's position, I wouldn't be giving a lot of opinions. I just wouldn't want to influence people too much.
I'd be like, you know, I forget conversations every time they happen, but I know I'm talking with like potentially millions of people who might be like really listening to what I say. I think I would just be like, I'm less inclined to give opinions. I'm more inclined to think through things or present the considerations to you or discuss your views with you.
But I'm a little bit less inclined to affect how you think because it feels much more important that you maintain autonomy there.
And kind of like walking that line between convincing someone and just trying to like talk at them versus like drawing out their views, like listening and then offering kind of counter considerations. Yeah. And it's hard.
I think it's actually a hard line where it's like, where are you trying to convince someone versus just offering them like considerations and things for them to think about so that you're not actually like influencing them. You're just like letting them reach wherever they reach. And that's like a line that is difficult, but that's the kind of thing that language models have to try and do.
Yeah, I think that most of the time when I'm talking with Claude, I'm trying to kind of map out its behavior in part. Like obviously I'm getting like helpful outputs from the model as well. But in some ways, this is like how you get to know a system, I think, is by like probing it and then augmenting like, you know, the message that you're sending and then checking the response to that.
So in some ways it's like how I map out the model. I think that people focus a lot on these quantitative evaluations of models. And this is a thing that I've said before, but I think in the case of language models, A lot of the time, each interaction you have is actually quite high information. It's very predictive of other interactions that you'll have with the model.
And so I guess I'm like, if you talk with a model hundreds or thousands of times, this is almost like a huge number of really high quality data points about what the model is like. in a way that lots of very similar but lower quality conversations just aren't, or questions that are just mildly augmented and you have thousands of them might be less relevant than 100 really well-selected questions.
I think it's almost like everything. Because I want like a full map of the model, I'm kind of trying to do... the whole spectrum of possible interactions you could have with it. So like one thing that's interesting about Claude, and this might actually get to some interesting issues with RLHF, which is if you ask Claude for a poem,
I think that a lot of models, if you ask them for a poem, the poem is like fine. You know, usually it kind of like rhymes and it's, you know, so if you say like, give me a poem about the sun, it will be like, yeah, it'll just be a certain length. It'll like rhyme. It will be fairly kind of benign.
Um, and I've wondered before, is it the case that what you're seeing is kind of like the average, it turns out, you know, if you think about people who have to talk to a lot of people and be very charismatic, um, One of the weird things is that I'm like, well, they're kind of incentivized to have these extremely boring views. Because if you have really interesting views, you're divisive.
And a lot of people are not going to like you. So if you have very extreme policy positions, I think you're just going to be less popular as a politician, for example. Um, and it might be similar with like creative work.
If you produce creative work that is just trying to maximize the kind of number of people that like it, you're probably not going to get as many people who just absolutely love it. Um, because it's going to be a little bit, you know, you're like, Oh, this is the app. Yes, this is decent.
And so you can do this thing where I have various prompting things that I'll do to get Claude to, I'll do a lot of like, this is your chance to be fully creative. I want you to just think about this for a long time. And I want you to create a poem about this topic that is really expressive of you, both in terms of how you think poetry should be structured, et cetera.
And you just give it this really long prompt. And its poems are just so much better. They're really good. And I don't think I'm someone who is... I think it got me interested in poetry, which I think was interesting. I would read these poems and just be like, I love the imagery. I love... And it's not trivial to get the models to produce work like that, but when they do, it's really good.
So I think that's interesting, that just encouraging creativity and for them to move away from the kind of standard, immediate reaction that might just be the aggregate of what most people think is fine can actually produce things that, at least to my mind, are probably a little bit more divisive, but I like them.
I really do think that philosophy has been weirdly helpful for me here more than in many other respects. So in philosophy, what you're trying to do is convey these very hard concepts. One of the things you are taught is like... And I think it is because... I think it is an anti-bullshit device in philosophy. Philosophy is an area where you could have people bullshitting and you don't want that.
Showing 41–60 of 254 · page 3 of 13 ← Previous Next →