Chris Olah
speaker
254 appearances
1 recordings
1 series
first heard Nov 2024
last heard Nov 2024
Chris Olah’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsNo recordings in the last 12 months.Older appearances are listed below; set an alert to hear about the next one.
Appearances
But I don't know that other humans are conscious. I think they are. I think there's a really high probability they are. But there's basically just a probability distribution that's usually clustered right around yourself. And then it goes down as things get further from you. And it goes immediately down. You're like, I can't see what it's like to be you.
I've only ever had this one experience of what it's like to be a conscious being. And so my hope is that we don't end up having to rely on like a very powerful and compelling answer to that question. I think a really good world would be one where basically there aren't that many trade-offs. Like it's probably not that costly to make Claude a little bit less apologetic, for example.
It might not be that costly to have Claude, you know, just like not take abuse as much, like not be willing to be like the recipient of that. In fact, it might just have benefits for both the person interacting with the model and if the model itself is like, I don't know, like extremely intelligent and conscious, it also helps it. So... That's my hope.
If we live in a world where there aren't that many trade-offs here and we can just find all of the kind of like positive sum interactions that we can have, that would be lovely. I mean, I think eventually there might be trade-offs and then we just have to do a difficult kind of like calculation.
Like it's really easy for people to think of the zero sum cases and I'm like, let's exhaust the areas where it's just basically costless to assume that if this thing is suffering, then we're making its life better.
Yeah, I think we added a thing at one point to the system prompt where basically if people were getting frustrated with Claude, it got the model to just tell them that it can do the thumbs down button and send the feedback to Anthropic.
And I think that was helpful because in some ways it's just like, if you're really annoyed because the model's not doing something you want, you're just like, just do it properly. Yeah. The issue is you're probably like, you know, you're maybe hitting some like capability limit or just some issue in the model and you want to vent.
And I'm like, instead of having a person just vent to the model, I was like, they should vent to us because we can maybe like do something about it.
Yeah. I mean, there's lots of weird responses you could do to this. Like if people are getting really mad at you, I don't try to diffuse the situation by writing fun poems, but maybe people wouldn't be happy with it.
I think that's feasible. I have wondered the same thing. And I could actually, not only that, I could actually just see that happening eventually where it's just like the model ended the chat.
Yeah, it feels very extreme or something. Like, the only time I've ever really thought this is, I think that there was like a, I'm trying to remember, this was possibly a while ago, but where someone just like kind of left this thing interact, like maybe it was like an automated thing interacting with Claude.
And Claude's like getting more and more frustrated and kind of like, why are we like having, and I was like, I wish that Claude could have just been like, I think that an error has happened and you've left this thing running. And I'm just like, what if I just stop talking now? And if you want me to start talking again...
actively tell me or do something but yeah it's like um it is kind of harsh like i'd feel really sad if like i was chatting with claude and claude just was like i'm done that would be a special touring test moment where claude says i need a break for an hour and it sounds like you do too and just leave close the window
I mean, obviously, it doesn't have a concept of time, but you can easily... I could make that right now, and the model would just... I could just be like, oh, here's the circumstances in which you can just say the conversation is done. And I mean, because you can get the models to be pretty responsive to prompts, you could even make it a fairly high bar.
It could be like, if the human doesn't interest you or do things that you find intriguing... and you're bored, you can just leave. And I think that like, um, it would be interesting to see where Claude utilized it, but I think sometimes it would, it should be like, oh, this is like this programming task is getting super boring.
Uh, so either we talk about, I don't know, like either we talk about fun things now or I'm just, I'm done.
I think that we're going to have to navigate a hard question of relationships with AIs, especially if they can remember things about your past interactions with them. I'm of many minds about this because I think the reflexive reaction is to be kind of like, this is very bad and we should sort of like prohibit it in some way.
I think it's a thing that has to be handled with extreme care for many reasons. Like one is, you know, like this is a, for example, if you have the models changing like this, you probably don't want people performing like long-term attachments to something that might change with the next iteration.
At the same time, I'm sort of like, there's probably a benign version of this where I'm like, if you like, you know, for example, if you are like unable to leave the house and you can't be like, you know, talking with people at all times of the day. And this is like something that you find nice to have conversations with.
You like it, that it can remember you and you genuinely would be sad if like you couldn't talk to it anymore. Yeah. there's a way in which I could see it being like healthy and helpful. So my guess is this is a thing that we're going to have to navigate kind of carefully.
Showing 201–220 of 254 · page 11 of 13
← Previous
Next →