Chris Olah
speaker
254 appearances
1 recordings
1 series
first heard Nov 2024
last heard Nov 2024
Chris Olah’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsNo recordings in the last 12 months.Older appearances are listed below; set an alert to hear about the next one.
Appearances
And I think it's also like, I don't see a good, like, I think it's just a very, it reminds me of all of the stuff where it has to be just approached with like nuance and thinking through what is, what are the healthy options here? And how do you encourage people to
towards those while you know respecting their right to you know like if someone is like hey i get a lot of chatting with this model um i'm aware of the risks i'm aware it could change um i don't think it's unhealthy it's just you know something that i can chat to during the day i kind of want to just like respect that i personally think there'll be a lot of really close relationships i don't know about romantic but friendships at least and then you have to i mean there's so many fascinating things there just like you said you have to
I think it's also the only thing that I've thought consistently through this as like a, maybe not necessarily a mitigation, but a thing that feels really important is that the models are always like extremely accurate with the human about what they are.
It's like a case where it's basically like, if you imagine, like I really like the idea of the models, like say knowing like roughly how they were trained. And I think Claude will often do this. I mean, for like, there are things like,
part of the traits training included like what Claude should do if people basically like explaining like the kind of limitations of the relationship between like an AI and a human that it like doesn't retain things from the conversation and so I think it will like just explain to you like hey here's like I won't remember this conversation Um, here's how I was trained.
It's kind of unlikely that I can have like a certain kind of like relationship with you. And it's important to, you know, that it's important for like, you know, your mental wellbeing that you don't think that I'm something that I'm not. And somehow I feel like this is one of the things where I'm like, oh, it feels like a thing that I always want to be true.
I kind of don't want models to be lying to people because I if people are going to have like healthy relationships with anything, it's kind of important. Yeah. Like I think that's easier if you always just like know exactly what the thing is that you're relating to. It doesn't solve everything, but I think it helps quite a lot.
Well, it depends partly on like the kind of capability level of the model. If you have something that is like capable in the same way that an extremely capable human is, I imagine myself kind of interacting with it the same way that I do with an extremely capable human with the one difference that I'm probably going to be trying to like probe and understand its behaviors.
But in many ways, I'm like, I can then just have like useful conversations with it. You know, so if I'm working on something as part of my research, I can just be like, oh, like, which I already find myself starting to do. You know, if I'm like, oh, I feel like there's this like thing in virtue ethics, I can't quite remember the term. Like, I'll use the model for things like that.
And so I can imagine that being more and more the case where you're just basically interacting with it much more like you would an incredibly smart colleague. and using it for the kinds of work that you want to do as if you just had a collaborator. Or the slightly horrifying thing about AI is as soon as you have one collaborator, you have a thousand collaborators if you can manage them enough.
um in a way that pushes its limits understanding where the limits are yep so i guess what would be a question you would ask to be like yeah this is agi that's really hard because it feels like in order to it has to just be a series of questions like if there was just one question like you can train anything to answer one question extremely well yeah um in fact you can probably train it to answer like you know 20 questions extremely well
It's a hard question because part of me is like, all of this just feels continuous. Like if you put me in a room for five minutes, I'm like, I just have high error bars, you know? And then it's just like, maybe it's like both the probability increases and the error bar decreases. I think things that I can actually probe the edge of human knowledge of. So I think this with philosophy a little bit.
Sometimes when I ask the models philosophy questions, I am like, this is a question that I think no one has ever asked. Like it's maybe like right at the edge of like some literature that I know.
And the models will just kind of like when they struggle with that when they struggle to come up with a kind of like novel like I'm like I know that there's like a novel argument here because I've just thought of it myself.
So maybe that's the thing where I'm like I've thought of a cool novel argument in this like niche area and I'm going to just like probe you to see if you can come up with it and how much like prompting it takes to get you to come up with it.
And I think for some of these like really like right at the edge of human knowledge questions, I'm like you could not in fact come up with the thing that I came up with. I think if I just took something like that where I like I know a lot about an area and I came up with a novel issue or a novel like solution to a problem.
and I gave it to a model and it came up with that solution that would be a pretty moving moment for me because I would be like this is a case where no human has ever like it's not and obviously we see these with this with like more kind of like you see novel solutions all the time especially to like easier problems I think people overestimate you know novelty isn't like it's completely different from anything that's ever happened it's just like this is it can be a variant of things that have happened and still be novel
But I think, yeah, if I saw... The more I were to see completely novel work from the models, that would be... And this is just going to feel iterative. It's one of those things where there's never... It's like... And, you know, people, I think, want there to be a lucky moment. And I'm like, I don't know. Like, I think that there might just never be a moment.
It might just be that there's just like this continuous ramping up.
Yeah.
Showing 221–240 of 254 · page 12 of 13
← Previous
Next →