Chris Olah
speaker
254 appearances
1 recordings
1 series
first heard Nov 2024
last heard Nov 2024
Chris Olah’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsNo recordings in the last 12 months.Older appearances are listed below; set an alert to hear about the next one.
Appearances
So I think of it as patching issues and slightly adjusting behaviors to make it better and more to people's preferences. So yeah, it's almost like the less robust but faster way of just like solving problems.
Yeah, no, I think that that is actually really interesting because I remember seeing this happen when people were flagging this on the internet. And it was really interesting because I knew that, at least in the cases I was looking at, it was like nothing has changed. Literally, it cannot. It is the same model with the same system prompt, same everything.
I think when there are changes, then it makes more sense. So one example is there... you know, you can have artifacts turned on or off on cloud.ai. And because this is like a system prompt change, I think it does mean that the behavior changes a little bit.
And so I did flag this to people where I was like, if you love cloud's behavior and then artifacts was turned from like the, I think you had to turn on to the default, just try turning it off and see if the issue you were facing was that change.
But it was fascinating because, yeah, you sometimes see people indicate that there's like a regression when I'm like, there cannot like I, you know, and like I'm like, I'm again, you know, you should never be dismissive. And so you should always investigate. You're like, maybe something is wrong that you're not seeing. Maybe there was some change made.
But then you look into it and you're like, this is just the same model doing the same thing. And I'm like, I think it's just that you got kind of unlucky with a few prompts or something. And it looked like it was getting much worse. And actually, it was just, yeah, it was maybe just like luck.
And you can get randomness is like the other thing. And just trying the prompt like, you know, four or 10 times, you might realize that actually like possibly, you know, like two months ago you tried it and it succeeded. But actually if you tried it, it would have only succeeded half of the time. And now it only succeeds half of the time. And that can also be an effect.
This feels like an interesting psychological question. I feel like a lot of responsibility or something. I think that's, you know, and you can't get these things perfect. So you can't like, you know, you're like, it's going to be imperfect. You're going to have to iterate on it. Yeah. I would say more responsibility than anything else.
Though I think working in AI has taught me that I like, I thrive a lot more under feelings of pressure and responsibility than I'm like, it's almost surprising that I went into academia for so long. Cause I'm like this, I just feel like it's like the opposite. Things move fast and you have a lot of responsibility and I quite enjoy it for some reason.
Yeah, I think that's the thing. It's something like if you do it well, like you're never going to get it perfect.
But I think the thing that I really like is the idea that like when I'm trying to work on the system prompt, you know, I'm like bashing on like thousands of prompts and I'm trying to like imagine what people are going to want to use Cloud for and kind of, I guess like the whole thing that I'm trying to do is like improve their experience of it. And so maybe that's what feels good.
I'm like, if it's not perfect, I'll like, you know, I'll improve it. We'll fix issues. But sometimes the thing that can happen is that you'll get feedback.
from people that's really positive about the model um and you'll see that something you did like like when i look at models now i can often see exactly where like a trait or an issue is like coming from and so when you see something that you did or you were like influential in like making like i don't know making that difference or making someone have a nice interaction it's like quite meaningful um
But yeah, as the systems get more capable, this stuff gets more stressful because right now they're like not smart enough to pose any issues. But I think over time it's going to feel like possibly bad stress over time.
I think I use that partly. And then obviously we have like, so people can send us feedback, both positive and negative about things that the model has done. And then we can get a sense of like areas where it's like falling short. Internally, people like work with the models a lot and try to figure out areas where there are like gaps.
And so I think it's this mix of interacting with it myself, seeing people internally interact with it, and then explicit feedback we get. And then I find it hard to know also, like, you know, if people are on the internet and they say something about Claude and I see it, I'll also take that seriously. I don't know.
I'm pretty sympathetic in that they are in this difficult position where I think that they have to judge whether some things actually seem risky or bad and potentially harmful to you or anything like that. So they're having to like draw this line somewhere.
And if they draw it too much in the direction of like, I'm going to, you know, I'm kind of like imposing my ethical worldview on you, that seems bad. So in many ways, like I like to think that we have actually seen improvements on this across the board, which is kind of interesting because that kind of coincides with like, For example, like adding more of like character training.
And I think my hypothesis was always like the good character isn't again one that's just like moralistic. It's one that is like like it respects you and your autonomy and your ability to like choose what is good for you and what is right for you. Within limits, this is sometimes this concept of like courageability to the user. So just being willing to do anything that the user asks.
And if the models were willing to do that, then they would be easily like misused. You're kind of just trusting. At that point, you're just saying the ethics of the model and what it does is completely the ethics of the user.
Showing 121–140 of 254 · page 7 of 13
← Previous
Next →