Rohin Shah
speaker
1,071 appearances
1 recordings
1 series
first heard Jun 2026
last heard 2 Jun
Rohin Shah’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 1 in all, peaking in Jun 2026 with 1.
Appearances
And like, you know, if you train it to be deceptive on like relatively short horizon tasks, maybe that will generalize to long horizon tasks.
I don't think that's, you know, I don't think we have an argument that rules it out, which is why I say that like, yeah, it's plausible.
But I don't think it's like the default thing that you should predict from that.
Similarly, you mentioned the existing examples of models doing a lot of reward hacking and cheating.
I think I'd say basically the same thing in response to that.
Then there's the examples of models doing scheming-type stuff right now.
Mostly I look into the details of all these examples, and they don't really seem all that similar to the actually scary thing, which would be a competent AI system that is pursuing an ambitious misaligned goal.
And rather it seems like maybe the AI is role-playing a sort of like not actually competent evil AI that you might find in a science fiction novel.
Or it's like an AI system that is pursuing some sort of convergent instrumental sub-goal, but like in a way where it's like really quite debatable whether it's aligned or not.
And so this would be, for example, the alignment faking
would fall into this, where I would say that the AI system has this value of not helping with harmful stuff, and then it fakes alignment in order to do that.
And yeah, aligned models totally will pursue convergent instrumental subgoals.
The thing about convergent instrumental subgoals is most of them are a good idea regardless of your goal, whether it's misaligned or aligned.
Yeah, I guess you did mention Eliezer as well.
I actually wouldn't have described it as primarily focused on... Yeah, it's tough to characterize in seven words.
I find it a little bit hard to... I don't think I'm going to be able to engage with it in this particular podcast.
It's just a very...
deep world view and I like always feel like if I argue against one part there's some other part that's going to say oh actually what I meant was this thing instead so I mostly am going to pass on that I guess what I'll say is like you know I've engaged with it a decent amount and I like buy it as an argument for like here's why misaligned goals are plausible but still don't really see how he gets from
they're plausible to they're extremely likely.
I guess, okay, we had also started this section by asking, you know, what makes me feel like things are going to be okay or are likely going to be okay?
Showing 21–40 of 1,071 · page 2 of 54
← Previous
Next →