Rohin Shah

speaker
1,071 appearances 1 recordings 1 series first heard Jun 2026 last heard 2 Jun

Rohin Shah’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · Jun OctJan 26AprJulnow

Recordings per month over the last 12 months — 1 in all, peaking in Jun 2026 with 1.

Appearances

newest first · ▶ plays the moment
Thanks a lot, Rob.
And that's a very generous intro.
And yeah, in the interest of being very opinionated, I do want to emphasize that these opinions are mine alone.
They're not meant to represent the opinions of Google or Google DeepMind.
There's a few different disjunctive reasons.
I don't feel like there's one particular thing.
Probably the highest level bit is that I don't feel like there is any particularly compelling argument that this is the thing that happens by default.
I think there's a lot of arguments that are suggestive that maybe it could happen, such that you should find it plausible.
I think that's sufficient to justify a significant amount of effort into averting it, which is why I work in the area that I do work in.
But none of them really rise to the level of like, oh, yeah, now I'm expecting this to happen by default.
I think they're like every argument that I've seen, they're like pretty significant holes one could poke if you try to take them as arguments for this is what happens likely as opposed to this is a plausible thing that could happen.
So if you take the...
Kotra and Carl Smith arguments of like you know we'll well they have like a variety of arguments but I think in fact one of the common ones which you pointed to is like you know we might accidentally train them to be deceptive totally true I agree that is pretty likely something that at least could happen pretty easily and like maybe it's even likely but
you know, we're not going to do reinforcement learning over the course of one year trajectories.
We're going to do, maybe we're going to do reinforcement learning over like a week or a month at most.
So like, it seems very plausible that like what the, like, I think the default prediction you should have for that is like what you train, what the AI system learns to do is like,
You know, I'm going to take opportunities to reward hack, seek reward as much as possible that would allow me to get a high score after a week or something like that, or whatever the time horizon actually was.
And this is very, very different from like the sort of ambitious misaligned goal that you need in order to motivate convergent instrumental sub goals to the point of like, now my job is to take over the world.
That's what I need in order to achieve my goal.
Like those really do seem like they need to be like significantly longer horizon goals.
Showing 1–20 of 1,071 · page 1 of 54 Next →