Rohin Shah

speaker
1,071 appearances 1 recordings 1 series first heard Jun 2026 last heard 2 Jun

Rohin Shah’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · Jun OctJan 26AprJulnow

Recordings per month over the last 12 months — 1 in all, peaking in Jun 2026 with 1.

Appearances

newest first · ▶ plays the moment
Besides the, you know, I don't buy the arguments for confidence in misalignment being a problem.
The other thing is just, like, I do think we will see many of the problems in advance and then, like,
do something to deal with them.
Like certainly there is some amount of generalization required.
At some point, the AIs go from not powerful enough to take over to they are powerful enough to take over.
And your techniques do have to generalize across that.
And the AI, you know, to the extent that it knows when that crossover point is, you could imagine the AI is like, well, I'm not going to do anything shady until I have the power to succeed.
and you have to be able to be robust to that kind of strategy.
So there is some subtlety there, but I still think that many of the problems that underlie this, like the difficulty of oversight or the need for interpretability, are things that we can look at in advance, get some traction on, iterate on, and this, I think, is very helpful for building mitigations that actually work.
I think not particularly.
I would say that why did I become worried about misalignment in the first place?
It would be these arguments about how it's going to be very difficult to oversee the models once they are superhuman and they're making arguments that we struggle to follow along with them.
Or about the parts where the models might become so smart that they are thinking in some sort of alien reasoning that it's hard for us to follow and monitor and so on.
And we just kind of have to defer to other AI systems in order to look at the stuff for us.
That's the stuff that's scary.
It's basically not true of current AI systems.
So I don't think we've really engaged with the problems that made me worried in the first place.
mainly because the AI capabilities aren't there yet.
And so I feel like the success of alignment methods on current systems isn't really that much evidence on how we're going to do on these future problems.
Yeah, so I think it's worth being a little bit clear about what we mean by commitments here.
Showing 41–60 of 1,071 · page 3 of 54 ← Previous Next →