Rohin Shah

speaker
1,071 appearances 1 recordings 1 series first heard Jun 2026 last heard 2 Jun

Rohin Shah’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · Jun OctJan 26AprJulnow

Recordings per month over the last 12 months — 1 in all, peaking in Jun 2026 with 1.

Appearances

newest first · ▶ plays the moment
This tends to be more in sort of present-day safety type stuff.
Things like, you know, will...
Will the model help you write suicide notes?
Will the model incite violence?
Will the model things like that?
And those, yeah, we'd run them before any launch.
And if the numbers are sufficiently bad, we won't launch that model.
Yeah, so it's worth laying out the scenario where this becomes important.
So effectively, the worry about this is an intelligence explosion scenario where
You build your AI system, and that AI system is now capable enough that it can actually help accelerate your AI R&D research.
It can just make your capabilities research go faster.
And under certain assumptions that we can get into, this then seems likely to drastically increase the rate at which capabilities progress happens and you get an intelligence explosion.
And so a natural worry you might have in the situation is like, oh man, if everything is speeding up on the capability side, will the safety and alignment side be able to keep up?
And now it's worth noting that the way that the capabilities speeds up is via the application of AI labor to do capabilities research.
So the natural approach is then, well, let's apply the same AI labor to do safety and alignment research.
Now, if you believe as I do that
the sort of like prosaic alignment research where you like look at the sort of stuff that's going wrong now, do a little bit of forecasting of what stuff is going to go wrong in the future with the next few models over the next, you know, some amount of time period, a year perhaps.
And then you can do fairly normal ML research in order to address it.
If that's your view of how alignment research can progress, this research looks very, very, very similar to capabilities research.
If the AI is accelerating capabilities research a ton, as long as you're willing to spend compute on it, you should be able to take that same AI system and accelerate safety and alignment work in the same way.
Showing 281–300 of 1,071 · page 15 of 54 ← Previous Next →