Rohin Shah

speaker
1,071 appearances 1 recordings 1 series first heard Jun 2026 last heard 2 Jun

Rohin Shah’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · Jun OctJan 26AprJulnow

Recordings per month over the last 12 months — 1 in all, peaking in Jun 2026 with 1.

Appearances

newest first · ▶ plays the moment
And so there's actually just quite a lot of increased defense you can get by making sure that you are making your security also apply to AI systems as untrusted insiders.
So things that you would want to do could include having...
infrastructure that requires agents to have separate identities from the human employees on whose behalf they're going to be doing things.
So this would be like, you are starting an agent on some task, it doesn't get all the permissions that you as a human employee have.
It gets some much smaller set of permissions.
And then as it will need some permissions in order to do its task, it has to request each of those individually, write a justification for why it needs that particular permission.
This allows you to incriminate it afterwards if it was lying about it.
Well, it helps you incriminate it, maybe doesn't guarantee it.
And then similarly, you can have some monitoring where you have a chain of thought monitor right now, for example, that looks at what the model is doing, sees if there is anything potentially suspicious about this.
If there is any things that are sufficiently highly suspicious, get flagged to humans to be audited.
This allows you to also generate evidence of misalignment.
If the AI system is actually misaligned,
and you audit it and you're like, oh my gosh, yeah, this AI system totally was trying to inject a security vulnerability and then would have exploited it in order to exfiltrate its weights.
That's a big deal.
I'm not sure I expect that to ever happen.
It seems quite plausible that models will never be misaligned.
So if we get that sort of evidence, I think that would change my mind a big deal and many other stakeholders as well.
So I think that's quite crucial to get.
So yeah, lots of stuff like this.
And I think it's important to do and very nascent.
Showing 701–720 of 1,071 · page 36 of 54 ← Previous Next →