Rohin Shah

speaker
1,071 appearances 1 recordings 1 series first heard Jun 2026 last heard 2 Jun

Rohin Shah’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · Jun OctJan 26AprJulnow

Recordings per month over the last 12 months — 1 in all, peaking in Jun 2026 with 1.

Appearances

newest first · ▶ plays the moment
And they will sometimes say things like, oh, it was really good to get more details about how exactly you evaluate for such and such risks.
And then I will say something like, great, and we also have this paper that goes into more details of the evaluation.
I think maybe, probably it was even cited in the model card in the Frontier Safety Report.
Did you read it?
And then usually they will not even have heard about this paper.
And so I kind of don't believe them.
when they say that these details actually matter to them.
Again, I do think publishing papers important, doesn't need to be tied to the model cards.
And like some people who I've talked to who have, who are especially more on the technical side and the one people actually building these evaluations do actually find the papers useful and helpful for them to build on.
And so I do want to continue publishing the papers where we can.
Yeah.
So I think actually now GDM overall is pretty evenly split between London and the Bay Area, but the safety team has historically been based in London.
We're expanding into the Bay Area, but I would still say that the locus of attention is in London.
And I think in practice, just a lot of research papers and ideas end up spreading in the community via word of mouth, which we are a little bit less plugged into, which is a bit unfortunate.
Yeah, so I think the pitch for this work is that you can train a model that doesn't do a particular kind of reward hacking, even if the overseer would not be able to notice that reward hack if it had happened.
And so in contrast to things like amplified oversight or debate, those are ones which rely on making the overseers smart enough or capable enough that they can detect the reward hack and penalize it.
This one doesn't rely on detection at all, so it's actually a bit more robust.
How does this work?
What is this magic that we managed to do?
This is basically doable based on the idea of myopic optimization.
Showing 561–580 of 1,071 · page 29 of 54 ← Previous Next →