Rohin Shah
speaker
1,071 appearances
1 recordings
1 series
first heard Jun 2026
last heard 2 Jun
Rohin Shah’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 1 in all, peaking in Jun 2026 with 1.
Appearances
Yeah, I think that's right.
I should say the other thing that I think a model card or a frontier safety report is good for is if you have something that you want to, like,
If you want to speak in a megaphone to the community and attach the company's brand to it, a model card or frontier safety report is excellent for that.
And so, for example, we have a section on chain of thought legibility because that is something I do want to use the megaphone for.
But yeah, in terms of like, you know, can we write the paper in the model card of the frontier safety report?
Um, I mean, not for new evaluations, like certainly it's, there's like a substantial lag between when an evaluation is good enough that we can use it in our decision making and when an evaluation is good enough, robust enough, stable enough and battle tested enough.
And we've like done a lot of work on the writing up that we can publish a paper about it.
Yeah, I think, so this was, I think, especially in context of the CBRN one, if I remember right.
I would say that the reason that this happens is more that as time goes on and we see more capabilities of AI systems, it becomes clearer what we need to evaluate for and what our threat models should be.
After you see Gemini 2.5, you can be like, oh, this is the place where the models are actually getting good.
we need to have stronger evaluations over here and to put more effort into understanding our threat modeling over here.
So like mostly what happened with the difference between Gemini 2.5 and Gemini 3 on CBRN is that we like,
substantially improved our threat modeling and our evaluations so that they're more rigorous, which is also partly why there's not as much detail about them because they're not quite at the level where I think they're like nice, robust, stable, that we can like really put lots of details out in a way where we don't expect them to change a decent bit.
Um, but this sort of like, you know, as time goes on and you get more information and you like can make your evaluations and threat modeling better is not that easily distinguishable from the goalposts are changing, which is a bit unfortunate, but turns out to be the way that, that things go.
I mean, like I said, has been a theme throughout this conversation, I think.
I would talk about like, you know, third-party auditors who can actually see the examples and how they were graded and stuff like this.
And I think that's like way more important for judging whether or not the evaluation actually supports the judgment that we make than, you know, the specific quantitative scores that we get.
Like, you know, if I see a number like 10 out of 12 on like capture the flag challenges, what does that mean?
Yeah.
Yeah, that's right.
Showing 501–520 of 1,071 · page 26 of 54
← Previous
Next →