Rob Wiblin

speaker
1,787 appearances 3 recordings 1 series first heard Sep 2024 last heard May 2025

Rob Wiblin’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
No recordings in the last 12 months.Older appearances are listed below; set an alert to hear about the next one.

Appearances

newest first · ▶ plays the moment
Basically, these are I I think m most listeners will be familiar with the idea of
Evals, we're we're trying to measure what capabilities does a model have, uh, as it's you know soon after it's being being trained, so we're not kind of blindsided later uh and we know what what safeguards are necessary.
Having accurate evals that can figure out what a model can and can't do uh really facilitates uh what's recently started being called if-then commitments.
So it's much easier to agree that if we have a model that is able to do X powerful or interesting thing, then
Then we will put in place a safeguard that is is appropriate for making sure that that doesn't end poorly.
And I think these are some of the evals that uh that uh GDM is developing to slot into the into Google's frontier safety framework, which is going to be the I guess the approach that uh that um that DMind takes to uh deploying very powerful, uh potentially at some point uh AGI um uh while making it go making it go well and and avoiding significant downside.
Sides.
The paper describes these evals and offers reports or sorry, offers results from the previous model that GDN had, which I guess is Gemini 1.0.
I think currently we're on Gemini, Gemini 1.5.
And I guess there's five different categories of evals that you're thinking about.
Persuasion and deception, cyber capabilities, self-proliferation and self-reasoning.
And there's a
that is around, I guess, biological threats and things like that, which uh which you don't cover in this paper because it's kind of it's kind of its own thing that you want to do separately.
What what's an eval in the paper that you think is kind of uh new or interesting that uh that people would would want to hear about?
Yeah.
I guess it's a relatively easier case than than the real world one.
But uh but the I suppose w we wanna we wanna have relatively equ easy cases now so we can see if we're incrementing closer to to the models being being more self aware and able to take these these sorts of interventions um in in in real world cases.
Yeah.
Um I think I I I I I read this paper uh about a month or two ago.
Um I guess I
Showing 921–940 of 1,787 · page 47 of 90 ← Previous Next →