Rob Wiblin

speaker
1,787 appearances 3 recordings 1 series first heard Sep 2024 last heard May 2025

Rob Wiblin’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
No recordings in the last 12 months.Older appearances are listed below; set an alert to hear about the next one.

Appearances

newest first · ▶ plays the moment
Not
well not very well.
Okay.
Pretty rubbish.
Um
I've heard you use this term evals in the wild, uh which uh s sounds related.
Uh w yeah, what are evals in the wild and maybe how could we get how could we get more out of them than we are now?
So you were saying earlier that
The challenge with all these evals is that it can be quite hard to fully elicit all of the capabilities that a model has.
It's like it's entirely possible to miss something because you haven't been eliciting that that skill in quite the right way.
And people might later on figure out that a model has an ability that wasn't picked up in the early evals.
So there's a significant uh false negative rate, basically.
And it seems unlikely that we're going to have such a such a science of evals that we're going to escape that issue any any time.
An issue that I've had with a lot of the different frameworks that are coming out that are related to using evals and then putting them into responsible scaling policies or whatever you call it is that everything ends up writing on the eval being accurate.
If and and if it fails to pick up a dangerous ability that's actually there, then the entire kind of the the the the the entire system falls apart.
I guess everyone everyone says it wouldn't you point this out.
Well, what we need is defense in depth.
We need to have different overlapping systems so that when one fails, uh, you know, the
Swiss cheese model, you get through one one one gap in the cheese, but uh but the next layer picks it up.
What depth are we thinking of adding?
Showing 961–980 of 1,787 · page 49 of 90 ← Previous Next →