Rob Wiblin

speaker
1,787 appearances 3 recordings 1 series first heard Sep 2024 last heard May 2025

Rob Wiblin’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
No recordings in the last 12 months.Older appearances are listed below; set an alert to hear about the next one.

Appearances

newest first · ▶ plays the moment
What what other procedures can we add that are complementary to this one?
So with with with the standard evals, you you find out the results uh at a at the point that you're training the model, or hopefully soon after.
Um I guess this is a challenge with the persuasion ones where you have to involve humans and that slows things down.
But generally like pretty soon after.
Then you've got evals in the wild, which are uh they have like better external validity, you were saying, but they tend to come later because you've already ha you have to have distributed the model to see how it's used in the real world.
Um and you and but you talk in the paper about this there's other things.
The ideal would be that we not only know uh
That we have an ability when we have it, but we would know exactly when it's going to arrive years into the future.
So we have this foresight, and we know that some um some preventative measures are not necessary now because we're not going to have this ability for five years time.
Um so we we avoid like not um what you know wasting our effort on something that's that's not necessary yet.
Um but we also know exactly what uh preventative measures we need to be putting in place now, uh, because we know exactly when an ability is gonna arrive.
Listen, so this show are probably going to be aware of super forecasting techniques about the possibility of using prediction markets and aggregating forecasts from different experts and lay people to try to try to predict what's going to happen when.
Is there much low-hanging fruit here on the forecasting, actual real-world app usefulness of AI models that is not being taken yet that you would like to see people uh uh you know actually actually get exploit?
Maybe it's surprising to think given the importance of this topic and its tractability or um you know the potential to to just do all kinds of different work figuring out when when AR models will actually be economically and like practically useful for all of these different tasks.
It it's surprising that people aren't doing this already, uh, 'cause you don't have to be inside DeepMind to to to to to run these tournaments or to or to or to aggregate these judgments.
But uh so I guess this is maybe an invitation for for listeners to to get involved with that.
Yeah.
I guess the the thing that's whack about forecasting abilities is that
With these scaling laws, we're we're we're so good at anticipating uh when we'll have a particular level of loss, this like measure of inaccuracy in the model.
But that hasn't translated into us being able to predict when will people actually want to use this model for for for for a given task.
Showing 981–1000 of 1,787 · page 50 of 90 ← Previous Next →