Yoshua Bengio
speaker
2,216 appearances
5 recordings
5 series
first heard Oct 2018
last heard 7 May
Yoshua Bengio’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 3 in all, peaking in May 2026 with 1.
Appearances
So once you have
this predictor, you can actually just produce a policy out of it by asking these questions about actions to achieve goals.
Yes, yes, exactly.
So the important point here is
to make sure that there's no reward hacking, like over-optimization of a policy.
The problem that could occur if you separately train a policy and a guardrail, if the policy is like very smart,
compared to the guardrail is that it could do the same thing as what jailbreaks do.
It could find questions or contexts, proposed actions for which the guardrail is simply going to produce a wrong answer, which means the policy is going to be able to bypass the guardrail.
And the reason is that neural nets are not, are never gonna be perfect.
They're always gonna make mistakes.
So how do we get around that?
Well, there are two aspects of this.
So one is that in the scientist AI,
we can not just produce those estimated probabilities, but also a confidence interval around the probabilities.
In other words, the system will estimate how much it trusts its own answers.
So why is that important?
Because
If the neural net asks a question for which its answer is not reliable, but it knows that its answer is not reliable, then it can just reject that question.
Now, there's another reason why the agentic scientist AI is going to be safe.
That has to do with
Showing 141–160 of 2,216 · page 8 of 111
← Previous
Next →