Yoshua Bengio

speaker
2,216 appearances 5 recordings 5 series first heard Oct 2018 last heard 7 May

Yoshua Bengio’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · May OctJan 26AprJulnow

Recordings per month over the last 12 months — 3 in all, peaking in May 2026 with 1.

Appearances

newest first · ▶ plays the moment
what are the best predictions or the best actions in the case of an agent, given the available information, like the data set and context that is available?
And then the second question is, if I were to do experiments in the world in order to acquire more knowledge, what are the right actions that will increase my understanding of the world and reduce my uncertainty about the world?
And by the way, this is how scientists in general, not AI scientists, people that are doing biology or chemistry or physics, this is how they think.
They ask themselves, if I were to do this experiment, would it help me to disambiguate between these two theories?
And you can quantify this mathematically with something that's called information gain.
And it turns out that once you have a good predictor, like good probabilistic predictor, you can also turn this into a good estimator of how much information you would gain if you were to do this experiment or that experiment.
And now you could build an agentic system on top of the scientist AI, for example, that would tell you which experiment to do in order to obtain good information gain.
But of course you could also use a guardrail.
So you would like experiments that help to disambiguate between explanations and theories at the same time as not harming people.
But that's easy to do in the scientist AI, right?
We have this guardrail notion.
So the user goal here is acquire information.
The safety goal is, you know, don't harm people.
I'm just, of course, this cartoon.
And so you could get both.
But of course, you now enter into the realm of agentic system.
And yeah, the whole plan of the scientist AI includes how we can develop agentic systems on top of a non-agentic trustworthy predictor.
So you could absolutely train it with trajectories of what happens when a particular agent did this and what was observed, what the consequences are.
This is how it would learn to learn good conditional probabilities of what actions to do in order to achieve particular goals, including the safety goal.
So it would be a different kind of training than the reinforcement learning training, but it would be using the same resource.
Showing 501–520 of 2,216 · page 26 of 111 ← Previous Next →