Yoshua Bengio
speaker
2,216 appearances
5 recordings
5 series
first heard Oct 2018
last heard 7 May
Yoshua Bengio’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 3 in all, peaking in May 2026 with 1.
Appearances
So whatever experience has been collected, by the way, it doesn't have to be
what people in RL call on policy.
So it could use the experience of any agent or anything that is observed that is not just agents doing things, but just observing things in the world.
All of that is data as far as it's concerned, that helps it both understand the world and construct the consequences.
And so one of the consequences, oh, what would happen if I do this action?
And what action will maximize the probability of achieving some goal?
In a way, it's closer to model-based reinforcement learning, where you are able to use your whole experience to come up with a policy by opposition to something that is fully interactive all the time.
So in the scientist AI, currently you would need to retrain it or fine-tune it with the new data if you use it and it produces new consequences and new observations.
But we can ride on the same research that companies and academia is working on, what's called continual learning.
So what happens when there's new information coming?
And of course, you could put it in the context, like in the input window, and the scientists there, you could do the same thing.
But at some point, you'd like it to be integrated into somehow the weights of the system.
And that's what continual learning is trying to do.
But the good news with the scientist AI is it's facing these same problems that current AIs are facing, and the solutions that are being explored will be applicable to the scientist AI as well.
Yeah, that is why I think we can do it pretty quickly.
And it's more like a matter of having the right resources for training.
And yeah, because the training objective is different, we do need to try it out and see how it works.
But fundamentally, it isn't so different, for example, from maximum likelihood training.
which is what we use in pre-training.
So in a way it's closer to the pre-training, except that, and we know that works really well, by the way, it's actually working better than RL, which is harder.
Showing 521–540 of 2,216 · page 27 of 111
← Previous
Next →