Francois Chaubard
speaker
744 appearances
1 recordings
1 series
first heard Jul 2026
last heard 17 Jul
Francois Chaubard’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 1 in all, peaking in Jul 2026 with 1.
Appearances
Unkit's drone is to try to hit me.
Right.
And so now, let's just call this the... This would be... Now, we're definitely not deterministic.
We're stochastic.
And stochastic and non-differentiable.
Yeah.
And in this case...
My state transition, what is st plus one?
It's going to be my state I'm in now, my thrust, and what unkit's going to do.
Right.
And these, it was all differentiable until this new variable.
Yeah, and I can't back prop through your brain to say what you're going to do with your little drone controller, right?
It's completely non-differentiable now.
And I'm resorting, and I have to resort to this awful area called reinforcement learning, which is just super brutal, and it's sprawling, and there's so many different things.
And you'll hear things like when you study initial...
reinforcement learning called value iteration or policy iteration.
Um, and there's DQN or deep Q learning or just Q learning.
Um, there's actor critic, there's all this bag of stuff.
Yeah, and so that's basically the main thing is you're going to start talking about this as a model, where I'm going to introduce this psi to say that this is going to be some model that's going to take in these things and then output this, and that we're going to train it over many, many instantiations of this.
And that's to get a better and better world model.
Showing 121–140 of 744 · page 7 of 38
← Previous
Next →