Francois Chaubard

speaker
744 appearances 1 recordings 1 series first heard Jul 2026 last heard 17 Jul

Francois Chaubard’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · Jul OctJan 26AprJulnow

Recordings per month over the last 12 months — 1 in all, peaking in Jul 2026 with 1.

Appearances

newest first · ▶ plays the moment
And the reason why that's important is because what's actually happening here is this is the discounted reward following policy pi.
Correct.
And that means that when I'm in this state, I will take this action, and then I'll end up in this to sc plus one, and then I'll take this action, and it's taking it greedy.
And so that's the value with respect to pi.
Right.
And so your standard kind of setup for this is what I'm always trying to get to at the end of the day is some joint distribution, which would be ST plus one given where I'm at now, where I'm at now.
And then this factorizes with chain rule simply to my pi, my policy, AT given ST.
Yep.
And my world model.
This is usually represented with theta.
And this is my world model, which would be ST plus 1 given ST and AT.
And so, and these are typically learned, uh, uh, separately and like, and like, you can imagine, in fact, actually, you can actually learn this.
This is a video generation model and I have the frame ST and I predict the next frame ST plus one.
Right.
And then, and we'll get into this.
Yeah.
And then what you can do, and this is like the in vogue thing to do since Dana jar and, and, um, the dreamer.
paper series from v1 to v4 is do action conditioning later like similar to clip where we will inject this like input head input tail to come into the model to influence and enable the world model to have embodiment what does that mean it means that not only can i predict like as a plant or tree on the on growing on the
the side of the building, I can like see the world passing by, but I can actually influence it and I can change the world and I can learn that with AT.
And it's far fewer samples to do this post action conditioning if I already have a really good ST to SC plus one world model.
Showing 161–180 of 744 · page 9 of 38 ← Previous Next →