Francois Chaubard

speaker
744 appearances 1 recordings 1 series first heard Jul 2026 last heard 17 Jul

Francois Chaubard’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · Jul OctJan 26AprJulnow

Recordings per month over the last 12 months — 1 in all, peaking in Jul 2026 with 1.

Appearances

newest first · ▶ plays the moment
S N A N R N. And if you won, then all of these, uh, all the moves that black, if black one, all the moves that black did get plus all the moves that white did were minus one.
And they just, that's how they create their, um, their rollouts.
Right.
So let's probably under this policy p theta t. And we're going to overload t. But this is that instantiation.
We froze that model.
We froze that model.
And we play, I think it's like 70 games.
And we treat all of those.
And we're going to subsample a bunch of these state action results, state action results to train our, to update our policy.
in our world model, our transition model.
And what it's actually doing is we take in an st, we give it to some theta, and it wants to output the probability of st plus 1 being played.
which is our transition function, and the value of the current state.
And how do we get the value?
And so the value of the current state, well, both of them are coming out of the model, but basically the loss function, l,
theta is going to equal, and it's going to be really close to this control problem one is we have some V theta minus this Z, which we'll just call it our here squared, and then plus, actually, it's minus this pi, which I'll explain in a second, log p theta.
And I think they, everyone includes this, but they include it in the paper, so I'll include it there as well, which is the weight decay.
And so this is basically what our loss function is.
Then we'll play a bunch of these games.
Let's try to be a little bit organized here.
And so this is our setup.
Showing 261–280 of 744 · page 14 of 38 ← Previous Next →