Francois Chaubard
speaker
744 appearances
1 recordings
1 series
first heard Jul 2026
last heard 17 Jul
Francois Chaubard’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 1 in all, peaking in Jul 2026 with 1.
Appearances
And so basically what the key idea in JIPA, if I have an image one and I have image two and I have image three,
I can do my world modeling, my world modeling of st plus one, given st and at in pixel space, and have this is, let's say, at time, t t plus one.
plus two, etc, etc.
And I have to actually predict now the full image that's extremely expensive from a computation standpoint.
And also from like a sample efficiency standpoint, what I can do instead is put this through some comnet, some encoder, some encoder, and then I'll get a latent for t. And I'll have a latent for t plus one of a latent for z t plus two.
So
And then I'll have from this, from z t, I want to predict z t plus one hat.
And my goal is to make this and this make my loss function will be something very simple, like
I want to minimize this.
That's it.
Now this doesn't work.
This collapses hard.
And so what happens is basically just, if you, if it just predicts zero, which the model will learn to do.
And I'm actually incorporating this into my current research right now.
And so what you need to do is something called SIGREG, or this is one technique, VICREG is another, where basically I add this another term that basically says, I want the, over a large enough batch size, I want the distribution of Z T plus one to follow a Gaussian distribution.
I mean, not in the same location.
And if it's zero, it can't be this, right?
Because then this is non-zero.
And so maybe I think that there's probably this or something like that.
But basically this prevents it from modal collapse and it makes it do something good.
Showing 561–580 of 744 · page 29 of 38
← Previous
Next →