Ankit Gupta

speaker
237 appearances 1 recordings 1 series first heard Jul 2026 last heard 17 Jul

Ankit Gupta’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · Jul OctJan 26AprJulnow

Recordings per month over the last 12 months — 1 in all, peaking in Jul 2026 with 1.

Appearances

newest first · ▶ plays the moment
I mean, you have bigger differences between embodiments than a Model 3 versus Y, and you have way bigger action spaces you have to sum up Model 1.
Okay, so now that we understand the basic setup here and why the action space problem is so big, why don't we talk a little bit about how world models actually fit into this?
You know, maybe first, I guess, what didn't work about the naive world models and how do we fix those?
And then let's kind of talk about some of the newest world modeling techniques.
And ideally, we would train it in that way you were describing of like somehow we would train a model just on these two.
And then later add this.
So the key thing there is you can basically use this
if you have some predictive model of this in that case, and eventually of this, you can use that as basically a synthetic training set to train your policy model and then basically fine tune it on real data later.
And the key unlock there, yeah, it uses synthetic data specifically on a model trained on just this sort of state transition type of thing.
Yes.
And this ends up being very convenient because it turns out
We as a society have a lot of this.
Yeah.
And we put out a video about diffusion models very recently in flow matching.
I imagine that now ties very closely to this, right?
Ultimately, the kind of current state of the art best way to do this on basically infinity data that we have available and can keep generating is using state of the art video diffusion slash flow matching.
And what I thought was really cool about this paper is that they do exactly this process where they have this joint model of state transitions and actions.
They train it by first instantiating it with the open source one video diffusion model.
And then it only takes them about 500 hours of teleop data, which is basically exactly this, to get it to be pretty good.
And they have a lot of clever tricks that allow it to be cross embodiment and working on scene tasks with relatively small amounts of data.
Showing 161–180 of 237 · page 9 of 12 ← Previous Next →