Ankit Gupta

speaker
237 appearances 1 recordings 1 series first heard Jul 2026 last heard 17 Jul

Ankit Gupta’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · Jul OctJan 26AprJulnow

Recordings per month over the last 12 months — 1 in all, peaking in Jul 2026 with 1.

Appearances

newest first · ▶ plays the moment
And so ultimately what it comes down to is we are trying to still find a new policy pie.
And along the way, we will use the learning models in various capacities.
This is standard RL to estimate the value function given the rewards we're receiving.
And then where world models come in is a way of incorporating all of those into some sort of joint modeling of the state and action distribution so that we can make more intelligent policies off of it.
Yeah.
For those of us who kind of saw our, our diffusion model series, often people these days use video diffusion for exactly this.
And so here you're saying, you know, what's also in vogue now is jointly training these versus separately training them.
Exactly.
Okay, so I think that's a really good segue.
I think, why don't we now motivate everything we just described through a series of increasingly complex environments?
So I'll contend that I think the right set of environments for us to consider is chess, followed by Go, followed by self-driving, followed by robotics.
In any given state, there's only eight-ish moves you could do.
You say it's tractable even though there's a really big state space here.
Yeah.
Why don't we talk about that for just a second?
I think this is a really important point.
When you say it's tractable, you're specifically referring to the action space being small because it affects the kind of like combinatorial expansion here.
Should we talk about that for just a second?
Or maybe we can add go and then kind of contrast the two.
Quite intractable.
Showing 41–60 of 237 · page 3 of 12 ← Previous Next →