Ankit Gupta
speaker
237 appearances
1 recordings
1 series
first heard Jul 2026
last heard 17 Jul
Ankit Gupta’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 1 in all, peaking in Jul 2026 with 1.
Appearances
intuitive ability to think as coming from some implicit world model we have in our heads encoded by genetics and our ability to learn and whatever else.
It seems like models can do surprisingly intelligent things despite not having an explicit world model when it comes to natural language.
When they're just talking, it seems like, you know, maybe under the hood, deep inside the weight somewhere, there's some kind of implicit understanding of the world, but there isn't an explicit representation of that.
But it seems like in certain domains, especially in robotics and self-driving, as we'll talk about, that sort of breaks down.
And maybe it would be helpful now to just think a little bit about and sort of define some of the pieces of what makes it challenging in these different domains.
And then we can use that to kind of build up to why it's particularly hard in things like self-driving and robotics to get these types of predictive models to work.
This is like a very fundamental for context, you know, this equivalent to a transition function you would think about in RL in general.
Exactly.
So we can solve this in closed form, basically, like, you know, we can, because we have this world model of Newtonian physics, we can say at every step exactly how this drone should fly so that it lands on the appropriate thing.
Exactly.
Under a set of constraints like max thrust.
So, yeah, let's have an example then of how you could make this non-differentiable.
Like, well, what's a scenario?
I guess even like this drone scenario where it now becomes non-differentiable.
And stop me from getting there.
Now, from the position of your drone, you don't know what actions I'm going to take.
All of this stuff ultimately comes down to ways to estimate, to, to model this non-differentiable stochastic process.
Exactly.
In a sense, the value gives you some expectation of future rewards, like the sum of future rewards you're getting.
And so if you're in a bad space, you would set the value to zero or negative infinity or something.
Showing 21–40 of 237 · page 2 of 12
← Previous
Next →