Kevin Wang
speaker
180 appearances
1 recordings
1 series
first heard Jan 2026
last heard 2 Jan
Kevin Wang’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 1 in all, peaking in Jan 2026 with 1.
Appearances
And then all of a sudden, like one day, like I ran this experiment, and there was like this one environment in which there was like like going from like like doubling the depth didn't really do anything, but like doubling the depth again with these different components, suddenly like skyrocketed performance in this one.
Environment.
Yeah, I think the best way that I would explain it is that we know that standard RL is not super scalable.
And so like why can this a different approach or different objective RL be scalable?
I think it's because we're fundamentally shifting the
Burden of learning from something like Q like Q learning or like regressing to like TD errors, which we know is quite spurious and noisy and biased, to fundamentally like a classification problem.
We're trying to classify whether future state is along the same trajectory or along a different trajectory.
And we do this with representation learning, right?
And we know that classification.
cross-entry loss and representation learning is scalable in the deep learning literature, right?
If we think about language um a and like some of the objectives there.
So in some sense, we're kind of
Blurring the the lines we're doing reinforcement learning.
It's still an actor-critic reinforcement learning algorithm.
It's like a goal condition reinforcement algorithm.
But the objective, the burden of like learning of the of solving that RL task shifts to something that's more similar to some objectives that you might see in language, envision that we know have scaled so much.
And so I think, yeah, I think that's like one of the fundamental insights that we've seen is that.
Um, it seems like by approaching RL and in this different approach, we were able to like s get so much more out of we were able to scale our networks like significantly beyond what is like standard used in RL.
Yeah.
So actually if you look at a lot of our tasks, they're particularly sort of like robotics tasks.
Showing 61–80 of 180 · page 4 of 9
← Previous
Next →