Kevin Wang
speaker
180 appearances
1 recordings
1 series
first heard Jan 2026
last heard 2 Jan
Kevin Wang’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 1 in all, peaking in Jan 2026 with 1.
Appearances
However, this is not so like within our paper, like
For most environments, um, we are able to like saturate, like get to like almost perfect performance within just, you know, we don't even need to get to like a thousand layers, like maybe just sixty-four layers, for example, is sufficient.
Um, and in in this regime, like like the legacy of the network is not necessarily actually even the uh not necessarily like a a significant bottleneck.
Like you can imagine there's a lot of tasks in which especially in RL that like collecting data might be the bottleneck, right?
And making
passes through our network, it may not be the bottleneck.
And so in our environment, we sp in our research, we specifically used the JAX G C R L environment, which is a JAX based GPU accelerator environment.
So we can collect like thousands of like environment trajectories like in parallel at the same time.
So that we're able to like make uh like oh this is built in.
Right, this is built in so that we can collect you know like like thousand a thousand trajectories at the same time al along all these environments and so um mak makes that make sure that like we have enough data to like saturate the learning from things.
Wow.
I guess even to build on that, like I I like drawing analogies to like successes in other areas of deep learning.
Like for example, in large language models, the reason why we're able to scale to such large networks is that we found a paradigm in which we can leverage the entire internet scale of data to learn, right?
And so data in RL traditionally has been hard to come by.
Um, but now with these like GPU accelerated environments, we can collect hundreds of millions of times of the data within just a few hours.
And so I think that this serves as like a really good test bed for us to be able to also find ways to scale up um like network capacity and get similar kind of gains.
Oh, I'm not saying that or change it.
I wanna leverage insights from that.
To apply to morale as
You think you should go the other way?
Showing 101–120 of 180 · page 6 of 9
← Previous
Next →