Yi Tay
speaker
925 appearances
1 recordings
1 series
first heard Jan 2026
last heard 23 Jan
Yi Tay’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 1 in all, peaking in Jan 2026 with 1.
Appearances
So I think on policyness is basically this idea of like model training on its own outputs and letting the model like generate its own trajectories and then let letting some reward verify it and then the model train its own outputs.
I think this is more generalizable in general.
I think there's still a lot of like science out still to be.
be done about the gap between SFT and and RL itself.
But I think basically on policy and off policy, right?
And I think bring this analogy back to real like life.
I mean we this on policy nurses is more like the humans, we are more on policy because we go around the world, we make mistakes and then we ah, okay, this is what
Like imitation learning is possibly somebody else.
No first principle.
It just tells you what to do and then you just copy.
So I think, yeah, this philosophy, bringing back this philosophy to life is quite like powerful.
Like when I like now I have a kid and everything, like one my kid to try stuff and then you tell them like, okay, this is like where this went wrong, where this went right and stuff, rather than okay, you just copy everything somebody else does.
Technically SFT is still imitation.
But I think for humans it's not a little bit of this, right?
Because if you basically like like sports, right?
When you play sports, you start off by imitating, like hardcore imitating.
But then you cannot imitate forever because you need to like imitation.
I don't know whether this is a good analogy, but watching a lot of tutorials and stuff is more like imitating.
You learn, try to learn sudden movements and stuff like that.
But then like on pause seen us, it's like going to the game itself and trying to
Showing 61–80 of 925 · page 4 of 47
← Previous
Next →