Yi Tay

speaker
925 appearances 1 recordings 1 series first heard Jan 2026 last heard 23 Jan

Yi Tay’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · Jan OctJan 26AprJulnow

Recordings per month over the last 12 months — 1 in all, peaking in Jan 2026 with 1.

Appearances

newest first · ▶ plays the moment
especially this year.
Yeah, I think reasoning these days reasoning and RL is like probably quite it's m RL reasoning comes from I spent a lot of my past life, I call it the past art, working on like architectures and pre training.
But I think now I more I would like transition more into RL recently.
I'm not like old school RL but again RL and the old school RL and and to be honest I had almost no RL background coming back.
But I think like like RL is
the main means of modeling these days.
And yeah, so I think it was pretty easy to jump back in.
And I think a lot of fundamental skills in research is general purpose and universal and and it's it's quite easy to innovate even in a tool set that you're not super used to.
And yeah, so I think RL is basically the main modeling tool set that we play around with.
Yeah.
These days.
But I know I understand it's basically the shift is objective and they have some like overlap, right?
Yeah.
I think it's just mainly like the on-policy and off-policiness of designing these things that change how like also the learning algorithm itself, right?
Yeah, so I think like the biggest analogy of the on policy and odd policy is basically odd policy is basically like when you SFT something, it's odd policy.
Basically you take some the other model larger model stuff and then it's basically like there's off polic somebody else's generated outputs, trajectories and whatever.
I think odd policy is mainly like the core idea of like modern LMRL where you like generate and then you reward the model based on its own generations and then the model trains on its own generation.
Yeah.
So it's more it's a bit like self-disolation to some extent.
You the model generates its own output and then you reward it and then trains on its own output.
Showing 41–60 of 925 · page 3 of 47 ← Previous Next →