Ashvin Nair

speaker
533 appearances 1 recordings 1 series first heard Dec 2025 last heard 30 Dec

Ashvin Nair’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · Dec OctJan 26AprJulnow

Recordings per month over the last 12 months — 1 in all, peaking in Dec 2025 with 1.

Appearances

newest first · ▶ plays the moment
Uh I think human feedback is kind of like a a bit of like a side branch because you can't really s pour that much compute into it, right?
It's like you t you take the model and you like elicit it to be a little bit better uh in terms of personality.
But like the people there are really convinced that like at some point, you know, it's not about copying the internet.
Like you can go um yeah, do RL and uh you like you know that's like the path to like getting
much better intelligence.
So I think it it there was kind of like a long line of kind of like returning to RL in for in like different ways.
And then it's just that around like yeah 2023 is when it started like really clicking.
And it was kind of interesting because even you know it's it's not like those initial models performed like um way better than the existing models because they're like smaller scale.
But people were very good at being like, oh like this is kind of interesting.
Like, you know, the the reasoning trace that you see here is kind of not something that you've really seen be so accurate um in other models like this, like this one.
Kind of similar to how
I think a lot of people didn't really think of GPT or GPT two as something that was like super compelling, probably.
I know that I personally didn't like uh GPT two that much of GPT two.
I was like, okay, whatever.
And then I think and then GPT three happened, I'm like, oh whoa, like I feel a lot of FOMO sitting in my PhD.
It's kind of that where I think it takes a bit of like first principles conviction to like um yeah, decide that like, oh, this this thing, like there's something here and we should really
go scale it up.
And open eye is really good about once you decide that something is good, then you just like scale it up all the way.
J just this like, you know, like uh
like running RL on even like a pretty small model, producing like very interesting reasoning traces and like getting like uh surprisingly good scores on math.
Showing 301–320 of 533 · page 16 of 27 ← Previous Next →