Anjney Midha
speaker
1,670 appearances
8 recordings
1 series
first heard Dec 2024
last heard Sep 2025
Anjney Midha’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsNo recordings in the last 12 months.Older appearances are listed below; set an alert to hear about the next one.
Appearances
Yeah, every couple of months OpeningI puts out a release and everyone goes, Oh, that's great.
But this RL thing is going to plateau.
We're gonna saturate the evals, the models won't generalize, or there's gonna be mode collapse because of too much synthetic data for whatever not everybody's got a laundry list of reasons to believe that the gains in performance from RL are going to tap out.
And somehow they just don't.
You guys just keep coming out and putting out continuous improvements.
Why is RL working so well?
And what if anything has surprised you?
About how well it works.
One of the hardest things about RL for folks who are not practitioners of RL is the idea of crafting the right reward model.
And so especially if you're a a business or an enterprise who wants to harness all this amazing progress you guys are putting out.
But doesn't even know where to start.
What do the next few years look like for a company like that?
What is the right mindset?
For somebody who's trying to make sense of RL.
to craft the right reward model.
Is there anything you've learned about the best practices or a of an approach of thinking to of using this latest sort of family of reasoning techniques?
What is the right way I should think about even approaching reward modeling?
As a biologist or a physicist.
I I that I have a question about that, which is what makes a great researcher?
Right.
Showing 21–40 of 1,670 · page 2 of 84
← Previous
Next →