Kyle Corbitt
speaker
658 appearances
1 recordings
1 series
first heard Oct 2025
last heard 16 Oct
Kyle Corbitt’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 1 in all, peaking in Oct 2025 with 1.
Appearances
You're especially as as you're getting into these agentic tasks um and training them to do that, like it seems very clear.
Well, obviously the big labs are paying like ridiculous amounts of money for these environments and everything, but also like they're actually getting really good results.
The the the models coming out, you know, we're seeing it especially on the coding model side, but like in other in other contexts as well, we're seeing the sort of especially agentic use is working way better because of this.
So I think like
Even
Late 2024, it was pretty clear that like RL was going to work in that context.
And then the question in our mind was like, can we apply this in a different segment of the business, which is kind of like task-specific customization?
And so the question is like, does that work well?
How much effort does that take?
Is it going to be something that ends up being unnecessary because oh, the the big labs can just like train on every single task and the base models are going to be just good at everything?
And so there's you know no benefit to it.
So those were kind of the open questions in in our
our mind, but it seemed like there was like at least a good enough bet that, you know, we wanted to try it out.
Told our team, and this was we decided to go in all in on RL in January of 2025.
And we've been doing some experience before that.
We released before that kind of like an RL model that had, you know, would generate like hacker news titles from articles, which is a fun project.
So we'd done a little bit before that, but that was kind of like we're like, hey, we're gonna bet the company on.
Uh not in a literal sense, like we we could have done something else later, but like this is like the thing that we're gonna spend all of our time working on for at least a few months.
And like what I told our
Team at that time in January 25 was like there's probably like a 25% chance that this is the right direction in the sense that like a year, two years from now, all the companies, you know, everyone doing inference should be doing RL and task-specific training so that like their model's just way, way better at their task is a relatively low chance.
Showing 161–180 of 658 · page 9 of 33
← Previous
Next →