Eiso Kant

speaker
1,158 appearances 1 recordings 1 series first heard Jul 2026 last heard 23 Jul

Eiso Kant’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · Jul OctJan 26AprJulnow

Recordings per month over the last 12 months — 1 in all, peaking in Jul 2026 with 1.

Appearances

newest first · ▶ plays the moment
I have a, I would say, a not commonly held opinion that reinforcement learning will move earlier and earlier into pre-training.
Not even mid-training.
Like mid-training today, right, is like if you look at, so we've been working on this for years already.
And I think the best, I think the first time we saw it out in public was the DeepSeek Zero paper.
That was a year and a half ago, I think, if I recall correctly.
Where, you know, you can very early on in a model as it starts capable of being able to use language, et cetera, induce reasoning.
And so the question that I kind of have is like, we have the data set that's the web.
And the web, I think we could arguably say probably has the totality of humanity's knowledge somewhere encoded in different places.
It's a huge variance degree of quality from garbage data.
And like, once you look at pre-training data, you really get humbled of like what the web is to like, you know, the most greatest scientific papers and best blog posts and like, you know, best transcripts and whatnot.
And so now what we are trying to figure out, and I've been doing a lot of work on, and it's a place where maybe not as open as we're on other things, but we will become more over time.
We've been spending a couple of years really doing research on how can we turn the web
into not just next token prediction, but into a way to teach the model to think earlier in its training.
And I think there's a huge amount of gold to be found there.
I think we are right now in, you know, we've got some drugs in the industry.
One of the drugs is distillation.
Another drug is, you know, more environments, like, and they're great and they make us feel good and they make the models better and like we're all addicted to them and we'll use them.
right, in various different ways.
But ultimately, I think we are still barely squeezing out of the web what we should be getting out of the web.
I think just next token prediction during pre-training is not enough.
Showing 441–460 of 1,158 · page 23 of 58 ← Previous Next →