Eiso Kant

speaker
1,158 appearances 1 recordings 1 series first heard Jul 2026 last heard 23 Jul

Eiso Kant’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · Jul OctJan 26AprJulnow

Recordings per month over the last 12 months — 1 in all, peaking in Jul 2026 with 1.

Appearances

newest first · ▶ plays the moment
And so for us, distilling down to a smaller model doesn't actually serve the purpose.
These models are kind of, it's not the right term, but to us, they're dual purpose models.
They are a progress for us to weigh to see, did we improve in the model factory and something to put out into the world.
And so that's why we don't do it.
We've done distillation experiments and there's really cool things you can do.
And I think if you have lots of user data, then you can go even further in that.
But I think there's something to be said in having a quick cadence of models trained end-to-end from scratch so that you as a research organization can learn the lessons and not wait.
That was actually one of the big lessons we learned over the years when we used to have a much
longer cadence between model trainings like six months and we would train just like a big big model wait six months train another bigger model uh you would be compounding so many changes of improvements that at the by the time you're training your next model it's a bit of a soup and you don't really know what ingredients led to the outcomes so when you are training far more frequently models uh
and this holds true for post-training and from pre-training from scratch, you are much more able to get an understanding of what led to the improvements.
And I think that's important.
Like ultimately, we are all still, there is no true science yet of, you know, deep learning for large language models.
But we are all, I think, you know, trying to gain insights from our experiments because it's those insights that lead to scaling laws, that lead to the kind of improvements that allow us to be, again, more compute efficient and get more capabilities.
I mean, it's wild, right?
I mean, it's like the golden age.
It's the fact that you can just take an idea and build something by waiting overnight for an agent to do the work.
I don't know, to me... How do you measure?
Because literally in a theory... Look, I think it's a good question.
It's one I haven't thought about in a long time.
No, it's a fair point.
Showing 1061–1080 of 1,158 · page 54 of 58 ← Previous Next →