Eiso Kant
speaker
1,158 appearances
1 recordings
1 series
first heard Jul 2026
last heard 23 Jul
Eiso Kant’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 1 in all, peaking in Jul 2026 with 1.
Appearances
But it made no sense to us intuitively.
Like, why do you have to have this so high?
Like, if you're just trying to avoid division by zero, why can't the value be extremely small?
And that was kind of like one of those moments where you realize like, okay, finding things out from scratch yourself builds a better intuition.
Because the one thing you learn very quickly with model building is that your intuitions that you start with are going to get beaten up.
So harsh, right?
Like the, it's such an experimental science, uh, that the things that seem obvious, you very quickly get to learn, like you were wrong and hopefully you figure out why.
And sometimes you don't even.
So I would say that our view from very early on in the company was that model building is ultimately 90% engineering.
And I think we all know it in the industry, because if you look at where every researcher is spending their time, they're spending their time writing code, right?
Looking at data and writing code.
And so we kind of said, okay, the state at the moment, like three years ago, was...
bash scripts and slurm and spaghetti code bases for training and like data pipelines that were kind of patched together and we kind of looked at this and said well ultimately model building is a process you're going from raw data right like pre-training raw material the web etc uh you're doing a whole bunch of filtering cleaning up transformations analyzing these days that's you know far more complex than it was three years ago uh
Then you're training a model, which is effectively a large distributed systems problem, right, across hardware that has still, it's become a lot more reliable, was extremely flaky back then.
And now with every new generation, we get our new sets of challenges.
And then you go into the next stages, right?
There was no mid-training back then, but like, you know, you're post-training and then you're reinforcement learning.
And so we kind of looked at this and we said, well, this looks like an industrialized process.
This looks like an end-to-end process that every single part of it kind of has its machinery, right?
If it's your big data pipelines, if it's your crawling ingestion of the web, if it's your large-scale distributed training, and then you've got your reliability.
Showing 161–180 of 1,158 · page 9 of 58
← Previous
Next →