Sebastian Raschka
speaker
1,024 appearances
1 recordings
1 series
first heard Feb 2026
last heard 1 Feb
Sebastian Raschka’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 1 in all, peaking in Feb 2026 with 1.
Appearances
And then it's also like the math approach.
How long is my model going to be on the market if I replace it in half a year?
Maybe it's not worth spending 5 million, 10 million, 100 million dollars on the training it longer.
Maybe it's just I will just do more inference scaling and get the performance from there.
It may cost me 2 million in terms of user queries.
It becomes a question of how many users you have and then doing the math.
And I think that's also where it's interesting.
where JGPD is in a position, I think they have a lot of users where they need to go a bit cheaper, where they have that JGPD 5 model that is a bit smaller.
Other companies that have, let's say if your customers have other trade-offs, for example, there was also the Math Olympiad or some of these math problems where
GGBT or OpenAI, they had a proprietary model, and I'm pretty sure it's just like a model that has been maybe fine-tuned a little bit more, but most of it was doing inference scaling to achieve this peak performance in certain tasks where you don't need that all the time.
But yeah, long story short, I do think all of these pre-training, mid-training, post-training, inference scaling, they are all still things you want to do.
It's just finding, at the moment, in this year, it's finding the right ratio that gives you the best bang for the buck, basically.
So pre-training is the classic training, one next token prediction at a time.
You have a big corpus of data.
And Nathan also has very interesting insights there because of OMO3.
It's a big portion of the paper focuses on the right data mix.
So pre-training is essentially just, you know, across entropy loss training on next token prediction on a vast corpus of internet data, books, papers, and so forth.
It has changed a little bit over the years in the sense people used to throw in everything they can.
Now, it's not just raw data, it's also synthetic data where people rephrase certain things.
So synthetic data doesn't necessarily mean purely
Showing 261–280 of 1,024 · page 14 of 52
← Previous
Next →