Sebastian Raschka

speaker
1,024 appearances 1 recordings 1 series first heard Feb 2026 last heard 1 Feb

Sebastian Raschka’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · Feb OctJan 26AprJulnow

Recordings per month over the last 12 months — 1 in all, peaking in Feb 2026 with 1.

Appearances

newest first · ▶ plays the moment
The quantity is not always better because, yeah, it's like being selective.
And the mid-training is being selective in terms of quality content at the end.
So the last thing the LM has seen is the quality stuff.
And then post-training is all the fine-tuning, supervised fine-tuning, DPO stuff.
um reinforcement learning with verifiable rewards with human feedback and so forth so the refinement stages and it's also interesting it's like the cost thing right i mean it's like pre-training you spend a lot of money on that right now rl a bit less rl you don't really i would say teach it knowledge it's more like unlocking the knowledge it's more like a skill learning like how to solve problems with the knowledge that it has from pre-training
There are actually three papers this year, or last year, 2025, on RL for pre-training.
But I mean, I don't think anyone does that in production.
Toy examples for now.
Toy examples, right.
But to generalize, RL post-training is more like the skill unlock, where pre-training is like soaking up the knowledge, essentially, and
One interesting question is, if I recall correctly, OMO3 was trained with less data than specifically some other open-weight models, maybe even OMO2, but you still got better performance.
And that might be one of the examples how the data helped.
It's mostly down to data quality.
At the same time, I think it's also one of the closest guarded secrets what your training data is for legal reasons.
And so there's also, I think, a lot of work that goes into hiding what your trading data was, essentially, like trying the model to not give away the sources because of legal reasons.
just purchase the license.
Let's say they buy a book online, let's say an Amazon Kindle book or let's say a money book or something, and then use that in the training data.
And that is like the gray zone because you paid for the content and you might want to train it.
But then there are also restrictions where even that shouldn't be allowed.
And so that is like where it gets a bit fuzzy and...
Showing 301–320 of 1,024 · page 16 of 52 ← Previous Next →