Sebastian Raschka

speaker
1,024 appearances 1 recordings 1 series first heard Feb 2026 last heard 1 Feb

Sebastian Raschka’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · Feb OctJan 26AprJulnow

Recordings per month over the last 12 months — 1 in all, peaking in Feb 2026 with 1.

Appearances

newest first · ▶ plays the moment
AI made up data, it's also taking something from an article, Wikipedia article, and then rephrasing it as a Q&A question or summarizing it, rewording it and making better data that way.
Because I think of it also like with humans, if someone, let's say, reads the book compared to a messy, I don't know, no offense, but like Reddit post or something like that.
I think that's the idea.
I think it's like if someone took that and rephrases that in a, let's say, more concise and structured way,
I think it's higher quality data that gets the LLM maybe the same.
You get the same LLM out of it at the end, but it gets there faster.
It trains faster because let's say if the grammar and the punctuation is correct, it already learns the correct way versus getting information from a messy way and then learning later how to correct that and stuff like that.
So I think
That is how pre-training evolved and why scaling still works is that it's not about just the amount of data, it's also the tricks to make that data better for you, in a sense.
And mid-training is, I mean, it used to be called pre-training.
I think it's called mid-training because it was awkward to have pre-training and post-training, but nothing in the middle, right?
It sounds a bit weird to have pre-training and post-training, but what's the actual training?
So the mid-training is...
usually similar to pre-training but you know it's a bit more i would say specialized in pre-training it's the same algorithm but what you do is you focus for example on long contact like it's one example you have long context documents the reason you don't do that during just pure pre-training is because you don't have that many long context documents you have a specific phase and one problem of lms is also still it's a neural network it has the problem of catastrophic forgetting so you teach it something it forgets other things and you want to
It's not 100% forgetting, but it's like no free lunch.
It's also the same with humans.
If you ask me some math I learned 10 years ago, I don't know.
I would have to look at it again.
I don't want to anthropomorphize LLMs, but it's, I think, the same kind of in that sense how humans learn.
I mean,
Showing 281–300 of 1,024 · page 15 of 52 ← Previous Next →