Sebastian Raschka

speaker
1,024 appearances 1 recordings 1 series first heard Feb 2026 last heard 1 Feb

Sebastian Raschka’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · Feb OctJan 26AprJulnow

Recordings per month over the last 12 months — 1 in all, peaking in Feb 2026 with 1.

Appearances

newest first · ▶ plays the moment
So you can convert one from one.
You can go from one into the other by just adding these changes, basically.
Yep.
So for example, you mentioned my book earlier, that's a GPT-2 model in the book because it's simple and it's very small.
So 124, 120 million parameters approximately.
But in the bonus materials, I do have almost three from scratch, Gemma 3 from scratch and other types of from scratch models.
And I always started with my GPT-2 model and just, you know, tweaked or added different components and you get from one to the other.
It's like, it's kind of like a lineage in a sense.
So there are the different stages where you develop the network or train the network.
You have the pre-training.
Now, back then, it was just pre-training with GPT-2.
Now you have pre-training, mid-training and post-training.
So I think right now we are in the post-training focus stage.
I mean, pre-training still gives you advantages if you scale it up to better, higher quality data.
But then we have capability unlocks that were not there with GPT-2.
For example, ChatGPT, it is basically a GPT-3 model.
And GPT-3 is the same as GPT-2 in terms of architecture.
What was new was adding the supervised fine-tuning and the reinforcement learning with human feedback.
So it's more on the algorithmic side rather than the architecture.
Like you said, they had, for example, in the mixture of experts, this NVFP4 optimization, for example, where you get more throughput.
Showing 201–220 of 1,024 · page 11 of 52 ← Previous Next →