Sebastian Raschka

speaker
1,024 appearances 1 recordings 1 series first heard Feb 2026 last heard 1 Feb

Sebastian Raschka’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · Feb OctJan 26AprJulnow

Recordings per month over the last 12 months — 1 in all, peaking in Feb 2026 with 1.

Appearances

newest first · ▶ plays the moment
There's a lot of, you know, that can go wrong, like collapse and everything.
So I think that's why Olmo 3 still uses dense.
I mean, you have, I think, Olmo models with a mixture of experts, but dense models where dense means also it's jargon.
There's a distinction between dense and sparse models.
So a mixture of experts is considered sparse because we have a lot of experts, but only a few of them are active.
So that's called sparse.
And then dense would be the opposite where you only have like one fully connected module and it's always, you know, utilized.
picture like the mixture of experts.
The attention mechanism in GPT-OSS, that would be the group query attention mechanism.
So it's a slight tweak from multi-head attention to group query attention.
So there we have two.
I think they replaced layer norm by RMS norm, but it's just like a different normalization layer.
Not a big change, it's just like a tweak.
The nonlinear activation function
People familiar with Deep New Networks, I mean, it's the same as changing Sigmoid with ReLU.
It's not changing the network fundamentally.
It's just like a tweak, a little tweak.
And that's about it, I would say.
It's not really fundamentally that different.
It's still the same architecture.
Showing 181–200 of 1,024 · page 10 of 52 ← Previous Next →