Sebastian Raschka

speaker
1,024 appearances 1 recordings 1 series first heard Feb 2026 last heard 1 Feb

Sebastian Raschka’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · Feb OctJan 26AprJulnow

Recordings per month over the last 12 months — 1 in all, peaking in Feb 2026 with 1.

Appearances

newest first · ▶ plays the moment
And if listeners are familiar with multi-layer perceptrons, you can think of a mini multi-layer perceptron, a fully connected neural network layer inside the transformer.
And it's very expensive because it's fully connected.
If you have 1,000 inputs, 1,000 outputs, that's like 1 million connections.
And it's a very expensive part in this transformer.
And the idea is
to kind of expand that into multiple feedforward networks.
So instead of having one, let's say you have 256, but it would make it way more expensive because now you have 256, but you don't use all of them at the same time.
So you now have a router that says, OK, based on this input token, it would be useful to use this fully connected network.
And in that context, it's called an expert.
So a mixture of experts means you have multiple experts.
And depending on what your input is, let's say it's more math heavy, it would use different experts compared to, let's say, translating input text from English to Spanish.
It would maybe consult different experts.
It's not quite clear.
I mean, it's clear cut to say, OK, this is only an expert for math and for Spanish is a bit more fuzzy.
But the idea is essentially that you pack more knowledge into the network, but not all the knowledge is used all the time.
That would be very wasteful.
So you're kind of like during the token generation, you're more selective.
There's a router that selects which tokens should go to which expert.
It's more complexity.
It's harder to train.
Showing 161–180 of 1,024 · page 9 of 52 ← Previous Next →