Sebastian Raschka

speaker
1,024 appearances 1 recordings 1 series first heard Feb 2026 last heard 1 Feb

Sebastian Raschka’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · Feb OctJan 26AprJulnow

Recordings per month over the last 12 months — 1 in all, peaking in Feb 2026 with 1.

Appearances

newest first · ▶ plays the moment
I would say like an RNN, you try to compress everything into one state.
You're a bit more selective there.
But then I think it's like this Goldilocks zone again.
With Nemotron 3, they found like a good ratio of how many attention layers do you need for the global information where everything is accessible compared to having these compressed states.
And I think that's how I think we will scale more by finding better, let's say, ratios in Goldilocks zone, like between...
like computing, making it cheap enough to run, but then also making it powerful enough to be useful.
And one more plug here, the recursive language model paper, that is one of the papers that tries to kind of address the long context thing.
So what they found is essentially instead of
stuffing everything into this long context um if you break it up into these smaller multiple smaller tasks so you save memory by having multiple smaller calls you can get actually better accuracy than having the llm try everything all at once i mean it's a new paradigm we will see you know there might be other flavors of that so i think with that we will still make improvement on long context but then also like nathan said i think the problem is
for pre-training itself we don't have as many long context documents as other documents so it's harder to study basically how LMs behave and stuff like that on that level like
One interesting also recent example would be DeepSeq version 3.2, where they had like the sparse attention mechanism where they have essentially like a very efficient, small, lightweight indexer.
And instead of attending to all the tokens, it selects, okay, what tokens do I actually need?
I mean, it almost comes back to the original idea of attention where you are selective, but attention is always on.
You have maybe zero weight on some of them, but you use them all.
But they are even more like, okay, let's just mask that out or like not even do that.
And even with sliding window attention, Olmo, that is also kind of like that idea.
You have that rolling window where you keep it fixed because you don't need everything all the time.
Occasionally, some layers you might, but it's wasteful.
But right now, I think, yeah, if you use everything, you're on the safe side.
It gives you the best bang for the buck because you never miss information.
Showing 741–760 of 1,024 · page 38 of 52 ← Previous Next →