Sebastian Raschka

speaker
1,024 appearances 1 recordings 1 series first heard Feb 2026 last heard 1 Feb

Sebastian Raschka’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · Feb OctJan 26AprJulnow

Recordings per month over the last 12 months — 1 in all, peaking in Feb 2026 with 1.

Appearances

newest first · ▶ plays the moment
But I do think this is for the speed.
This is true, but it doesn't give the model new capabilities in a sense.
It's just how much can we make the computation coarser without suffering in terms of model performance degradation.
But I do think, I mean, there are alternatives popping up to the transformer.
There's text diffusion models, completely different paradigm.
And there's also, I mean, the text diffusion models might use transformer architectures, but it's not an auto-regressive transformer.
And also Mamba models, it's a state-space model, but they do have trade-offs.
And what's right is there's nothing that has replaced the auto-regressive transformer as state-of-the-art model.
So like for state-of-the-art, you would still do that,
go with that thing but there are no alternatives for the cheaper and like alternatives that are kind of um making compromises but it's not just one architecture anymore there are little ones coming up but if we talk about the state of the art it's pretty much still the the transformer architecture autoregressive derived from gpt2 essentially i guess the big question here is we talked quite a bit here on the architecture behind the pre-training
Yeah, so that's a big can of worms here.
But so basically two of the knobs are the training and the inference scaling where you can get gains.
And so in a world where we had, let's say, infinite compute resources, you want to do all of them.
So you have training, you have inference scaling, and training is like a hierarchy.
It's pre-training, mid-training, post-training, changing the model size, more training data, making training a bigger model.
gives you more knowledge in the model than the model, let's say, has a better, it's like a better base model back in the day, or still we call it foundation model.
And it unlocks, but you don't, let's say, have the model be able to solve your most complex tasks during pre-training or after pre-training.
You still have these other unlock phases where you have mid-training or long context, for example, post-training with LRVR that unlocks capabilities that the model has in terms of just knowledge in the pre-training.
And I think, sure, if you do more pre-training, you get a better base model that you can unlock later.
But like Nathan said, it just becomes too expensive.
Showing 221–240 of 1,024 · page 12 of 52 ← Previous Next →