Sebastian Raschka
speaker
1,024 appearances
1 recordings
1 series
first heard Feb 2026
last heard 1 Feb
Sebastian Raschka’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 1 in all, peaking in Feb 2026 with 1.
Appearances
And then you train on that.
But at some point, yeah, you learn that average preferred answer.
And there's no, I think, reason to keep training longer on it because, you know, it's just a style where with RLVR, you literally give the model, well, you let the model solve more and more complex, difficult problems.
And so I think that it makes more sense to allocate more budget long-term to LRVR.
And also that right now we are in LRVR 1.0 land where it's still like that simple model
thing where we have a question and answer but we don't do anything with the one stuff in between so there was a i mean multiple research papers also by google for example on process reward models that also give scores for the explanation how correct is the explanation and i think that will be the next thing let's say our lvr 2.0 for this year
focusing in between question and answer, like how to leverage that information, the explanation to improve the explanation and help it to get better accuracy.
But then, so that's one angle.
And there was a DeepSeq math version two paper where they also had interesting inference scaling there, where first they had developed models that grade themselves a separate model.
And I think that that will be one aspect.
And the other, like Nathan mentioned, it will be for LRVR branching into other domains.
So I would personally start, like you said, implementing a simple model from scratch that you can run on your computer.
The goal is not, if you build a model from scratch, to have something you use every day for your personal projects.
It's not going to be your personal assistant replacing an existing Openweight model or Chachapiti.
It's to see what exactly goes into the LLM, what exactly comes out of the LLM, how the pre-training works in that sense, on your own computer preferably.
And then you learn about the pre-training, the supervised fine-tuning, the attention mechanism.
You get a solid understanding of how things work.
But at some point you will reach a limit because small models can only do so much.
And the problem with learning about LLMs at scale is, I would say it's exponentially more complex to make a larger model because...
It's not that the model just becomes larger.
Showing 481–500 of 1,024 · page 25 of 52
← Previous
Next →