Sebastian Raschka

speaker
1,024 appearances 1 recordings 1 series first heard Feb 2026 last heard 1 Feb

Sebastian Raschka’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · Feb OctJan 26AprJulnow

Recordings per month over the last 12 months — 1 in all, peaking in Feb 2026 with 1.

Appearances

newest first · ▶ plays the moment
And so you can see that basically the RL, it's not teaching the model any new knowledge about math.
You can't do that in 50 steps.
So the knowledge is already there in the pre-training.
You're just unlocking it.
But if it weren't true, I would say distillation wouldn't work, right?
I mean, distillation can work to some extent.
But the thing is, that is, I think, the biggest problem in LLM research, this contamination, because we don't know what's in the data.
Unless you have a new data set, it's really impossible.
And the same, you mentioned math, the math data set, which is you have a question and an answer and an explanation is given.
But then also even something simpler like MMLU, which is a multiple choice benchmark, which
If you just change the format slightly, like, I don't know, you use a dot instead of a parenthesis or something like that, the model accuracy will vastly differ.
It's not even malicious by the developers of the LLM, like, hey, we want to cheat at that benchmark.
It's just, it has seen something at some point.
And I think the only fair way to evaluate an LLM is to have a new benchmark that is after the cutoff date when the LLM was deployed.
So RLVR is more, let's say, unlimited how much you can train and get still benefit where RLHF, because it's a preference tuning, you reach a certain point where it doesn't really make sense to spend more RL budget on that.
So just a step back with preference tuning.
So there are multiple people that can give multiple, let's say, explanations for the same thing and they can both be correct.
But at some point you learn a certain style and it doesn't make sense to, you know, iterate on it.
My favorite example is like if relatives ask me what laptop they should buy.
i give them an explanation or ask them like yeah what is your um use case like they for example prioritize battery life and storage other people like us for example we would prioritize ram and compute and so but both both answers are correct but different people require different answers and with preference tuning well you're trying to average somehow like you are asking the data labelers to give you the right or not the right the preferred answer
Showing 461–480 of 1,024 · page 24 of 52 ← Previous Next →