Steve Gibson
speaker
17,616 appearances
14 recordings
1 series
first heard May 2026
last heard 2 Sep
Steve Gibson’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 14 in all, peaking in Jul 2026 with 5.
Appearances
That is, it wasn't giving it new knowledge, it was it was reformatting it and giving it behavior for the first time.
So
To state that a bit differently for clarity.
The breakthrough about four years ago was not the use of a larger model than we had at that time.
That is one of the things that's been happening since then.
But the breakthrough was the unexpected discovery that a comparatively infinitesimal dose of human provided, here's what
a good answer looks like and here's which of these answers
Is better than the other.
Training.
Would ре shape.
a massive already trained model's entire behavior, turning it from something that completes patterns into something that acts like it's trying to be helpful.
So
Once the massive model learned what helpful looked like, and that its trainers wanted it to look like that, all of its stored knowledge was immediately available in that new helpful format.
And we got helpful chat EI AI.
That's how it happened.
Um, and as we know, those first steps made by Chat GPT were more than a little shaky.
You know, uh, you know, researchers realized that their new, their newly birthed chatbot would need some additional post-training alignment, as it's now being called.
And this began as that RLHF, the reinforcement learning from human feedback that we talked about.
But then a year later in 2023, a handful of AI researchers at Stanford University published a paper titled Direct Preference Optimization.
The paper's title is Direct Preference Optimization.
Showing 1901–1920 of 17,616 · page 96 of 881
← Previous
Next →