Steve Gibson

speaker
17,616 appearances 14 recordings 1 series first heard May 2026 last heard 2 Sep

Steve Gibson’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
5 · Jul OctJan 26AprJulnow

Recordings per month over the last 12 months — 14 in all, peaking in Jul 2026 with 5.

Appearances

newest first · ▶ plays the moment
That has its roots in AI circa 2020.
Way back then, if you were to carefully phrase a question to GPT 3, such as the capital of France is
it would have been able to complete the sentence by emitting the next expected word.
Paris.
Because the Maddal th that model, the GPT three model's vast statistical data set.
um made uh made Paris the next most likely word.
The capital of France is Paris.
But if you used question style phrasing back then, what is the capital of France?
That would not have been met with the same success.
So what happened?
Obviously, we have that now.
How did we turn these models from autocomplete engines, that is, that's where that's all they could do, into conversationalists?
So it took a few years of experimentation, but AI researchers first used something that's now known as instruction tuning.
They took the knowledge-trained model and fine-tuned it on a, and this is again another surprise, a surprisingly small set of human-written examples of what good responses would look like.
After that, then something known as RLHF, reinforcement learning from human feedback, was used.
So this applied preference rankings, where again, humans compared and ranked multiple outputs from best to worst.
And that ranking was used to train a reward which pushed the model toward the better behavior.
Now here's the astonishing part to me.
Researchers then found that they could employ a comparatively small model of these query and ranked response samples.
Which then created reward feedback.
Showing 1861–1880 of 17,616 · page 94 of 881 ← Previous Next →