Steve Gibson

speaker
17,616 appearances 14 recordings 1 series first heard May 2026 last heard 2 Sep

Steve Gibson’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
5 · Jul OctJan 26AprJulnow

Recordings per month over the last 12 months — 14 in all, peaking in Jul 2026 with 5.

Appearances

newest first · ▶ plays the moment
your language model is secretly a reward model.
Um, and since all of the descendants of this, it's now known as DPO system, direct preference optimization was sort of the granddaddy.
Now we have refin further refinements of that, which have occurred in the last three years.
Uh, something known as IPO, there's KTO, ORPO, and SIMPO.
They all descend from DPO.
I want to share just the abstract.
of the Stanford researchers original paper, which which was the the breakthrough beyond that earlier reinforcement learning from human feedback.
So they explain their invention by writing
While large scale, unsupervised language models learn broad world knowledge and some reasoning skills, achieving precise control of their behavior is difficult due to the completely unsupervised nature of their training.
Existing methods for gaining such steerability, collect human labels of the you know feedback of the relative quality of model generations, you know, model output, and fine-tune the unsupervised language model to align with these preferences, often with reinforcement learning from human feedback, RLHF.
However, they write, RLHF is a complex and often unstable procedure, first fitting a reward model that reflects the human preferences, and then fine-tuning the large unsupervised LM using reinforcement learning to maximize this estimated reward without drifting too far from the original model.
In this paper
We introduce a new parameterization of the reward model in RLHF that enables extraction of the corresponding optimal policy in closed form.
I've no idea what that means, but you get us, you'll get a sense for this, allowing us to solve the standard RLHF problem with only a simplification.
Classific a simple classification loss.
The resulting algorithm, which we call direct preference optimization, is stable, performant, and computationally lightweight, eliminating the need for sampling from the language model during fine tuning or performance significant hyperparameter tuning, whatever that is.
Our experiments show that DPO can fine-tune language models to align with human preferences as well as or better than existing methods.
Actually, it's vastly better.
The completely obsolete at everything that came before.
Notably, fine-tuning with DPO exceeds PPO-based RLHF inability to control.
Showing 1921–1940 of 17,616 · page 97 of 881 ← Previous Next →