DPO

concept
4 mentions 4 recordings first heard Nov 2024 last heard 21 May ±0 vs the 12 months before

Direct Preference Optimization, recent RL method for LLM fine‑tuning

Trend

mentions per week · public audio 7d30d6M12M
1 · 18 May SepDecMarJunnow

Peak in the week of 18 May — 1 mentions across 1 series.

Moments

newest first · ▶ plays the moment
Yann DuboisThe MAD Podcast with Matt Turck · OpenAI's Yann Dubois: Why AI Progress Suddenly Feels Real · 42:28 · 21 May
…People used to have different methods like PPO and and DPO and like people seem to have really converged to this one.…
· Open transcript →
UnknownLatent Space: The AI Engineer Podcast · [State of Post-Training] From GPT-4.1 to 5.1: RLVR, Agent & ... · 8:39 · 31 Dec
…I would say like uh a lot of people at the time, this there's still like this whole PPO versus DPO discussion that's there.…
· Open transcript →
Jesse Hoogland"The Cognitive Revolution" · Embryology of AI: How Training Data Shapes AI Development w/... · 29:14 · Jun 2025
…RLHF, constitutional AI, DPO, deliberative alignment, all of these techniques are basically just modifications on the same deep learning processes with different data.…
· Open transcript →
Nathan Lambert"The Cognitive Revolution" · Everything You Wanted to Know About LLM Post-Training, with ... · 43:50 · Nov 2024
…And I don't know the intuitions off the top of……
4 moments · sign in to read them all