DPO
concept
4 mentions
4 recordings
first heard Nov 2024
last heard 21 May
±0 this month
Direct Preference Optimization, recent RL method for LLM fine‑tuning
Trend
No mentions in the last 30d.Widen the range to see when DPO was said.
Moments
newest first · ▶ plays the momentYann DuboisThe MAD Podcast with Matt Turck · OpenAI's Yann Dubois: Why AI Progress Suddenly Feels Real · 42:28 · 21 May
…People used to have different methods like PPO and and DPO and like people seem to have really converged to this one.…
UnknownLatent Space: The AI Engineer Podcast · [State of Post-Training] From GPT-4.1 to 5.1: RLVR, Agent & ... · 8:39 · 31 Dec
…I would say like uh a lot of people at the time, this there's still like this whole PPO versus DPO discussion that's there.…
Jesse Hoogland"The Cognitive Revolution" · Embryology of AI: How Training Data Shapes AI Development w/... · 29:14 · Jun 2025
…RLHF, constitutional AI, DPO, deliberative alignment, all of these techniques are basically just modifications on the same deep learning processes with different data.…
4 moments · sign in to read them all