FlowRL: Distribution Matching for LLM Reasoning

Audio

Description

This September 2025 paper introduces FlowRL, a novel reinforcement learning (RL) algorithm for large language models (LLMs) that shifts the optimization objective from reward maximization to reward distribution matching via flow balancing. Traditional RL methods like PPO and GRPO tend to over-optimize high-reward paths, leading to limited solution diversity and mode collapse, particularly in complex tasks like long Chain-of-Thought (CoT) reasoning. FlowRL addresses this by minimizing the reverse KL divergence between the policy and a reward-weighted target distribution, which is shown to be equivalent to the trajectory balance loss from GFlowNets, thereby jointly promoting reward and entropy maximization. Through experiments on math and code reasoning benchmarks, FlowRL demonstrates significant performance gains—an average of 10.0% over GRPO and 5.1% over PPO on math tasks—by generating substantially more diverse and generalizable reasoning trajectories.Source:https://arxiv.org/pdf/2509.15207

Transcription

This episode hasn't been transcribed yet

Help us prioritize this episode for transcription by upvoting it.

0 upvotes

🗳️ Sign in to Upvote

Popular episodes get transcribed faster

AI Post Transformers

This episode hasn't been transcribed yet

Other recent transcribed episodes

13:00H | 21 DIC 2025 | Fin de Semana

10:00H | 21 DIC 2025 | Fin de Semana

12:00H | 20 DIC 2025 | Fin de Semana

2ª PARTE | 06 ENE 2026 | EL PARTIDAZO DE COPE

3ª PARTE | 22 ENE 2026 | EL PARTIDAZO DE COPE

3ª PARTE | 04 MAR 2026 | EL PARTIDAZO DE COPE

Sign in to Audioscrape

Share this moment