RLTL;DR: Learning from Self-Generated Feedback
This presentation explores RLTL;DR, a breakthrough approach to reinforcement learning that solves a critical failure mode in verifiable-reward systems. When standard RL methods fail because all rollouts produce zero reward, RLTL;DR generates compact, self-written feedback insights from failed attempts and backpropagates through them to improve future performance. Remarkably, training solely on these short insight labels nearly matches the full system's performance while using 99% fewer backward tokens, suggesting that the useful learning target may be a compact representation of the correction rather than the successful trajectory itself.Script
Standard reinforcement learning with verifiable rewards hits a wall when all rollouts fail. You collect trajectories, the verifier says no to every one, and the policy receives no useful gradient. RLTL;DR breaks through by turning those failures into compact, self-generated lessons.
The method replaces independent sampling with sequential retries on the same task. After each failure, the policy analyzes the verifier output and writes a single-sentence insight describing what should change. Subsequent attempts receive these accumulated insights, and critically, the training objective backpropagates through those insight tokens even though they never appear at test time.
Before any parameter update, sequential insights already change the game. On tasks where 128 independent attempts find essentially zero solutions, sequential feedback solves 57 percent of tasks compared to just 6 percent with independent sampling. The exploration distribution becomes dramatically richer.
Here is where the paper delivers its sharpest result. Training solely on task to insight pairs, without any GRPO loss on solution trajectories, recovers nearly the full performance of the complete system. The deduplicated variant achieves competitive accuracy while backpropagating through only 68 thousand tokens instead of 12 million. The useful learning target appears to be the compact correction, not the full trajectory.
More detailed feedback can assist the conditioned rollout while harming unaided learning. Diagnostic paragraphs and corrected code improve immediate task success but reduce transfer when internalized, because they contain episode-specific details that are hard to predict from the task alone. Short procedural rules suppress this noise and provide a cleaner supervision signal.
The training mechanism is conceptually simple but practically powerful. Gradient updates on task to insight mappings alter the policy so it applies the relevant correction while generating a normal solution, even though it never predicts the insight at test time. Visit EmergentMind.com to explore this paper further and create your own video summaries of the research you care about.