A Clean Slate for Offline Reinforcement Learning

Offline reinforcement learning promises to train agents from fixed datasets without costly or risky exploration, but the field lacks a clear foundation for evaluating progress. This paper exposes how inconsistent evaluation protocols, hidden hyperparameter tuning, and implementation differences have obscured algorithmic comparisons. The authors introduce a transparent evaluation framework that explicitly accounts for limited online tuning budgets, provide clean single-file implementations of major algorithms, and propose Unifloral—a unified compositional framework that reveals how existing methods share underlying components. Within this framework, they develop two new algorithms, TD3-AWR and MoBRAC, demonstrating that systematic recombination of existing ideas can outperform established baselines when evaluated under realistic deployment constraints.
Script
Offline reinforcement learning lets robots learn from past experience without touching the real world during training. But here's the problem: when researchers report that algorithm A beats algorithm B, they're often comparing different amounts of hidden online tuning, different code implementations, and different subsets of benchmarks.
The authors make the online tuning budget explicit. They sample multiple hyperparameter configurations, train policies for each, then simulate what happens when you can only afford a limited number of real-world test episodes to pick the best one. Eight candidate policies compete in a bandit algorithm that balances exploring uncertain options and exploiting promising performers.
They organize algorithms into a genealogy based on shared components and implement them in clean, single-file code that changes only the lines necessary for each variant. This Unifloral framework reveals that methods like ReBRAC, IQL, and TD3-BC differ mainly in how they combine critic objectives, actor regularization, and model-based rollouts. The implementations run over 100 times faster than prior libraries.
No existing algorithm dominates across all datasets. ReBRAC and IQL perform best overall, but each still fails on some tasks. Model-based methods like MOPO collapse outside locomotion environments, often underperforming even simple behavioral cloning because their hyperparameters were overfit to MuJoCo tasks and their policy optimizers are weak.
Within Unifloral, the authors build two new algorithms by recombining existing components. TD3-AWR merges ReBRAC's value optimization with IQL's advantage-weighted behavior cloning and outperforms both parents on most datasets. MoBRAC replaces weak model-based policy optimizers with ReBRAC, making model-based methods competitive beyond locomotion for the first time.
The paper argues that offline reinforcement learning needs transparent evaluation before the field can identify real algorithmic progress. When tuning budgets, implementation details, and benchmark selection remain hidden, reported improvements may reflect engineering effort rather than algorithmic insight. You can explore the clean implementations and unified framework yourself at EmergentMind.com, where you can create your own video explanations of this and other research.