- The paper unifies Bellman-error minimization and f-divergence-regularized linear programming, showing that methods such as IQL, OptiDICE, and RelaxDICE are special cases of a common framework.
- The proposed piecewise Flex-f divergence adapts asymmetric penalties using learned scaling factors and thresholds, enabling offline RL algorithms to adjust conservatism to behavior-policy stochasticity and dataset diversity.
- Flex-f-DICE improves performance over OptiDICE in challenging datasets, including Hopper 4-policy data (99.0 vs. 77.9) and Pen-human data (39.6 vs. 10.2), while Flex-f-Q generally matches IQL.
Motivation: dataset characteristics that break standard offline RL
Offline RL algorithms are typically designed around a pessimism principle: improve over the behavior policy while constraining the learned policy to remain within the support of the static dataset. The authors of this paper argue that this uniform pessimism is mismatched to two properties common in practical data collection but underrepresented in benchmarks like D4RL: (i) behavior policies with very low stochasticity, often near-deterministic or rule-based due to safety and cost constraints, and (ii) datasets aggregated from many behavior policies spanning diverse expertise levels. Limited exploration impairs value estimation, while constraints tied to low-expertise behavior policies can be overly conservative.
The paper's empirical motivation rests on two observations. First, existing dataset-level statistics — Positive Scaled Variance (PSV) and Normalized Expected Return (NER) for returns, and SACo for exploration — overlap heavily across datasets that differ in the number of behavior policies and their stochasticity. Consequently, these distribution-level measurements cannot identify what setting an individual dataset comes from, even though algorithm performance differs drastically across those settings. Second, standard algorithms (IQL, OptiDICE) suffer substantial performance degradation as behavior policy stochasticity drops or the number of mixed behavior policies grows. The paper's response is not a new algorithm class but a more general theoretical account of where the constraint enters RL optimization, plus a parameterized divergence family that adapts it.
The paper builds on the Linear Programming formulation of MDPs, using the primal problem minimizing (1−γ)p0(s)ν(s) subject to Bellman-inequality constraints, its dual maximizing discounted reward over occupancy measures satisfying the Bellman flow constraint, and the density ratio ζ(s,a)=d(s,a)/dD(s,a).
The first technical contribution is an alternative form for the primal objective term LP. Beyond the standard α(s)ν(s), the authors show via the performance difference lemma that −eν(s,a) — the negated TD error (advantage estimate) evaluated on the dataset — is also a valid choice, since (1−γ)VD(s) is a fixed constant for a static dataset. Substituting this into the Lagrangian and eliminating ζ through the convex conjugate yields the unconstrained objective minν−eν+g(eν).
A theorem then characterizes when this objective is a proper loss: for convex g whose conjugate satisfies (1−γ)p0(s)ν(s)0 and (1−γ)p0(s)ν(s)1, the function (1−γ)p0(s)ν(s)2 is convex with minimum value zero at (1−γ)p0(s)ν(s)3. This means any RL algorithm minimizing Bellman error with a compatible loss implicitly performs an (1−γ)p0(s)ν(s)4-divergence-regularized LP; e.g., the MSE loss of Q-learning corresponds exactly to the (1−γ)p0(s)ν(s)5-divergence.
Two relaxations distinguish this derivation from prior DICE-family work. First, the constraint (1−γ)p0(s)ν(s)6 is removed by converting the primal inequality into an equality constraint — justified because the optimal value function satisfies the Bellman equation exactly. This provides theoretical grounding for prior relaxed-positivity methods and admits divergences defined over all of (1−γ)p0(s)ν(s)7 rather than (1−γ)p0(s)ν(s)8, at the cost of losing the importance-sampling interpretation of negative (1−γ)p0(s)ν(s)9 (the paper points to quasiprobability likelihood-ratio estimation as a possible interpretation). Second, the paper derives a unified form showing duality between penalizing the Bellman residual ζ(s,a)=d(s,a)/dD(s,a)0 on the primal side and the Bellman occupancy residual ζ(s,a)=d(s,a)/dD(s,a)1 on the dual side: applying an equality constraint on one residual makes the primal equivalent to an unconstrained dual with penalty ζ(s,a)=d(s,a)/dD(s,a)2, and vice versa. Existing dual algorithms such as OptiDICE, RelaxDICE, and PORelDICE, and IQL's expectile regression, are shown to be special cases of this framework.
Flexible ζ(s,a)=d(s,a)/dD(s,a)3-divergence
The core methodological proposal is a piecewise divergence function ζ(s,a)=d(s,a)/dD(s,a)4 formed by joining two scaled base functions from the valid ζ(s,a)=d(s,a)/dD(s,a)5-divergence family (ζ(s,a)=d(s,a)/dD(s,a)6, ζ(s,a)=d(s,a)/dD(s,a)7), separated at a threshold ζ(s,a)=d(s,a)/dD(s,a)8: ζ(s,a)=d(s,a)/dD(s,a)9 scales the penalty below LP0 and LP1 above it. Continuity of the function and its conjugate impose two conditions, satisfied by choosing the linear coefficient LP2 and offset LP3 accordingly. A linear difference added to one branch is shown to act only as a constant bias under expectation over LP4, so it does not affect the learning objective. For algorithms that need the inverse map (e.g., OptiDICE-style estimation of LP5 from LP6), closed-form expressions for LP7 and its inverse are derived.
The intuition is asymmetric regularization: larger derivative magnitude on the positive side of LP8 penalizes overestimation more aggressively, and vice versa. Notably, IQL's expectile regression with parameter LP9 is recovered exactly by α(s)ν(s)0 bases with α(s)ν(s)1 and α(s)ν(s)2, situating expectile-based methods inside this family.
Because no automatic selection procedure is proposed, the paper supplies heuristics: α(s)ν(s)3 are estimated from the cosine similarity between a behavior-cloned policy's action probabilities and α(s)ν(s)4 across sampled batches (smoothed with an EMA), and α(s)ν(s)5 is obtained by mapping the EMA-smoothed mean α(s)ν(s)6 through the inverse derivative at default parameters. These heuristics introduce new hyperparameters and depend on numerical-stability clamps (e.g., restricting α(s)ν(s)7 to α(s)ν(s)8 for α(s)ν(s)9 estimation); the authors explicitly defer fully optimizable −eν(s,a)0 and −eν(s,a)1 to future work.
Experimental validation
Two algorithms instantiate the framework. Flex-−eν(s,a)2-Q approximates −eν(s,a)3 as −eν(s,a)4, updates −eν(s,a)5 by semi-gradient descent on the direct objective, and uses −eν(s,a)6 as −eν(s,a)7 — conceptually IQL with the flexible divergence substituted. Flex-−eν(s,a)8-DICE is OptiDICE with its divergence replaced by Flex-−eν(s,a)9. Experiments cover MuJoCo locomotion (Hopper, Walker2d, Ant, HalfCheetah), Fetch manipulation (Push, PickAndPlace) with HER resampling, and AdroitHand (Pen, Hammer) with cloned and human D4RL datasets. New MuJoCo/Fetch datasets were collected with SAC behavior policies at fixed variance 0.0 (fully deterministic), mixed across 2, 4, and 10 behavior policies of varying expertise; results average five seeds.
Three findings stand out:
| Setting |
Result |
| Flex-(1−γ)VD(s)0-Q vs. IQL |
Performance within (1−γ)VD(s)1 almost everywhere, occasionally higher |
| Flex-(1−γ)VD(s)2-DICE vs. OptiDICE |
Improved in nearly all settings; e.g., Hopper 4-p: 99.0 vs. 77.9; Pen-human: 39.6 vs. 10.2 |
| Base-function ablation |
Best combination varies by environment and algorithm |
Flex-(1−γ)VD(s)3-Q's parity with IQL empirically validates the general LP form despite the swapped (1−γ)VD(s)4 and removed positivity constraint. Flex-(1−γ)VD(s)5-DICE's gains are strongest precisely where OptiDICE collapses — deterministic, multi-policy datasets — including a nearly fourfold improvement on Pen-human (39.6 vs. 10.2). Since Flex-(1−γ)VD(s)6-DICE shares OptiDICE's soft-(1−γ)VD(s)7 base functions in MuJoCo, those gains isolate the contribution of the adaptively estimated (1−γ)VD(s)8 and (1−γ)VD(s)9.
The ablation over base-function pairs (ζ0/KL for ζ1; ζ2/KL/Le-Cam/Hellinger for ζ3) shows no universal best pair, supporting the paper's hypothesis that constraint level must be environment- and algorithm-dependent, though Hellinger–ζ4 is relatively consistent for Flex-ζ5-DICE. One notable observation: Flex-ζ6-Q diverged only when both branches used KL, suggesting that a non-KL lower branch can mitigate Q-value overestimation in semi-gradient optimization.
Limitations and open questions
The paper is candid about several constraints on its claims. The heuristic estimators for ζ7 and ζ8 are ad hoc, rely on EMA smoothing and stability clamps, and are not derived from an optimization principle; integrating them directly into the objective remains open. The experiments do not compare against architecturally different algorithms (only TD3BC and CQL appear in supplementary material), so the contribution should be read as improving compatible constrained-optimization bases rather than establishing state-of-the-art across offline RL. Hammer results are near-zero for all methods, indicating the hardest dexterous tasks remain unsolved in this regime. Finally, removing ζ9 sacrifices the probabilistic interpretation of the density ratio, and the paper leaves open how quasiprobability frameworks might restore it. Whether a principled, fully automated adaptation mechanism for the divergence shape exists — beyond the proposed heuristics — is the central question the paper does not answer.
Conclusion
This paper makes two contributions: a generalized LP view of RL in which Bellman-error minimization is shown to carry an implicit minν−eν+g(eν)0-divergence penalty, unifying primal and dual formulations under interchangeable residuals and penalty functions; and a flexible, piecewise minν−eν+g(eν)1-divergence family with adaptive scaling and thresholding, instantiated in Flex-minν−eν+g(eν)2-Q and Flex-minν−eν+g(eν)3-DICE. Empirically, the framework matches IQL and substantially improves OptiDICE on datasets with deterministic behavior policies and heterogeneous expertise mixtures, validating both the theory and the practical value of dataset-adaptive constraint levels.