Papers
Topics
Authors
Recent
Search
2000 character limit reached

RankQ: Offline-to-Online Reinforcement Learning via Self-Supervised Action Ranking

Published 11 May 2026 in cs.AI and cs.RO | (2605.11151v1)

Abstract: Offline-to-online reinforcement learning (RL) improves sample efficiency by leveraging pre-collected datasets prior to online interaction. A key challenge, however, is learning an accurate critic in large state--action spaces with limited dataset coverage. To mitigate harmful updates from value overestimation, prior methods impose pessimism by down-weighting out-of-distribution (OOD) actions relative to dataset actions. While effective, this essentially acts as a behavior cloning anchor and can hinder downstream online policy improvement when dataset actions are suboptimal. We propose RankQ, an offline-to-online Q-learning objective that augments temporal-difference learning with a self-supervised multi-term ranking loss to enforce structured action ordering. By learning relative action preferences rather than uniformly penalizing unseen actions, RankQ shapes the Q-function such that action gradients are directed toward higher-quality behaviors. Across sparse reward D4RL benchmarks, RankQ achieves performance competitive with or superior to seven prior methods. In vision-based robot learning, RankQ enables effective offline-to-online fine-tuning of a pretrained vision-language-action (VLA) model in a low-data regime, achieving on average a 42.7% higher simulation success rate than the next best method. In a high-data setting, RankQ improves simulation performance by 13.7% over the next best method and achieves strong sim-to-real transfer, increasing real-world cube stacking success from 43.1% to 84.7% relative to the VLA's initial performance.

Authors (2)

Summary

  • The paper presents RankQ, a novel reinforcement learning framework that uses self-supervised action ranking to improve offline-to-online learning, addressing issues with uniform pessimism in offline RL.
  • RankQ stands out by achieving the most consistent performance across various benchmarks, including significant improvements in both low and high-data regimes, such as a 43.1% to 84.7% success rate in real-world cube-stacking.
  • RankQ demonstrated remarkable improvements in value estimation, speed, and sample efficiency, showing that value shaping benefits extend to more decisive policies and utilizes only 4 batched critic evaluations per iteration, making it approximately 2.8 times faster in its backward pass compared to methods like CQL and Cal-QL

Motivation and problem statement

Offline-to-online reinforcement learning (RL) combines offline pretraining on static datasets with online fine-tuning to improve sample efficiency, but its effectiveness hinges on the quality of the learned critic. When dataset coverage is limited, temporal-difference (TD) learning systematically overestimates the value of out-of-distribution (OOD) actions, producing harmful actor updates during fine-tuning. The dominant remedy—pessimism, as instantiated by Conservative Q-Learning (CQL) and Calibrated Q-Learning (Cal-QL)—penalizes policy-sampled actions relative to dataset actions. The authors argue that this uniform pessimism has two structural defects: it is unbounded in CQL's case, yielding excessively large ∂Q/∂a\partial Q/\partial a gradients that destabilize actor updates, and it implicitly assumes dataset actions are superior to all unseen ones, which fails when the offline data is dominated by suboptimal or failed rollouts. RankQ addresses both issues by replacing uniform pessimism with a self-supervised multi-term ranking objective that enforces structured ordering among action qualities.

Method

RankQ augments the standard TD loss with pairwise ranking terms built from a softplus-based ranking function R(s,a+,a−)=sp(Qθ(s,a−)−Qθ(s,a+))\mathcal R(s, a^+, a^-) = \mathrm{sp}(Q_\theta(s,a^-) - Q_\theta(s,a^+)). The dataset is partitioned into success (Dsuccess\mathcal D_{\text{success}}) and failure (Dfailure\mathcal D_{\text{failure}}) trajectories, and for each successful transition four classes of self-supervised suboptimal actions are constructed: noisy perturbations aϵ=a+ϵa_\epsilon = a + \epsilon, doubly noisy a2ϵa_{2\epsilon}, uniformly random ara_r, and permuted actions drawn from unrelated states. Three ordering constraints are imposed: (i) successful actions must outrank all four suboptimal variants; (ii) a "chain" ordering among negatives, Q(s,aϵ)>Q(s,a2ϵ)>Q(s,ar)Q(s,a_\epsilon) > Q(s,a_{2\epsilon}) > Q(s,a_r), using action-space proximity to success as a heuristic for relative quality; and (iii) failure-dataset actions must merely outrank random actions, with no stronger constraint due to the absence of a meaningful quality notion among failures. The full objective weights the success and failure terms by α0\alpha_0 and α1\alpha_1 (defaults of 1, with R(s,a+,a−)=sp(Qθ(s,a−)−Qθ(s,a+))\mathcal R(s, a^+, a^-) = \mathrm{sp}(Q_\theta(s,a^-) - Q_\theta(s,a^+))0 for antmaze where successes constitute under 5% of mini-batch samples) plus R(s,a+,a−)=sp(Qθ(s,a−)−Qθ(s,a+))\mathcal R(s, a^+, a^-) = \mathrm{sp}(Q_\theta(s,a^-) - Q_\theta(s,a^+))1 for the noise scale.

The key mechanistic claim is that these constraints shape the Q-value landscape so that action gradients R(s,a+,a−)=sp(Qθ(s,a−)−Qθ(s,a+))\mathcal R(s, a^+, a^-) = \mathrm{sp}(Q_\theta(s,a^-) - Q_\theta(s,a^+))2 consistently point toward higher-quality regions, rather than toward arbitrary dataset actions. A toy example visualizing 2D Q-manifolds shows CQL producing values and gradient magnitudes several orders of magnitude larger than other methods, while Cal-QL's calibration floor bounds pessimism but still directs many sampled actions' ascent trajectories toward failure actions rather than successes. Notably, RankQ requires only 4 batched critic evaluations per iteration versus 20 for CQL/Cal-QL, which translates into roughly a R(s,a+,a−)=sp(Qθ(s,a−)−Qθ(s,a+))\mathcal R(s, a^+, a^-) = \mathrm{sp}(Q_\theta(s,a^-) - Q_\theta(s,a^+))3 faster backward pass on VLA-scale models—a practical advantage beyond the algorithmic one.

D4RL benchmark results

The method is evaluated against seven baselines (CQL, CQL+SAC, Cal-QL, Cal-QL+SAC, SAC+OFF, SAC, Hybrid RL) on antmaze-medium/large and adroit-pen/door/relocate sparse-reward tasks, using the Cal-QL codebase with tuned parameters. Several findings stand out:

  • Only the full offline-to-online variants (CQL, Cal-QL, RankQ) achieve non-zero success across all environments; switching to plain SAC after offline training (the "+SAC" variants) collapses performance entirely on antmaze-large, indicating that naive online switching destroys capabilities acquired offline.
  • RankQ achieves the most consistent overall performance, matching the strongest baseline on antmaze-medium, adroit-pen, and adroit-door while outperforming all baselines on antmaze-large (0.912 success on large-play versus 0.828 for CQL, and 0.847 on large-diverse versus 0.740 for Cal-QL) and adroit-relocate.
  • Pure online SAC fails completely on antmaze, underscoring the long-horizon difficulty.

An important caveat: the authors report that several baseline algorithms diverged with the original released code, and they introduced gradient clipping, larger mini-batches, and lower learning rates that raised baseline performance above originally published numbers. This strengthens the fairness of comparisons but means results are not directly comparable to prior literature.

VLA fine-tuning in low-data regimes

The paper's most distinctive contribution is applying offline-to-online RL to a pretrained vision-language-action model (R(s,a+,a−)=sp(Qθ(s,a−)−Qθ(s,a+))\mathcal R(s, a^+, a^-) = \mathrm{sp}(Q_\theta(s,a^-) - Q_\theta(s,a^+))4, flow-matching head, frozen VLM backbone), using an off-policy PPOFlow-derived formulation called SACFlow with action chunks of size 4 at 5 Hz control and a single integration step (R(s,a+,a−)=sp(Qθ(s,a−)−Qθ(s,a+))\mathcal R(s, a^+, a^-) = \mathrm{sp}(Q_\theta(s,a^-) - Q_\theta(s,a^+))5). In the low-data regime—200 offline self-rollouts and only 8 online rollouts per update across three WidowX manipulation tasks—the initial success rates for cube-stacking and spoon-into-bowl are roughly 20%, meaning about 80% of the offline data consists of failures. This is precisely the regime where uniform pessimism toward dataset actions should be counterproductive, and the results confirm it: RankQ is the only method that significantly improves over the baseline, achieving final success rates of 0.930, 0.940, and 0.800 on carrot-onto-plate, cube-stacking, and spoon-into-bowl respectively—exceeding the next best non-RankQ method by 9.5%, 64.3%, and 54.2%. All pessimistic baselines largely plateau near their imitation starting points.

In the high-data regime (800 rollouts, ~8% initial success, extensive domain randomization, 192 rollouts per update), CQL and Cal-QL do improve over the baseline but remain far less sample-efficient than RankQ, which reaches a training success rate of 0.880 versus 0.743 for Cal-QL and reduces task completion time to 8.36 seconds versus 10.51 for Cal-QL. The improvement in execution speed indicates that value shaping benefits extend beyond binary success to more decisive policies.

Sim-to-real transfer

Zero-shot deployment of one RankQ cube-stacking policy on a real WidowX 250S over 144 trials across a 3×3 configuration grid yields a real-world success rate of 84.7% versus 43.1% for the baseline VLA—a 41.6 percentage-point gain—with partial success rate rising from 77.8% to 98.6% and time-to-finish dropping from 14.81 s to 12.34 s with reduced variance. This constitutes direct evidence that the ranking-shaped critic transfers outside simulation rather than exploiting simulator artifacts.

Analysis

Three supporting analyses reinforce the mechanism. First, ablations show component contributions are environment-dependent: removing the chain loss degrades antmaze-large-play and adroit-relocate, removing permuted-action ranking hurts adroit-door, and increasing R(s,a+,a−)=sp(Qθ(s,a−)−Qθ(s,a+))\mathcal R(s, a^+, a^-) = \mathrm{sp}(Q_\theta(s,a^-) - Q_\theta(s,a^+))6 to 0.30 hurts antmaze-large-play—though the easiest environments are insensitive, and the authors concede the optimal configuration may vary by task. Second, R(s,a+,a−)=sp(Qθ(s,a−)−Qθ(s,a+))\mathcal R(s, a^+, a^-) = \mathrm{sp}(Q_\theta(s,a^-) - Q_\theta(s,a^+))7 statistics show CQL exhibits the largest gradient spikes during offline training (consistent with unbounded pessimism), Cal-QL produces smaller spikes, and RankQ remains stable with minimal distributional shift at the offline-to-online transition. Third, critic accuracy measurements confirm RankQ ranks successful actions above all four suboptimal categories most accurately; notably, CQL and Cal-QL implicitly learn similar rankings but far more slowly, suggesting RankQ's explicit objective aligns with—and accelerates—an implicit behavior of existing state-of-the-art methods.

Limitations and open questions

Several limitations deserve plain acknowledgment. The method presupposes reliable success/failure labels for partitioning the dataset, which holds for rule-based reward detection but would require extension for dense or ambiguous rewards. The heuristic that action-space proximity to a successful action proxies relative quality is an assumption, not a derived property, and the permuted-action construction assumes state-independent action plausibility. Hyperparameter sensitivity exists (R(s,a+,a−)=sp(Qθ(s,a−)−Qθ(s,a+))\mathcal R(s, a^+, a^-) = \mathrm{sp}(Q_\theta(s,a^-) - Q_\theta(s,a^+))8 is specifically motivated by antmaze's class imbalance), and the ablations indicate no single configuration dominates across environments. D4RL coverage was limited to seven environments because the reused codebase omitted others such as franka-kitchen. Finally, the sim-to-real result covers a single task family (cube stacking) on one robot platform, leaving open whether the ranking objective generalizes across contact-rich or deformable-object manipulation, and how the approach behaves when success detection is noisy or when the offline dataset contains neither clean successes nor clean failures.

Conclusion

RankQ replaces uniformly pessimistic OOD penalization with structured self-supervised action ranking, shaping the Q-landscape so that action gradients point toward higher-quality behaviors. The empirical case is strong: competitive-or-better performance against seven baselines on sparse-reward D4RL tasks, uniquely effective VLA fine-tuning in a severely data-limited regime (+42.7% average over the next best method), faster and more successful high-data training (+13.7% success, +25.7% completion speed), and a substantial sim-to-real gain (43.1% → 84.7% real-world success). The results position explicit relative-action-quality modeling as a viable alternative to pessimistic value estimation, particularly when offline datasets are dominated by failures.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 1 tweet with 0 likes about this paper.