Characterize the competence threshold for residual reinforcement learning

Characterize the intermediate reference-policy competence regime in which residual reinforcement learning becomes effective, and determine whether an approximately 50% base-policy success rate is a universal threshold across tasks, objects, and embodiments.

Background

The paper reports that residual reinforcement learning provides little measurable improvement when the frozen supervised-fine-tuning or DAgger reference policy succeeds on fewer than approximately 20% of evaluation trials, but becomes consistently useful once the success rate exceeds approximately 50%. The authors interpret this behavior as evidence that residual learning requires the reference policy to visit task-relevant states often enough for local exploration to obtain informative experience.

The unresolved issue is whether the observed intermediate regime has been systematically characterized and whether the approximately 50% threshold generalizes across different tasks, objects, and robot embodiments. Resolving this would clarify when latent residual reinforcement learning can be expected to improve an imperfect imitation policy.

References

We have not systematically characterized the intermediate regime or established that 50\% is a universal threshold across tasks, objects, or embodiments.

— Towards High-DoF Dexterous Manipulation through VLA Post-Training  (2609.19666 - Zhu et al., 17 Sep 2026) in Section 'Limitations and Discussion'