Characterize the competence threshold for residual reinforcement learning
Characterize the intermediate reference-policy competence regime in which residual reinforcement learning becomes effective, and determine whether an approximately 50% base-policy success rate is a universal threshold across tasks, objects, and embodiments.
References
We have not systematically characterized the intermediate regime or established that 50\% is a universal threshold across tasks, objects, or embodiments.
— Towards High-DoF Dexterous Manipulation through VLA Post-Training
(2609.19666 - Zhu et al., 17 Sep 2026) in Section 'Limitations and Discussion'