Determine the optimal hand-latent dimensionality for online residual reinforcement learning

Determine whether nine latent dimensions per hand is optimal for online residual reinforcement learning under different pretrained foundation policies, robot embodiments, and interaction budgets, independently of the 32-dimensional action-interface constraint of the pretrained \(\pi_{0.5}\) checkpoint.

Background

The proposed codec compresses each 20-dimensional dexterous hand action to nine latent dimensions because the pretrained π0.5\pi_{0.5} policy accepts a 32-dimensional action interface. Fourteen dimensions are allocated to the two 7-DoF arms, leaving 18 dimensions for the two hands under the paper's symmetric layout.

Although the paper evaluates reconstruction quality at several latent widths, it does not separate the effect of latent dimensionality from the fixed action-interface capacity of the pretrained policy, nor does it establish that nine dimensions per hand yields the best online residual-RL performance. The open problem is therefore to evaluate independently variable action capacities across policies, embodiments, and interaction budgets.

References

Although the codec sweep characterizes reconstruction quality at several latent widths, we have not decoupled this interface constraint from the effect of latent dimensionality on online RL. Consequently, our results do not establish that nine dimensions per hand is optimal for other foundation policies, embodiments, or interaction budgets; evaluating latent residual learning with independently variable action capacity remains an important direction for future work.

— Towards High-DoF Dexterous Manipulation through VLA Post-Training  (2609.19666 - Zhu et al., 17 Sep 2026) in Section 'Limitations and Discussion'