Determine the optimal hand-latent dimensionality for online residual reinforcement learning
Determine whether nine latent dimensions per hand is optimal for online residual reinforcement learning under different pretrained foundation policies, robot embodiments, and interaction budgets, independently of the 32-dimensional action-interface constraint of the pretrained \(\pi_{0.5}\) checkpoint.
References
Although the codec sweep characterizes reconstruction quality at several latent widths, we have not decoupled this interface constraint from the effect of latent dimensionality on online RL. Consequently, our results do not establish that nine dimensions per hand is optimal for other foundation policies, embodiments, or interaction budgets; evaluating latent residual learning with independently variable action capacity remains an important direction for future work.