Mechanism of self-supervision benefits

Determine whether the benefits of self-supervision in vision-, text-, and state-conditioned generative control arise primarily from the states traversed by executing the base policy’s actions or from the distillation action targets themselves.

Background

The paper introduces self-supervised generative replay by executing a frozen pretrained VLA on a target robot and using the resulting observation–action trajectories as rehearsal data during fine-tuning. Although the method improves retention and adaptation, the authors do not establish whether its effectiveness is attributable to the distribution of states visited during online execution, the teacher’s action labels, or both. Resolving this distinction could guide the design of more efficient replay procedures.

References

Why does self-supervision help, is it due to the states traversed by executing the base policy's actions, % off-policy action execution or the distillation action targets themselves?

Fine-Tuning VLAs with Self-Demonstrated Generative Control for Multi-Task Manipulation  (2608.19490 - Garg et al., 19 Aug 2026) in Section 6, “Future Work and Limitations”