Systematic investigation of model size effects relative to pretraining on forgetting

Investigate systematically how model size influences forgetting and Negative Backward Transfer in Vision-Language-Action models for continual robot learning, and determine how model size interacts with pretraining to affect resistance to forgetting.

Background

The authors provide initial evidence that larger models trained from scratch can also reduce forgetting, and that increasing the vision backbone and language-action component sizes can lower NBT. However, this is presented as a preliminary investigation with limited scope.

They explicitly state that a more systematic study is needed to disentangle the effects of model size from pretraining in shaping continual learning behavior and forgetting dynamics.

References

While this serves an initial investigation, we will leave more systematic study of model size (in relation to pretraining) in future work.

Pretrained Vision-Language-Action Models are Surprisingly Resistant to Forgetting in Continual Learning  (2603.03818 - Liu et al., 4 Mar 2026) in Appendix C (Study on Other Factors that Contribute to VLA's Continual Learning Behavior)

Finally, while we demonstrate the effect from co-observation across supervised and self-supervised paradigms on moderate-scale vision tasks, representational dynamics can shift at massive scales. Future work must investigate whether billion-parameter over-parameterization naturally mitigates or exacerbates the need for simultaneous data observation.

Forgetting, plasticity, and co-observation: a third facet of continual learning  (2608.18803 - Hess et al., 19 Aug 2026) in Section Limitations