Determine the effects of surgical pretraining catalogue, sampling, and masking choices

Determine the effects of training-catalogue composition, sampling strategy, and masking objective on the downstream transfer of OmniRAS V-JEPA-2.1 encoders when evaluated at production-scale budgets beyond the tested small-budget screening regime.

Background

OmniRAS investigates continued self-supervised pretraining of 1B- and 2B-parameter V-JEPA-2.1 video encoders on a multi-source surgical-video catalogue. The authors conduct smaller screening experiments to examine catalogue composition, sampling temperature, and masking-objective choices, but report that these experiments do not establish consistent effects at the tested budgets.

The production-scale results are also confounded: the longest 36.86-million-sample run changes both the training budget and the catalogue composition, preventing attribution of its improved transfer to compute alone. Consequently, the general effects of catalogue composition, sampling, and masking remain unresolved and require larger, better-controlled experiments.

References

The experiments also clarify where these gains come from. Within the continued-pretraining recipe, the largest production configuration shows the clearest transfer gains, although its simultaneous catalog change prevents attributing them to compute alone; catalog, sampling, and masking effects remain unresolved at the tested screening budget.

— OmniRAS: Standardizing Foundation Model Training and Evaluation in Robot-Assisted Surgery  (2608.31048 - Borgioli et al., 31 Aug 2026) in Section Conclusion