Scaling DART and evaluating cross-domain effectiveness

Determine whether DART—Depth-as-Target Pretraining for Surgical Vision Foundation Models—remains effective when applied to larger vision-transformer backbones and to domains other than surgical vision.

Background

The experiments evaluate DART on surgical data using ViT-S and ViT-B backbones. The method incorporates pseudo-depth reconstruction at masked iBOT patch locations while retaining RGB-only fine-tuning and inference.

The paper explicitly identifies the generality of the method beyond the tested backbone scales and surgical domain as unresolved. Answering this question would establish whether the observed representation-learning gains are specific to the experimental setting or persist with larger models and in other application domains.

References

Our evidence is also confined to surgical data and to ViT-{S,B} backbones, so whether DART is effective at larger scales or in other domains remains an open question.

DART: Depth-as-Target Pretraining for Surgical Vision Foundation Models  (2609.04555 - Han et al., 3 Sep 2026) in Section 5, “Conclusion and Future Work”