Sensitivity of DART representations to pseudo-depth quality

Characterize how sensitive the representations learned by DART—Depth-as-Target Pretraining for Surgical Vision Foundation Models—to the quality of pseudo-labeled depth maps and to the choice of the monocular depth estimator used to generate them.

Background

DART uses pseudo-labeled depth generated by the off-the-shelf Depth Anything V2-Large monocular depth estimator as a supervisory target during DINOv2 pretraining. The depth maps are imperfect because monocular depth estimation is ill-posed and because surgical images may contain blur, blood, borders, and other challenging visual conditions.

Although the reported experiments show that DART improves downstream surgical vision tasks despite noisy targets, the paper does not determine how these improvements vary with pseudo-depth accuracy or with alternative depth estimators. Establishing this sensitivity would clarify the robustness and practical reliability of the depth-as-target pretraining strategy.

References

We do not characterize how sensitive the learned representations are to depth quality or the choice of the estimator.

DART: Depth-as-Target Pretraining for Surgical Vision Foundation Models  (2609.04555 - Han et al., 3 Sep 2026) in Section 5, “Conclusion and Future Work”