Scaling DINOcular to larger and more diverse datasets

Determine whether the DINOcular self-supervised RGB-D representation preserves its reported performance gains when trained and evaluated at substantially larger data scales and on more diverse data.

Background

DINOcular is evaluated at a moderate data scale and is reported to outperform comparable RGB-D and RGB-only foundation models in that regime. The paper does not establish whether these gains persist when the framework is scaled to substantially larger and more diverse datasets, leaving the scalability of its visuospatial representations unresolved.

References

Our study focuses on moderate data scale, and we have not yet validated whether the proposed method preserves the same gains when scaled to substantially larger and diverse data.

DINOcular: Self-Supervised Visuospatial Representations  (2608.27226 - Almukhamedov et al., 27 Aug 2026) in Section Limitations