Verify persistence of LeVLJEPA’s advantages at larger model and data scales

Establish whether the dense semantic feature advantages and competitive global-feature performance observed for LeVLJEPA with a ViT-B/16 backbone trained on the Datacomp-L scale persist when scaling to larger vision encoders and larger multimodal datasets.

Background

All main experiments use a ViT-B/16 backbone and Datacomp-L-scale pretraining. While LeVLJEPA trains stably and is competitive at this scale, it remains to be shown whether its observed advantages—particularly on dense token features—carry over as model capacity and data scale increase.

References

Several questions remain open. The contrastive objectives retain a stronger mechanism for zero-shot image-text alignment, and we do not close this gap; whether the dense-feature advantage of non-contrastive pretraining can be combined with competitive alignment within a single objective is a natural direction for future work. Our experiments are conducted with a ViT-B/16 backbone, and while LeVLJEPA trains stably and remains competitive at the scale of Datacomp-L, establishing that these advantages persist at larger model and data scales remains important to verify.

— LeVLJEPA: End-to-End Vision-Language Pretraining Without Negatives  (2607.00784 - Kuhn et al., 1 Jul 2026) in Discussion (Section 7)

The data-scaling experiment of Section~\ref{sec:scaling_data} is consistent with this trajectory, with both appearance-centric and motion-centric accuracy improving under a growing corpus; a controlled comparison at internet scale remains an open question.

— LeVJEPA: Efficient & Scalable Video Pretraining without the Heuristics  (2608.27395 - Kuhn et al., 27 Aug 2026) in Discussion, paragraph beginning “The comparison to image-based pretraining suggests a broader implication.”

Our controlled comparisons are conducted on a restricted corpus at up to ViT-L scale; the behavior of the objective at the model and data scales of recent video foundation models, and the interaction of SIGReg with very large batch and model regimes, remain to be characterized, and the favorable scaling observed in Section~\ref{sec:scaling_data} makes this a promising rather than merely open question.

— LeVJEPA: Efficient & Scalable Video Pretraining without the Heuristics  (2608.27395 - Kuhn et al., 27 Aug 2026) in Discussion, paragraph beginning “Several directions remain open.”