Papers
Topics
Authors
Recent
Search
2000 character limit reached

Scalable Patch-Level Self-Supervised Learning

Published 7 Oct 2026 in cs.CV | (2610.10013v1)

Abstract: Self-supervised learning (SSL) at scale produces powerful visual representations. However, most scalable SSL methods rely on ad hoc combinations of multiple objectives and stabilization mechanisms. Taking a step back, we ask if we can design a high-performing, yet principled SSL algorithm. Starting from the multi-view assumption, stipulating that task-relevant content is captured by the information common to different views, we construct an information-theoretic objective decomposing into interpretable terms. This derivation yields JEM, a student-teacher method that learns by aligning corresponding patch representations across views, explicitly regularized by information and structure preservation losses. JEM trains stably from 300M to 7B parameters, and, to our knowledge, is the first latent-space patch-level method demonstrated at 7B scale. Across all scales, JEM reaches strong performance on both global and dense probing tasks, on segmentation benchmarks consistently surpassing the DINOv2 algorithm, an influential foundation for today's strongest visual SSL methods. Notably, at 7B parameters, it exceeds the performance of DINOv3 on panoptic segmentation, despite being trained on 12×12\times less data without refinement stages. These results demonstrate that we can indeed design an SSL algorithm that learns strong representations, is principled and stable.

Summary

  • The paper introduces JEM, a joint-embedding self-supervised learning method that achieves robust dense visual representations without combining multiple independent losses by aligning student and teacher croppings and using mutual information regularizers and multilayer supervision.
  • JEM's robust dense representation stemming from mutual information maximization boosts performance in semantic, instance, and panoptic segmentation tasks structures across different encoders: ViT-L (303M) demonstrates a 48.1 mIoU on ADE20k; ViT-g (1.1B) reaches 51.2 mIoU; and ViT-7B gets the top 27.3 PQ on COCONut dataset and 23.9 ADE20k.
  • Ablation studies show that local crops and masking are crucial for maintaining performance in semantic segmentation tasks whereas patch alignment has a substantial impact on representational outcomes.

The paper introduces a patch-level self-supervised learning method, referred to in the experiments as JEM, whose objective is derived from a multi-view InfoMax formulation rather than assembled from independently motivated losses. Its central claim is that strong dense visual representations can be obtained with a single patch-level joint-embedding objective, provided that cross-view alignment is accompanied by explicit controls for both statistical and structural collapse. The method is evaluated with ViT-L, ViT-g, and ViT-7B encoders and is reported to scale stably from 300 million to 7 billion parameters (2610.10013).

Problem setting and motivation

Contemporary visual SSL systems frequently combine global self-distillation, masked patch prediction, entropy or redundancy regularization, and additional stabilization mechanisms. DINOv2, for example, combines image-level DINO loss, patch-level iBOT prediction, and KoLeo regularization; DINOv3 adds Gram anchoring and high-resolution refinement to preserve spatial structure during extended large-scale training. The paper questions whether this accumulation of objectives is necessary, or whether the same functional requirements can be expressed through a more coherent information-theoretic construction.

The proposed starting point is the multi-view assumption: task-relevant information is assumed to be present in the information shared by different views of the same image. The views are generated through cropping, masking, and photometric transformations. Under this assumption, the representation should retain information about a teacher target that is predictable from the student view. The paper emphasizes that the assumption is conditional on the view construction and downstream task; it is therefore used as an idealized design principle rather than as an unconditional statement about visual information (2610.10013).

The resulting method differs from image-level self-distillation in two ways. First, the learning signal is applied to spatially corresponding patch representations rather than to a class token. Second, the anti-collapse mechanism is formulated explicitly in terms of per-element mutual information and cross-element redundancy. This formulation is intended to address failure modes that are specific to dense representations and are not controlled by image-level entropy balancing alone.

Information-theoretic derivation

Let XSX_S and XTX_T denote paired student and teacher views, let R=f(XS)R=f(X_S) be the student representation, and let TT be a stochastic teacher target generated from XTX_T. The paper seeks to maximize I(R;T)I(R;T), the mutual information between the student representation and the teacher target. By the data-processing inequality,

I(XS;XT)≥I(XS;T)≥I(R;T).I(X_S;X_T) \geq I(X_S;T) \geq I(R;T).

A variational lower bound yields

I(R;T)≥I(XT;T)−EXS,XTDKL ⁣(p(T∣XT) ∥ q(T∣XS)).I(R;T) \geq I(X_T;T) - \mathbb{E}_{X_S,X_T} D_{\mathrm{KL}}\!\left( p(T\mid X_T)\,\|\,q(T\mid X_S) \right).

This decomposition has a direct algorithmic interpretation. The KL term is a cross-view alignment objective, while I(XT;T)I(X_T;T) is an anti-collapse term requiring the teacher targets to remain informative about the input. The teacher is implemented as an EMA of the student, so the anti-collapse quantity is not optimized directly through teacher parameters. Instead, the method applies corresponding regularization to student predictions and relies on the EMA dynamics to transmit the constraint to the teacher.

The target TT is factorized into elements indexed by patch position and feature layer. Under the factorization assumptions used in the derivation, the target mutual information decomposes into a sum of per-element information terms minus a total-correlation term:

XTX_T0

This decomposition identifies two distinct collapse modes. Per-element information can vanish, producing global collapse in which categorical predictions become uniform or constant. Alternatively, each element can retain information while different patches encode largely redundant information, producing structural collapse in which spatial features become indistinguishable within an image. The method assigns one regularizer to each mode.

The derivation is conceptually useful, but it does not make the practical objective assumption-free. In particular, the factorization of the conditional target distribution and the use of a structural proxy instead of direct total-correlation estimation are modeling choices. The information-theoretic lower bound motivates the losses, but does not establish that optimizing the implemented proxy is equivalent to maximizing the original mutual information.

JEM objective and architecture

JEM uses three components:

  1. Patch-wise alignment: Student and teacher views are placed on a common patch lattice. For overlapping patches, the student categorical distribution is matched to the corresponding teacher distribution using a KL divergence. The objective is applied to patch-layer pairs, allowing intermediate teacher layers and final-layer representations to contribute to training.
  2. Global collapse regularization: For every patch position, the method estimates mutual information between the input batch and the student categorical predictions. This estimate is the difference between the entropy of the batch-averaged categorical distribution and the average conditional entropy of individual predictions. Maximizing it encourages broad marginal usage of the XTX_T1 categories while keeping predictions input-dependent and relatively sharp.
  3. Structural regularization: The method does not estimate XTX_T2 directly. Instead, the final student representation is projected through learned nonlinear maps and trained to reproduce pairwise cosine similarities from early and intermediate teacher layers. The structural loss therefore preserves information about the teacher’s spatial dependency structure without forcing the final representation to equal an intermediate-layer embedding.

This design distinguishes structural regularization from a simple diversity penalty. The objective permits semantically related patches to remain similar, but discourages a degenerate representation in which all patches encode the same image-level signal. In the authors’ interpretation, intermediate-layer geometry provides a proxy for local and spatial information that must remain recoverable in the final embedding.

The view construction combines mechanisms that are often treated separately. Each sample produces one teacher view, one global student view, and four local student views. Global and teacher crops are XTX_T3 pixels, while local crops are XTX_T4 pixels. Crops are shifted on a patch-aligned grid so that overlapping patches have exact one-to-one correspondences. Eighty percent of global and local student views are masked, with mask ratios of 65% and 50%, respectively. Student and teacher views receive independent photometric transformations. The method can therefore express ordinary masked prediction as a special case, but does not require fully overlapping crops or masking in principle.

The encoders are ViTs without class tokens and with four register tokens. All models use XTX_T5 patches. The principal scaling experiments use ViT-L with approximately 303 million parameters, ViT-g with approximately 1.13 billion, and ViT-7B with approximately 6.7 billion parameters.

Ablation evidence

The ablations clarify which parts of the method are essential and which primarily improve robustness or scaling.

The full ViT-L model trained on ImageNet-22k reaches 82.8% ImageNet-1k attention-probe accuracy and 48.1 mIoU on ADE20k. Removing local crops reduces these scores to 81.2% and 46.2 mIoU. Removing masking has a smaller effect on classification, 82.3%, but reduces ADE20k performance to 45.3 mIoU. Removing the student-teacher crop shift leaves classification unchanged at 82.8% but lowers ADE20k mIoU to 46.6. By contrast, removing independent photometric augmentation is highly damaging: performance falls to 74.8% on ImageNet-1k and 31.4 mIoU on ADE20k. This result implies that the method’s dense representations depend strongly on invariance to low-level appearance changes, even though the geometric view transformations are not individually indispensable.

The composition of the alignment loss produces a similarly uneven pattern. Removing multi-layer supervision changes ImageNet accuracy from 82.8% to 81.7% and slightly improves ADE20k from 48.1 to 48.3 mIoU. Removing masked-patch alignment produces 81.3% and 46.0 mIoU. However, removing visible-patch alignment causes a severe degradation to 70.8% ImageNet accuracy and 28.1 mIoU. The authors attribute this to the role of visible-patch alignment in preserving spatial correspondence. Thus, the result contradicts the common assumption that masked patches are the defining or indispensable component of patch-level SSL: in this setup, visible-patch supervision is substantially more important.

The objective ablation isolates the collapse mechanisms. Removing alignment reduces performance to 63.1% on ImageNet-1k and 22.4 mIoU on ADE20k. Removing global collapse regularization is catastrophic, producing 10.3% classification accuracy and 0.5 mIoU. This confirms that patch-wise alignment alone admits trivial solutions. The coefficient of the global regularizer is relatively insensitive once it is nonzero across a broad range, but setting it to zero produces degeneration.

Structural regularization has little impact at ViT-L scale: removing it changes ImageNet accuracy from 82.8% to 82.3% and ADE20k mIoU from 48.1 to 48.0. Its role becomes more consequential for larger models. At ViT-g scale, disabling the term causes patch representations to become progressively more spatially uniform after approximately 250,000 iterations and reduces ADE20k performance by 2.7 mIoU. The implication is that structural collapse is a scale-dependent failure mode: it may be nearly invisible in small-model ablations while becoming important during prolonged training of larger encoders.

Benchmark performance

The evaluation uses frozen backbones and lightweight probes for image classification, semantic segmentation, instance segmentation, panoptic segmentation, and monocular depth. Comparisons are made against published SSL methods and against retrained DINOv2 baselines using a closely matched implementation and training setup.

On ImageNet-1k-trained ViT-L, JEM obtains 82.8% classification accuracy, compared with 83.8% for the controlled DINOv2 baseline. The small classification deficit is accompanied by consistent dense-task gains:

Model ImageNet-1k ADE20k mIoU VOC mIoU Cityscapes mIoU COCO mask AP ADE20k PQ COCONut PQ
DINOv2* ViT-L 83.8 45.7 83.3 66.5 12.1 18.0 22.1
JEM ViT-L 82.8 46.4 84.9 67.4 13.5 20.1 24.3

At ViT-g scale trained on ImageNet-22k, JEM again trails DINOv2 slightly on classification, 85.5% versus 86.5%, but improves substantially on dense prediction. It reaches 51.2 versus 50.0 ADE20k mIoU, 85.9 versus 83.6 VOC mIoU, 69.6 versus 68.5 Cityscapes mIoU, 14.7 versus 13.0 COCO mask AP, 22.4 versus 18.1 ADE20k PQ, and 25.7 versus 21.1 COCONut PQ. The four-point panoptic improvements are particularly important because they indicate that the method improves representations requiring both semantic recognition and instance-level spatial organization.

The results support a specific trade-off rather than uniform dominance. JEM is stronger on dense tasks while DINOv2 remains stronger on global classification in the controlled comparisons. This pattern is consistent with the method’s exclusive emphasis on patch representations and its omission of a class-token objective.

Scaling to 7 billion parameters

The main scalability claim is that JEM trains stably at 7 billion parameters without adding a separate refinement stage. ViT-L, ViT-g, and ViT-7B models are trained for 500,000 iterations on the A140M dataset, a curated 140-million-image corpus constructed for comparability with prior work. JEM improves task-averaged performance as model size increases and outperforms the matched DINOv2 baseline at all three scales.

At 7B, JEM uses approximately 1.85 trillion student tokens, compared with approximately 4.06 trillion for DINOv3. It also uses the A140M dataset rather than DINOv3’s LVD-1689M dataset and does not use DINOv3’s Gram-anchoring or high-resolution refinement stages. Nevertheless, the reported dense-task results are strong:

Model Parameters Training tokens ImageNet-1k ADE20k mIoU VOC mIoU Cityscapes mIoU COCO mask AP ADE20k PQ COCONut PQ
DINOv3 7B 4.06T 89.0 55.9 86.7 75.2 16.0 21.6 27.0
JEM 7B 1.85T 86.9 54.4 87.1 73.6 15.6 23.9 27.3

The most consequential comparisons are panoptic segmentation: JEM reaches 23.9 PQ on ADE20k, compared with 21.6 for DINOv3, and 27.3 PQ on COCONut, compared with 27.0. It also exceeds DINOv3 on VOC semantic segmentation, 87.1 versus 86.7, and KITTI depth estimation, RMSE 2.125 versus 2.242. DINOv3 remains clearly superior on ImageNet classification, ADE20k semantic segmentation, COCO instance segmentation, and NYUv2 depth. Accordingly, the paper’s claim is not that JEM dominates every benchmark, but that a simpler patch-only objective can achieve competitive or superior dense representations under substantially less data and compute.

The comparison with DINOv3 should nevertheless be interpreted carefully. Although the training data and refinement stages favor DINOv3 as a more comprehensive foundation-model recipe, the models are not identical in architecture, data composition, or optimization history. The results establish strong empirical competitiveness under the reported protocols, but they do not isolate the contribution of the objective from every difference in data and training configuration.

Limitations and open questions

The multi-view assumption is the principal conceptual limitation. The method assumes that each view retains sufficient task-relevant information and that the intersection of view information is the appropriate target for representation learning. Strong cropping, masking, or photometric transformations can violate this condition for tasks requiring view-specific details, and the paper does not provide a formal characterization of when the assumption is valid.

The structural regularizer is also a proxy rather than a direct optimization of the derived objective. Total correlation over a large collection of discrete patch variables is intractable, so JEM matches pairwise similarity structures from selected teacher layers. This depends on the assumption that early and intermediate layers encode useful spatial dependencies and that those dependencies should be preserved in the final representation. The experiments support this proxy empirically, especially at large scale, but do not establish that it controls total correlation in a quantitatively identifiable way.

The reported evaluation is based on frozen-backbone probes, and several probes include hyperparameter searches and task-specific constructions. These protocols are appropriate for measuring representation quality, but they do not determine how the features perform after end-to-end fine-tuning or in settings outside the evaluated image datasets. The paper also leaves open whether the structural regularizer remains necessary under other architectures, patch sizes, resolutions, or training corpora.

Finally, the method retains substantial engineering choices despite its principled derivation: view overlap, patch-grid alignment, independent augmentations, categorical target dimensionality, regularizer schedules, EMA momentum, and layer-selection policies all matter. The claim that multi-objective SSL is unnecessary should therefore be understood narrowly. JEM replaces several semantically distinct losses with one derived objective plus two explicitly motivated regularizers; it does not eliminate the need for careful data augmentation, optimization schedules, and architectural inductive biases.

Conclusion

“Scalable Patch-Level Self-Supervised Learning” (2610.10013) presents JEM as a patch-only SSL framework derived from a variational multi-view InfoMax objective. Its three-part design—cross-view patch alignment, per-patch mutual-information regularization, and structural similarity preservation—targets the principal collapse modes of dense joint-embedding learning. The empirical results show a consistent advantage over controlled DINOv2 baselines on semantic, instance, and panoptic segmentation, while maintaining competitive classification and depth performance. The most notable result is stable training at 7 billion parameters with 12 times less data than DINOv3 and higher reported performance on several dense benchmarks. The remaining technical question is how broadly the proposed information-theoretic decomposition and structural proxy transfer across view distributions, architectures, and downstream adaptation regimes.

Whiteboard

Explain it Like I'm 14

1. What is this paper about?

This paper introduces a new way for computers to learn useful information from images without being given human labels.

Usually, teaching an AI to recognize objects requires many labeled examples, such as pictures marked “dog,” “car,” or “tree.” In self-supervised learning, the computer creates a learning task for itself. This allows it to study huge collections of unlabeled images.

The paper presents a method called JEM, short for Joint-Embedding Multi-view learning. JEM is designed to learn useful information from small parts of images, called patches, while also understanding the whole image.

The main goal is to create a method that is:

  • accurate,
  • stable when made very large,
  • simpler and more logically designed than some existing methods.

2. What questions are the researchers asking?

The researchers focus on several main questions:

  1. Can an AI learn useful visual features by comparing different versions of the same image? For example, can it recognize that two cropped, recolored, or partly hidden versions still show the same object?
  2. Can the AI learn from image patches instead of only looking at the entire image? This is important for tasks such as finding each object in a picture or understanding where objects are located.
  3. Can one carefully designed learning goal replace several complicated goals? Earlier systems often combine many different training tricks. The researchers ask whether a simpler, more principled method can work just as well or better.
  4. Does the method continue to work when the AI becomes extremely large? They test models ranging from about 300 million to 7 billion adjustable numbers, called parameters.
  5. Can the method avoid “collapse”? Collapse happens when the AI produces nearly the same representation for every image or every patch. If this happens, the AI has not learned anything useful.

3. How does the method work?

Learning from different views

The researchers make several views of each image. A view is a changed version of the original picture. It may be:

  • cropped,
  • resized,
  • given different colors or lighting,
  • partly covered by masks,
  • smaller or larger than another view.

Imagine looking at a cat through two different windows. One window shows the cat’s head, and another shows most of its body. Even though the views are different, some information is shared: both views are about the same cat.

The basic idea is:

Information that appears in several views is probably important.

The AI should learn this shared information while ignoring unimportant changes, such as a change in color or the exact crop.

A student and a teacher

JEM uses two copies of the image-processing model:

  • The student looks at one view and tries to make useful predictions.
  • The teacher looks at another view and provides a target for the student.

The teacher is not given human answers. Instead, it produces its own predictions. The teacher changes slowly by following an average of the student’s earlier versions. This is called an exponential moving average, or EMA.

A simple analogy is a student learning from a slightly older, more experienced version of themselves.

Comparing corresponding patches

Images are divided into small squares, or patches, much like cutting a photograph into tiles.

When two views overlap, the method identifies patches that refer to the same part of the original image. The student must produce a representation similar to the teacher’s representation for those matching patches.

For example, if one view contains a patch showing a dog’s ear, and another view also contains that same area, the student and teacher should describe those patches similarly.

This is called patch-level alignment.

Three important training goals

JEM uses three main parts in its training objective.

1. Alignment

The student tries to match the teacher’s predictions for corresponding patches.

This teaches the model that different views of the same image should have related meanings.

2. Global anti-collapse protection

The model is encouraged to use a variety of outputs instead of giving the same answer for every patch.

For example, if the model always predicted “category 1,” no matter what it saw, it would technically make consistent predictions but would learn nothing. The anti-collapse term prevents this by encouraging the model’s outputs to contain information about the input.

3. Structural protection

The model is also encouraged to preserve relationships between patches.

Suppose two patches are close together or have similar visual content. Their representations should reflect this relationship. This helps the AI keep spatial information instead of turning every patch into an almost identical feature.

This is especially important for tasks such as:

  • identifying each object separately,
  • outlining objects,
  • understanding the layout of a scene.

The information-theory idea

The paper uses a concept called mutual information. In simple terms, mutual information measures how much knowing one thing tells us about another thing.

Here, the researchers want the model’s representation to contain information that is shared between different views of the same image.

They also use entropy, which can be thought of as a measure of uncertainty or variety. The model should have enough variety in its predictions to avoid collapse, while still making predictions that are clear and useful.

The equations in the paper turn these ideas into specific training losses. In everyday language, the equations say:

Make matching patches agree, keep the outputs informative, and preserve the relationships between different patches.

4. How did the researchers test JEM?

The researchers used Vision Transformers, or ViTs. A Vision Transformer is a type of neural network that studies an image as a collection of patches and learns how those patches relate to one another.

They tested three model sizes:

  • about 300 million parameters,
  • about 1 billion parameters,
  • about 7 billion parameters.

The models were trained on image datasets, including ImageNet and a much larger collection containing about 140 million images.

After training, the main model was kept fixed. The researchers then tested whether its learned features could help with several tasks. This is called a frozen-backbone evaluation: the main visual system is frozen, and only a small extra system is trained for the new task.

They tested:

Task What it means
Image classification Choosing the main object or category in an image
Semantic segmentation Giving every pixel a category, such as “road” or “person”
Instance segmentation Separating individual objects of the same type
Panoptic segmentation Combining object separation with pixel-level categories
Depth estimation Predicting how far away parts of an image are

They also performed ablation studies. In an ablation study, researchers remove one part of a method at a time to see whether it is important. This is similar to removing one ingredient from a recipe and checking whether the result still tastes good.

5. What did the researchers find?

JEM performed especially well on detailed visual tasks

JEM often performed better than DINOv2, an influential self-supervised learning method, on tasks that require understanding image details and locations.

These tasks included:

  • semantic segmentation,
  • instance segmentation,
  • panoptic segmentation.

This suggests that JEM learns strong information about individual patches and their positions.

On image classification, JEM was usually close to DINOv2 but was sometimes slightly worse. This is still a strong result because DINOv2 was specifically designed to perform very well on whole-image classification.

JEM remained stable at very large sizes

Many self-supervised systems become difficult to train when they grow larger. They may become unstable or lose useful spatial information.

JEM trained successfully from approximately 300 million parameters up to 7 billion parameters. According to the paper, this makes it the first method focused only on patch-level learning in a hidden feature space to demonstrate successful training at this scale.

The structural regularizer became more important in large models

For smaller models, the structural protection term made only a small difference. For larger models, it helped prevent different patches from becoming too similar.

Without this protection, the model gradually lost information about where patches were located and how they differed. With it, the model kept more useful spatial detail.

The anti-collapse term was essential

Removing the global anti-collapse term caused the model to fail badly. Its accuracy dropped dramatically because many outputs became almost identical.

This shows that simply making student and teacher predictions agree is not enough. The system must also be encouraged to produce varied and informative representations.

Some common techniques were not strictly necessary

The researchers found that JEM could still work well even when certain features were removed, such as:

  • masking some patches,
  • using smaller local crops,
  • shifting the student and teacher crops.

However, using different visual changes for each view, especially color and lighting changes, was very important. These changes stopped the model from relying on unhelpful shortcuts, such as memorizing exact pixel colors.

Comparison with DINOv3

At the 7-billion-parameter scale, JEM performed better than DINOv3 on some panoptic segmentation measurements, even though DINOv3:

  • used about 12 times more training data,
  • trained for longer,
  • used extra refinement stages.

This is an impressive result, although comparisons between models trained with different data and procedures should be interpreted carefully.

6. Why are these results important?

The results suggest that self-supervised learning does not always need many separate objectives and complicated training tricks.

JEM uses one central idea—matching information across different views—and adds two protections:

  1. stop the model from producing identical outputs;
  2. preserve the structure and differences between image patches.

This relatively unified design works well for both small and very large models.

The strong segmentation results are particularly important. Classification only asks, “What is in this picture?” Segmentation asks much more:

  • Where is each object?
  • Which pixels belong to it?
  • Are there several separate objects?
  • How are the objects arranged?

Because JEM learns useful patch-level information, it performs well on these more detailed questions.

7. What could this research lead to?

The method could help create general-purpose computer vision systems that learn from large collections of unlabeled images. Such systems might later be used in:

  • self-driving vehicles,
  • medical image analysis,
  • robots,
  • image and video search,
  • maps and satellite images,
  • augmented reality.

A major possible benefit is reducing the need for humans to label millions of images by hand. The AI can first learn general visual knowledge without labels, and then be adapted to a specific task using only a smaller labeled dataset.

The paper also suggests that future AI systems may be made simpler and more principled without losing performance. Instead of adding many unrelated training tricks, researchers may be able to build systems around a few clear ideas.

However, JEM is still a large and expensive model to train. It also depends on the “multiple views share the important information” assumption. That assumption may not always be true. For example, cropping or masking could remove information that is essential for a particular task.

Overall, the paper’s main message is:

A carefully designed system can learn powerful visual knowledge by comparing different views of images, focusing on image patches, avoiding collapse, and preserving spatial structure—even when the model becomes extremely large.

Knowledge Gaps

Knowledge gaps, limitations, and open questions

  • The paper does not empirically test whether the multi-view assumption holds across downstream tasks; views may omit task-relevant information, particularly for fine-grained recognition, text reading, rare objects, or geometric reasoning.
  • The relationship between the proposed mutual-information objective and actual downstream task information remains unverified; high estimated I(XT;T)I(X_T;T) may not imply that the representation preserves labels or other task-relevant factors.
  • The variational lower bound is not measured directly during training, and the paper does not quantify the gap between the theoretical objective I(R;T)I(R;T) and the implemented losses.
  • The anti-collapse derivation assumes factorization of p(T∣XT)p(T\mid X_T) and uses TC(T∣XT)=0\mathrm{TC}(T\mid X_T)=0, but the validity and practical consequences of this assumption are not investigated.
  • The global regularizer estimates mutual information from minibatch class frequencies, leaving unresolved how sensitive training is to batch size, class imbalance, batch composition, and distributed-batch statistics.
  • The discrete target space is fixed to K=4096K=4096 classes; the effect of the number of categories, alternative target parameterizations, and continuous targets is not systematically studied.
  • The structural regularizer is only a proxy for controlling TC(T)\mathrm{TC}(T); the paper does not estimate total correlation or establish theoretically that matching teacher-layer similarities reduces target redundancy.
  • The assumption that different teacher layers encode complementary structural information is not validated; layer choice, number of layers, and the learned transformations gkg_k are not comprehensively ablated.
  • The structural loss is computationally quadratic in the number of patch elements, but its memory and runtime scaling, pair-sampling alternatives, and behavior at higher image resolutions are not reported.
  • The paper does not determine why the structural regularizer becomes important at larger model sizes or predict the model scale at which it, or additional regularization, becomes necessary.
  • The experiments use only ViT architectures with 16×1616\times16 patches and learned register tokens, so the method’s effectiveness with other architectures, patch sizes, tokenization schemes, or hierarchical backbones remains unknown.
  • The claim of stable scaling is tested only up to 7B parameters and 500k iterations; stability over substantially larger models, longer training, higher resolutions, or different compute/data regimes is unresolved.
  • The training-data analysis is limited primarily to ImageNet-1k, ImageNet-22k, and A140M; robustness to substantially different domains, web-scale noise, domain shifts, and non-natural imagery is not established.
  • Comparisons involving DINOv3, LingBot-Vision, and other models are not fully controlled because datasets, training duration, resolution, refinement stages, and data quality differ across methods.
  • The method is evaluated mainly with frozen-backbone linear or attention probes; its transfer under full fine-tuning, few-shot learning, retrieval, detection, tracking, video understanding, and multimodal tasks is not examined.
  • Classification performance remains below DINOv2 and DINOv3 in several comparisons, but the paper does not investigate whether patch-only training inherently limits global semantic representations or whether a patch-level mechanism could close this gap.
  • The benchmarks are dominated by standard vision datasets and do not test robustness to distribution shift, corruptions, adversarial perturbations, long-tail categories, or out-of-distribution detection.
  • The paper does not analyze which visual information is discarded by crop, mask, and photometric invariances, including color, texture, viewpoint, scale, illumination, and spatial relationships.
  • Independent photometric augmentation is shown to be crucial, but the paper does not identify which individual augmentations are responsible or whether the learned invariances are beneficial across different downstream domains.
  • The view-construction study covers selected ablations but does not systematically explore overlap ratios, crop scales, mask ratios, mask patterns, local-view counts, or the number of teacher and student views.
  • The effect of the EMA teacher is evaluated only through a limited ablation; the interaction among teacher momentum, target temperature, optimization schedule, and collapse dynamics is not characterized.
  • The reported collapse indicators—per-patch mutual information and distance from the mean patch—are proxies whose relationship to representation quality, spatial locality, and downstream performance is not quantitatively validated.
  • The paper does not provide a detailed analysis of the learned features’ semantic content, spatial correspondence quality, object boundaries, or sensitivity to patch alignment beyond aggregate benchmark scores.
  • Reproducibility is potentially limited because several implementation details, data-curation procedures, training configurations, and exact compute budgets are delegated to appendices or inherited from prior systems.
  • The environmental and computational costs of training 7B-parameter models are not reported, leaving the method’s practical efficiency relative to multi-objective alternatives unclear.
  • The method’s behavior with very small datasets, low-resource settings, noisy labels or metadata, and limited computational budgets is not evaluated.
  • The paper does not test whether the method can learn useful representations without masking, local crops, or shifted views simultaneously, despite suggesting that these components are individually nonessential.
  • The theoretical framework does not explain why visible-patch alignment is especially important, although its removal causes a large segmentation degradation; the mechanism behind this effect remains an open question.
  • It remains unclear whether the method’s gains on dense prediction arise from the principled objective itself, the absence of a class token, the view construction, the structural regularizer, or inherited architectural and optimization choices.
  • The paper does not investigate whether the proposed objective can be extended beyond images to video, 3D data, audio-visual inputs, or other modalities where view correspondence and spatial structure differ.

Practical Applications

Immediate Applications

  • General-purpose visual backbone for computer-vision products (software, media, retail, security) JEM can be used as a pretrained Vision Transformer backbone for image classification, semantic segmentation, instance segmentation, panoptic segmentation, and monocular depth estimation. Its patch-level representations are particularly suitable for systems that need object boundaries and spatial detail rather than only image-level labels. Potential products/workflows: reusable embedding APIs, visual search, catalog tagging, image moderation, defect detection, scene parsing, and content-organization pipelines. Dependencies: access to trained JEM checkpoints or sufficient compute to pretrain them; task-specific fine-tuning or probing; validation on the target domain, since the reported results are primarily based on ImageNet-derived data.
  • Label-efficient industrial inspection (manufacturing, logistics, agriculture) A frozen or lightly adapted JEM encoder could provide dense feature maps for detecting surface defects, missing components, damaged packaging, crop disease symptoms, or irregular product geometry. The self-supervised pretraining reduces dependence on manually pixel-labeled datasets. Potential workflow: pretrain or reuse JEM, extract patch embeddings from production images, train a small segmentation or anomaly-detection head, and deploy it at inspection stations. Dependencies: domain shift from natural images to factory, agricultural, or warehouse imagery; calibration of false positives and false negatives; adequate image resolution and representative unlabeled data.
  • Robotics and autonomous systems perception (robotics, autonomous vehicles, drones) JEM’s strong dense and panoptic features can support object localization, scene understanding, obstacle segmentation, and depth estimation for mobile robots, warehouse robots, and autonomous vehicles. Patch-level spatial preservation is useful for identifying navigable areas and separating instances of nearby objects. Potential tools: perception modules, bird’s-eye-view preprocessing, obstacle-mapping systems, robotic grasp-region proposals, and camera-based navigation pipelines. Dependencies: the paper evaluates still images rather than complete robotic or video-control loops; temporal consistency, latency, sensor fusion, adverse-weather performance, and safety certification require additional testing.
  • Medical-image representation learning (healthcare) JEM can serve as an initialization for segmentation and localization of anatomical structures, lesions, instruments, or tissue regions in radiology, pathology, dermatology, and endoscopy. Its dense features may be useful where pixel- or region-level outputs matter more than a single diagnostic label. Potential workflow: pretrain on institution-specific unlabeled images, freeze most of the encoder, and fine-tune a small task head with limited expert annotations. Dependencies: medical images differ substantially from ImageNet photographs; modality-specific augmentations must preserve clinically relevant information; patient privacy, regulatory approval, clinical validation, and robustness across hospitals are essential. The paper does not establish diagnostic accuracy.
  • Geospatial and environmental image analysis (climate, agriculture, public infrastructure) Dense JEM features could support land-cover segmentation, road and building extraction, flood or wildfire mapping, crop monitoring, and aerial-image change detection. The method’s ability to learn from unlabeled images is valuable where annotation by experts is expensive. Dependencies: satellite and aerial imagery may require multispectral, temporal, or very high-resolution adaptations; the multi-view assumption must not remove information needed for tasks such as seasonal or spectral analysis.
  • Data annotation and labeling assistance (software, research operations) JEM embeddings can be used to initialize interactive segmentation, cluster visually similar images, identify duplicate or near-duplicate samples, and prioritize examples for human labeling. Patch-level features can help annotators refine object boundaries and discover visually coherent groups. Potential tools: active-learning dashboards, image-dataset browsers, annotation pre-labelers, and embedding-based search systems. Dependencies: similarity in the learned feature space may reflect dataset bias rather than semantic equivalence; human review remains necessary for high-stakes or ambiguous cases.
  • Academic benchmarking and reproducible SSL research (academia) The paper provides a comparatively simple, information-theoretically motivated alternative to multi-objective methods such as DINOv2. Researchers can use JEM to study patch-level representation quality, collapse prevention, view construction, scaling laws, and transfer across dense prediction tasks. Potential outputs: open pretrained checkpoints, standardized dense-feature benchmarks, and ablation studies comparing alignment, global-collapse, and structural-collapse regularization. Dependencies: reproducibility depends on implementation details, large-batch training, data curation, augmentation choices, and substantial GPU resources, especially at 1B–7B scale.
  • More efficient domain adaptation workflows (enterprise AI, cloud ML) Because JEM reportedly performs strongly without the additional refinement stages used by some competing systems, organizations could use it as a foundation for a simpler transfer-learning pipeline: frozen backbone, lightweight probe, and selective fine-tuning. Dependencies: the reported efficiency advantage is not a complete cost-of-ownership analysis; inference memory, training energy, checkpoint availability, and performance on proprietary data must be measured directly.
  • Everyday image organization and accessibility tools (consumer software) JEM-based local models could improve photo search, automatic object or scene labeling, background separation, image cropping, and visual similarity search. Dense features may also support assistive descriptions that identify regions or objects in an image. Dependencies: consumer deployment may require distillation or quantization because a 300M-parameter model or larger is resource-intensive; privacy-preserving on-device inference and bias testing are necessary.

Long-Term Applications

  • Foundation models for video understanding and embodied AI (robotics, autonomous driving, video analytics) JEM’s multi-view framework could be extended from spatial image patches to spatiotemporal tokens, enabling self-supervised learning of object persistence, motion-aware segmentation, depth, and scene dynamics. This could support robots that learn from large collections of unlabeled video. Dependencies: temporal view construction, motion-specific collapse modes, long-range correspondence, compute and memory requirements, and evaluation in closed-loop interaction remain unresolved.
  • Multimodal and cross-sensor representation learning (healthcare, robotics, remote sensing) The alignment principle could be applied to paired views from RGB, depth, thermal, LiDAR, radar, ultrasound, or text. Shared representations could support sensor substitution, cross-modal retrieval, and robust perception when one sensor is degraded. Dependencies: the modalities must contain sufficiently shared task-relevant information; imperfect synchronization and missing modalities complicate patch correspondence; aligning modalities indiscriminately could suppress modality-specific signals that are operationally important.
  • Large-scale 3D and spatial foundation models (construction, mapping, robotics, digital twins) Patch-level structural regularization could be generalized to point clouds, multi-view 3D reconstruction, or neural rendering. A resulting model might provide reusable features for indoor mapping, infrastructure inspection, object pose estimation, and digital-twin construction. Dependencies: 2D patch correspondence does not directly transfer to irregular 3D tokens; geometric invariances, occlusion handling, coordinate systems, and large 3D datasets require new methods.
  • Clinical segmentation and decision-support foundation models (healthcare) With suitable medical pretraining, a JEM-like model could become a general-purpose encoder for organ segmentation, tumor delineation, surgical video understanding, and longitudinal disease monitoring. Its dense representations could reduce annotation requirements across related clinical tasks. Dependencies: prospective clinical studies, expert agreement standards, explainability, fairness across populations and institutions, privacy-preserving training, and regulatory clearance. Representation quality on benchmark segmentation datasets alone would not demonstrate clinical utility.
  • Low-label learning for underrepresented languages, regions, and domains (public policy, education, global development) Governments, universities, and NGOs could pretrain JEM-like models on locally collected unlabeled imagery for disaster response, infrastructure inventories, biodiversity monitoring, or educational-content indexing where labeled datasets are scarce. Dependencies: data governance, consent, community participation, geographic bias, and the validity of local multi-view augmentations. Cropping or photometric transformations may destroy information important for culturally or geographically specific tasks.
  • Adaptive and continual visual learning (industrial AI, edge computing) The alignment and anti-collapse formulation could support continual updating from new unlabeled images while preserving spatially useful features. This might enable inspection systems to adapt to new products, cameras, lighting conditions, or seasonal environments without full relabeling. Dependencies: catastrophic forgetting, stability of the EMA teacher under distribution shifts, monitoring for representation drift, and safeguards against learning corrupted or adversarial data.
  • Energy- and compute-efficient visual foundation models (cloud infrastructure, edge AI, sustainability) Since JEM achieves strong dense performance with fewer reported training tokens than the compared DINOv3 setup, future work could investigate whether its simpler objective reduces training cost, data requirements, or refinement stages. Distilled JEM models could provide dense perception on mobile and embedded hardware. Dependencies: the paper does not provide a full energy, wall-clock, or hardware-cost comparison; large ViT models remain expensive, and compression may reduce the spatial fidelity that motivates JEM.
  • Policy and standards for self-supervised visual AI (public policy, AI governance) JEM offers a concrete case for evaluating foundation models beyond classification accuracy, emphasizing segmentation, depth, spatial locality, collapse behavior, and scaling stability. These metrics could inform procurement standards for public-sector computer vision and safety-critical perception systems. Dependencies: benchmarks must become more representative of real-world populations and environments; policy use requires reporting dataset provenance, demographic performance, uncertainty, robustness, and environmental cost.
  • Human-computer interaction based on spatial visual embeddings (education, accessibility, productivity software) Future applications could use JEM-like patch representations for region-aware image editors, visual question answering interfaces, educational object identification, augmented-reality assistance, and tools that let users search or manipulate image regions by meaning. Dependencies: JEM alone is not a language or reasoning model; these systems require multimodal interfaces, grounding mechanisms, user studies, accessibility validation, and protection against incorrect or misleading visual interpretations.

Glossary

  • Anti-collapse term: A component of an objective designed to prevent representations from becoming constant or uninformative. “an anti-collapse term that maintains view-specific information”
  • Attention probe: A lightweight evaluation method that uses attention-based readouts from a frozen representation. “we report attention probe accuracy on ImageNet-1k”
  • Categorical distribution: A probability distribution over a finite set of discrete classes. “We use K=4096K=4096 classes for the categorical distributions.”
  • Class token: A special transformer token intended to summarize an entire input, often used for classification. “our model has no class token”
  • Conditional entropy: The uncertainty of one random variable given another. “maintaining targets with high marginal entropy H(T)H(T) and low conditional entropy H(T∣XT)H(T \mid X_T)”
  • Contrastive learning: A representation-learning approach that brings related examples together and separates unrelated examples. “Contrastive methods \citep{he2020moco,chen2020simple} pull together representations of different views of the same image while pushing apart representations of other images.”
  • Cosine similarity: A measure of similarity based on the angle between two vectors. “where cos⁡\operatorname{cos} denotes the cosine similarity”
  • Data processing inequality: An information-theoretic result stating that processing data cannot increase its mutual information with another variable. “By the data processing inequality, we have I(XS;XT)≥I(XS;T)≥I(R;T)I(X_S; X_T) \geq I(X_S; T) \geq I(R; T).”
  • Dense prediction: A task requiring an output for many or all spatial locations in an input, such as segmentation. “strong performance on both global and dense probing tasks”
  • Discrete representation target: A target representation whose values come from a finite or countable set. “our use of discrete representation targets can be seen as an architectural bottleneck”
  • Exponential moving average (EMA): A weighted running average that gives greater weight to recent values, commonly used to update teacher models. “the teacher an EMA of the student”
  • Foundation model: A large pretrained model intended to support many downstream tasks. “DINOv3 7B foundation model”
  • Global collapse: A failure mode in which representations become constant or lose information across inputs. “preventing global collapse of the targets”
  • Gram anchoring: A regularization strategy that preserves relationships among feature vectors through their Gram matrix. “DINOv3 even adds a further term, Gram anchoring”
  • High-resolution fine-tuning: Additional training on inputs with greater spatial resolution to improve visual detail. “gram and high-resolution fine-tuning”
  • Image-level objective: A learning objective applied to a representation of an entire image rather than to local regions. “Unlike methods with separate global and local objectives”
  • Information bottleneck: A principle that preserves information predictive of a target while discarding irrelevant information. “the information bottleneck principle suggests retaining predictive information while compressing superfluous information”
  • Information collapse: The loss of informative variation in learned representations or targets. “one preventing information collapse”
  • Information-theoretic objective: An optimization objective formulated using quantities from information theory. “we construct an information-theoretic objective decomposing into interpretable terms.”
  • Inductive bias: A preference introduced by a model architecture, training procedure, or data construction that influences what it learns. “learning good representations also relies on inductive biases from architecture and view construction.”
  • Instance segmentation: A vision task that identifies and delineates each individual object instance. “mask AP for instance segmentation on COCO”
  • Joint embedding: A representation space in which inputs from different views or modalities are mapped for comparison. “a novel joint-embedding multi-view method”
  • Kullback–Leibler divergence: A measure of how one probability distribution differs from another. “where H(T)H(T) is the entropy of the target and istheKullback−Leiblerdivergence”</li><li><strong>Latent−spacemethod</strong>:Amethodthatpredictsoralignslearnedrepresentationsratherthanrawinputdata.“thefirstexclusivelypatch−levellatent−spacemethod”</li><li><strong>Marginalentropy</strong>:Theentropyofavariableconsideredwithoutconditioningonanothervariable.“marg.entropy”</li><li><strong>Maskedprediction</strong>:Aself−supervisedtaskinwhichhiddenpartsofaninputarepredictedfromthevisibleparts.“Maskedpredictioninsteadremovespartoftheinputandpredictsthemissinginformation”</li><li><strong>Meanintersectionoverunion(mIoU)</strong>:Theaverageintersection−over−unionscoreusedtoevaluatesegmentationquality.“segmentationmIoUonADE20k”</li><li><strong>Mutualinformation</strong>:Ameasureoftheamountofinformationsharedbytworandomvariables.“Theinformationwewouldliketocaptureintherepresentation is the Kullback-Leibler divergence”</li> <li><strong>Latent-space method</strong>: A method that predicts or aligns learned representations rather than raw input data. “the first exclusively patch-level latent-space method”</li> <li><strong>Marginal entropy</strong>: The entropy of a variable considered without conditioning on another variable. “marg. entropy”</li> <li><strong>Masked prediction</strong>: A self-supervised task in which hidden parts of an input are predicted from the visible parts. “Masked prediction instead removes part of the input and predicts the missing information”</li> <li><strong>Mean intersection over union (mIoU)</strong>: The average intersection-over-union score used to evaluate segmentation quality. “segmentation mIoU on ADE20k”</li> <li><strong>Mutual information</strong>: A measure of the amount of information shared by two random variables. “The information we would like to capture in the representation Risthestudent−teachersharedmutualinformation(MI) is the student-teacher shared mutual information (MI) I(X_S; X_T)$.”
  • Panoptic segmentation: A vision task combining semantic segmentation and instance segmentation. “PQ for panoptic segmentation on COCONut”
  • Patch-level objective: A training objective applied to local image patches rather than only to a whole-image representation. “our model has no class token and applies the objective exclusively to patch representations”
  • Photometric transformation: An image transformation that changes appearance-related properties such as color, brightness, or contrast. “Student and teacher views are also independently transformed with a standard set of photometric transformations”
  • Projection head: A learned mapping that transforms an encoder representation into a prediction or target space. “a projection head maps the representation to a predictive distribution”
  • Redundancy reduction: A learning strategy that discourages different representation dimensions or elements from encoding the same information. “Redundancy-reduction methods”
  • Representation collapse: A degeneracy in which different inputs receive identical or nearly identical representations. “the representation avoids the collapse to a single, redundant vector”
  • Self-clustering: A self-supervised approach that learns cluster assignments and makes them consistent across views. “Self-clustering approaches such as SwAV~\cite{caron2021swav} and DINO~\citep{caron2021emerging} learn cluster assignments and align those across views.”
  • Self-distillation: A training method in which a student learns to match predictions produced by a teacher model without external labels. “global DINO self-distillation”
  • Semantic segmentation: A task that assigns a semantic class label to each image pixel. “For semantic segmentation, we report linear probe mIoU”
  • Spatial locality: The property that nearby spatial positions retain locally meaningful and distinct information. “dense feature maps progressively lose spatial locality.”
  • Structural collapse: A failure mode in which different spatial elements encode redundant or overly similar information. “which we term a structural collapse”
  • Student–teacher paradigm: A framework in which a student network learns to match targets generated by a teacher network. “the variational approximation which naturally fits the student-teacher paradigm”
  • Total correlation: A multivariate measure of redundancy among several random variables. “the total correlation $\TC(T)$ measures the redundancy among elements”
  • Variational approximation: An approximation that uses an optimized auxiliary distribution to estimate an otherwise difficult probabilistic quantity. “we use a variational approximation which naturally fits the student-teacher paradigm”
  • Variational lower bound: A tractable quantity that provides a lower bound on an objective such as mutual information. “We can maximize this objective using a variational lower bound”
  • View invariance: The property that representations remain similar when an input is altered through transformations or different viewpoints. “enforce view invariance while explicitly preventing collapse across dimensions.”
  • Vision Transformer (ViT): A transformer architecture that processes an image as a sequence of patch tokens. “Our encoder is a Vision Transformer (ViT)”
  • Visual representation: A learned numerical encoding of visual input used by downstream tasks. “Self-supervised learning (SSL) at scale produces powerful visual representations.”

Open Problems

We found no open problems mentioned in this paper.