Scalable Patch-Level Self-Supervised Learning
Abstract: Self-supervised learning (SSL) at scale produces powerful visual representations. However, most scalable SSL methods rely on ad hoc combinations of multiple objectives and stabilization mechanisms. Taking a step back, we ask if we can design a high-performing, yet principled SSL algorithm. Starting from the multi-view assumption, stipulating that task-relevant content is captured by the information common to different views, we construct an information-theoretic objective decomposing into interpretable terms. This derivation yields JEM, a student-teacher method that learns by aligning corresponding patch representations across views, explicitly regularized by information and structure preservation losses. JEM trains stably from 300M to 7B parameters, and, to our knowledge, is the first latent-space patch-level method demonstrated at 7B scale. Across all scales, JEM reaches strong performance on both global and dense probing tasks, on segmentation benchmarks consistently surpassing the DINOv2 algorithm, an influential foundation for today's strongest visual SSL methods. Notably, at 7B parameters, it exceeds the performance of DINOv3 on panoptic segmentation, despite being trained on less data without refinement stages. These results demonstrate that we can indeed design an SSL algorithm that learns strong representations, is principled and stable.
Paper Prompts
Sign up for free to create and run prompts on this paper.
Top Community Prompts
Explain it Like I'm 14
1. What is this paper about?
This paper introduces a new way for computers to learn useful information from images without being given human labels.
Usually, teaching an AI to recognize objects requires many labeled examples, such as pictures marked “dog,” “car,” or “tree.” In self-supervised learning, the computer creates a learning task for itself. This allows it to study huge collections of unlabeled images.
The paper presents a method called JEM, short for Joint-Embedding Multi-view learning. JEM is designed to learn useful information from small parts of images, called patches, while also understanding the whole image.
The main goal is to create a method that is:
- accurate,
- stable when made very large,
- simpler and more logically designed than some existing methods.
2. What questions are the researchers asking?
The researchers focus on several main questions:
- Can an AI learn useful visual features by comparing different versions of the same image? For example, can it recognize that two cropped, recolored, or partly hidden versions still show the same object?
- Can the AI learn from image patches instead of only looking at the entire image? This is important for tasks such as finding each object in a picture or understanding where objects are located.
- Can one carefully designed learning goal replace several complicated goals? Earlier systems often combine many different training tricks. The researchers ask whether a simpler, more principled method can work just as well or better.
- Does the method continue to work when the AI becomes extremely large? They test models ranging from about 300 million to 7 billion adjustable numbers, called parameters.
- Can the method avoid “collapse”? Collapse happens when the AI produces nearly the same representation for every image or every patch. If this happens, the AI has not learned anything useful.
3. How does the method work?
Learning from different views
The researchers make several views of each image. A view is a changed version of the original picture. It may be:
- cropped,
- resized,
- given different colors or lighting,
- partly covered by masks,
- smaller or larger than another view.
Imagine looking at a cat through two different windows. One window shows the cat’s head, and another shows most of its body. Even though the views are different, some information is shared: both views are about the same cat.
The basic idea is:
Information that appears in several views is probably important.
The AI should learn this shared information while ignoring unimportant changes, such as a change in color or the exact crop.
A student and a teacher
JEM uses two copies of the image-processing model:
- The student looks at one view and tries to make useful predictions.
- The teacher looks at another view and provides a target for the student.
The teacher is not given human answers. Instead, it produces its own predictions. The teacher changes slowly by following an average of the student’s earlier versions. This is called an exponential moving average, or EMA.
A simple analogy is a student learning from a slightly older, more experienced version of themselves.
Comparing corresponding patches
Images are divided into small squares, or patches, much like cutting a photograph into tiles.
When two views overlap, the method identifies patches that refer to the same part of the original image. The student must produce a representation similar to the teacher’s representation for those matching patches.
For example, if one view contains a patch showing a dog’s ear, and another view also contains that same area, the student and teacher should describe those patches similarly.
This is called patch-level alignment.
Three important training goals
JEM uses three main parts in its training objective.
1. Alignment
The student tries to match the teacher’s predictions for corresponding patches.
This teaches the model that different views of the same image should have related meanings.
2. Global anti-collapse protection
The model is encouraged to use a variety of outputs instead of giving the same answer for every patch.
For example, if the model always predicted “category 1,” no matter what it saw, it would technically make consistent predictions but would learn nothing. The anti-collapse term prevents this by encouraging the model’s outputs to contain information about the input.
3. Structural protection
The model is also encouraged to preserve relationships between patches.
Suppose two patches are close together or have similar visual content. Their representations should reflect this relationship. This helps the AI keep spatial information instead of turning every patch into an almost identical feature.
This is especially important for tasks such as:
- identifying each object separately,
- outlining objects,
- understanding the layout of a scene.
The information-theory idea
The paper uses a concept called mutual information. In simple terms, mutual information measures how much knowing one thing tells us about another thing.
Here, the researchers want the model’s representation to contain information that is shared between different views of the same image.
They also use entropy, which can be thought of as a measure of uncertainty or variety. The model should have enough variety in its predictions to avoid collapse, while still making predictions that are clear and useful.
The equations in the paper turn these ideas into specific training losses. In everyday language, the equations say:
Make matching patches agree, keep the outputs informative, and preserve the relationships between different patches.
4. How did the researchers test JEM?
The researchers used Vision Transformers, or ViTs. A Vision Transformer is a type of neural network that studies an image as a collection of patches and learns how those patches relate to one another.
They tested three model sizes:
- about 300 million parameters,
- about 1 billion parameters,
- about 7 billion parameters.
The models were trained on image datasets, including ImageNet and a much larger collection containing about 140 million images.
After training, the main model was kept fixed. The researchers then tested whether its learned features could help with several tasks. This is called a frozen-backbone evaluation: the main visual system is frozen, and only a small extra system is trained for the new task.
They tested:
| Task | What it means |
|---|---|
| Image classification | Choosing the main object or category in an image |
| Semantic segmentation | Giving every pixel a category, such as “road” or “person” |
| Instance segmentation | Separating individual objects of the same type |
| Panoptic segmentation | Combining object separation with pixel-level categories |
| Depth estimation | Predicting how far away parts of an image are |
They also performed ablation studies. In an ablation study, researchers remove one part of a method at a time to see whether it is important. This is similar to removing one ingredient from a recipe and checking whether the result still tastes good.
5. What did the researchers find?
JEM performed especially well on detailed visual tasks
JEM often performed better than DINOv2, an influential self-supervised learning method, on tasks that require understanding image details and locations.
These tasks included:
- semantic segmentation,
- instance segmentation,
- panoptic segmentation.
This suggests that JEM learns strong information about individual patches and their positions.
On image classification, JEM was usually close to DINOv2 but was sometimes slightly worse. This is still a strong result because DINOv2 was specifically designed to perform very well on whole-image classification.
JEM remained stable at very large sizes
Many self-supervised systems become difficult to train when they grow larger. They may become unstable or lose useful spatial information.
JEM trained successfully from approximately 300 million parameters up to 7 billion parameters. According to the paper, this makes it the first method focused only on patch-level learning in a hidden feature space to demonstrate successful training at this scale.
The structural regularizer became more important in large models
For smaller models, the structural protection term made only a small difference. For larger models, it helped prevent different patches from becoming too similar.
Without this protection, the model gradually lost information about where patches were located and how they differed. With it, the model kept more useful spatial detail.
The anti-collapse term was essential
Removing the global anti-collapse term caused the model to fail badly. Its accuracy dropped dramatically because many outputs became almost identical.
This shows that simply making student and teacher predictions agree is not enough. The system must also be encouraged to produce varied and informative representations.
Some common techniques were not strictly necessary
The researchers found that JEM could still work well even when certain features were removed, such as:
- masking some patches,
- using smaller local crops,
- shifting the student and teacher crops.
However, using different visual changes for each view, especially color and lighting changes, was very important. These changes stopped the model from relying on unhelpful shortcuts, such as memorizing exact pixel colors.
Comparison with DINOv3
At the 7-billion-parameter scale, JEM performed better than DINOv3 on some panoptic segmentation measurements, even though DINOv3:
- used about 12 times more training data,
- trained for longer,
- used extra refinement stages.
This is an impressive result, although comparisons between models trained with different data and procedures should be interpreted carefully.
6. Why are these results important?
The results suggest that self-supervised learning does not always need many separate objectives and complicated training tricks.
JEM uses one central idea—matching information across different views—and adds two protections:
- stop the model from producing identical outputs;
- preserve the structure and differences between image patches.
This relatively unified design works well for both small and very large models.
The strong segmentation results are particularly important. Classification only asks, “What is in this picture?” Segmentation asks much more:
- Where is each object?
- Which pixels belong to it?
- Are there several separate objects?
- How are the objects arranged?
Because JEM learns useful patch-level information, it performs well on these more detailed questions.
7. What could this research lead to?
The method could help create general-purpose computer vision systems that learn from large collections of unlabeled images. Such systems might later be used in:
- self-driving vehicles,
- medical image analysis,
- robots,
- image and video search,
- maps and satellite images,
- augmented reality.
A major possible benefit is reducing the need for humans to label millions of images by hand. The AI can first learn general visual knowledge without labels, and then be adapted to a specific task using only a smaller labeled dataset.
The paper also suggests that future AI systems may be made simpler and more principled without losing performance. Instead of adding many unrelated training tricks, researchers may be able to build systems around a few clear ideas.
However, JEM is still a large and expensive model to train. It also depends on the “multiple views share the important information” assumption. That assumption may not always be true. For example, cropping or masking could remove information that is essential for a particular task.
Overall, the paper’s main message is:
A carefully designed system can learn powerful visual knowledge by comparing different views of images, focusing on image patches, avoiding collapse, and preserving spatial structure—even when the model becomes extremely large.
Knowledge Gaps
Knowledge gaps, limitations, and open questions
- The paper does not empirically test whether the multi-view assumption holds across downstream tasks; views may omit task-relevant information, particularly for fine-grained recognition, text reading, rare objects, or geometric reasoning.
- The relationship between the proposed mutual-information objective and actual downstream task information remains unverified; high estimated may not imply that the representation preserves labels or other task-relevant factors.
- The variational lower bound is not measured directly during training, and the paper does not quantify the gap between the theoretical objective and the implemented losses.
- The anti-collapse derivation assumes factorization of and uses , but the validity and practical consequences of this assumption are not investigated.
- The global regularizer estimates mutual information from minibatch class frequencies, leaving unresolved how sensitive training is to batch size, class imbalance, batch composition, and distributed-batch statistics.
- The discrete target space is fixed to classes; the effect of the number of categories, alternative target parameterizations, and continuous targets is not systematically studied.
- The structural regularizer is only a proxy for controlling ; the paper does not estimate total correlation or establish theoretically that matching teacher-layer similarities reduces target redundancy.
- The assumption that different teacher layers encode complementary structural information is not validated; layer choice, number of layers, and the learned transformations are not comprehensively ablated.
- The structural loss is computationally quadratic in the number of patch elements, but its memory and runtime scaling, pair-sampling alternatives, and behavior at higher image resolutions are not reported.
- The paper does not determine why the structural regularizer becomes important at larger model sizes or predict the model scale at which it, or additional regularization, becomes necessary.
- The experiments use only ViT architectures with patches and learned register tokens, so the method’s effectiveness with other architectures, patch sizes, tokenization schemes, or hierarchical backbones remains unknown.
- The claim of stable scaling is tested only up to 7B parameters and 500k iterations; stability over substantially larger models, longer training, higher resolutions, or different compute/data regimes is unresolved.
- The training-data analysis is limited primarily to ImageNet-1k, ImageNet-22k, and A140M; robustness to substantially different domains, web-scale noise, domain shifts, and non-natural imagery is not established.
- Comparisons involving DINOv3, LingBot-Vision, and other models are not fully controlled because datasets, training duration, resolution, refinement stages, and data quality differ across methods.
- The method is evaluated mainly with frozen-backbone linear or attention probes; its transfer under full fine-tuning, few-shot learning, retrieval, detection, tracking, video understanding, and multimodal tasks is not examined.
- Classification performance remains below DINOv2 and DINOv3 in several comparisons, but the paper does not investigate whether patch-only training inherently limits global semantic representations or whether a patch-level mechanism could close this gap.
- The benchmarks are dominated by standard vision datasets and do not test robustness to distribution shift, corruptions, adversarial perturbations, long-tail categories, or out-of-distribution detection.
- The paper does not analyze which visual information is discarded by crop, mask, and photometric invariances, including color, texture, viewpoint, scale, illumination, and spatial relationships.
- Independent photometric augmentation is shown to be crucial, but the paper does not identify which individual augmentations are responsible or whether the learned invariances are beneficial across different downstream domains.
- The view-construction study covers selected ablations but does not systematically explore overlap ratios, crop scales, mask ratios, mask patterns, local-view counts, or the number of teacher and student views.
- The effect of the EMA teacher is evaluated only through a limited ablation; the interaction among teacher momentum, target temperature, optimization schedule, and collapse dynamics is not characterized.
- The reported collapse indicators—per-patch mutual information and distance from the mean patch—are proxies whose relationship to representation quality, spatial locality, and downstream performance is not quantitatively validated.
- The paper does not provide a detailed analysis of the learned features’ semantic content, spatial correspondence quality, object boundaries, or sensitivity to patch alignment beyond aggregate benchmark scores.
- Reproducibility is potentially limited because several implementation details, data-curation procedures, training configurations, and exact compute budgets are delegated to appendices or inherited from prior systems.
- The environmental and computational costs of training 7B-parameter models are not reported, leaving the method’s practical efficiency relative to multi-objective alternatives unclear.
- The method’s behavior with very small datasets, low-resource settings, noisy labels or metadata, and limited computational budgets is not evaluated.
- The paper does not test whether the method can learn useful representations without masking, local crops, or shifted views simultaneously, despite suggesting that these components are individually nonessential.
- The theoretical framework does not explain why visible-patch alignment is especially important, although its removal causes a large segmentation degradation; the mechanism behind this effect remains an open question.
- It remains unclear whether the method’s gains on dense prediction arise from the principled objective itself, the absence of a class token, the view construction, the structural regularizer, or inherited architectural and optimization choices.
- The paper does not investigate whether the proposed objective can be extended beyond images to video, 3D data, audio-visual inputs, or other modalities where view correspondence and spatial structure differ.
Practical Applications
Immediate Applications
- General-purpose visual backbone for computer-vision products (software, media, retail, security) JEM can be used as a pretrained Vision Transformer backbone for image classification, semantic segmentation, instance segmentation, panoptic segmentation, and monocular depth estimation. Its patch-level representations are particularly suitable for systems that need object boundaries and spatial detail rather than only image-level labels. Potential products/workflows: reusable embedding APIs, visual search, catalog tagging, image moderation, defect detection, scene parsing, and content-organization pipelines. Dependencies: access to trained JEM checkpoints or sufficient compute to pretrain them; task-specific fine-tuning or probing; validation on the target domain, since the reported results are primarily based on ImageNet-derived data.
- Label-efficient industrial inspection (manufacturing, logistics, agriculture) A frozen or lightly adapted JEM encoder could provide dense feature maps for detecting surface defects, missing components, damaged packaging, crop disease symptoms, or irregular product geometry. The self-supervised pretraining reduces dependence on manually pixel-labeled datasets. Potential workflow: pretrain or reuse JEM, extract patch embeddings from production images, train a small segmentation or anomaly-detection head, and deploy it at inspection stations. Dependencies: domain shift from natural images to factory, agricultural, or warehouse imagery; calibration of false positives and false negatives; adequate image resolution and representative unlabeled data.
- Robotics and autonomous systems perception (robotics, autonomous vehicles, drones) JEM’s strong dense and panoptic features can support object localization, scene understanding, obstacle segmentation, and depth estimation for mobile robots, warehouse robots, and autonomous vehicles. Patch-level spatial preservation is useful for identifying navigable areas and separating instances of nearby objects. Potential tools: perception modules, bird’s-eye-view preprocessing, obstacle-mapping systems, robotic grasp-region proposals, and camera-based navigation pipelines. Dependencies: the paper evaluates still images rather than complete robotic or video-control loops; temporal consistency, latency, sensor fusion, adverse-weather performance, and safety certification require additional testing.
- Medical-image representation learning (healthcare) JEM can serve as an initialization for segmentation and localization of anatomical structures, lesions, instruments, or tissue regions in radiology, pathology, dermatology, and endoscopy. Its dense features may be useful where pixel- or region-level outputs matter more than a single diagnostic label. Potential workflow: pretrain on institution-specific unlabeled images, freeze most of the encoder, and fine-tune a small task head with limited expert annotations. Dependencies: medical images differ substantially from ImageNet photographs; modality-specific augmentations must preserve clinically relevant information; patient privacy, regulatory approval, clinical validation, and robustness across hospitals are essential. The paper does not establish diagnostic accuracy.
- Geospatial and environmental image analysis (climate, agriculture, public infrastructure) Dense JEM features could support land-cover segmentation, road and building extraction, flood or wildfire mapping, crop monitoring, and aerial-image change detection. The method’s ability to learn from unlabeled images is valuable where annotation by experts is expensive. Dependencies: satellite and aerial imagery may require multispectral, temporal, or very high-resolution adaptations; the multi-view assumption must not remove information needed for tasks such as seasonal or spectral analysis.
- Data annotation and labeling assistance (software, research operations) JEM embeddings can be used to initialize interactive segmentation, cluster visually similar images, identify duplicate or near-duplicate samples, and prioritize examples for human labeling. Patch-level features can help annotators refine object boundaries and discover visually coherent groups. Potential tools: active-learning dashboards, image-dataset browsers, annotation pre-labelers, and embedding-based search systems. Dependencies: similarity in the learned feature space may reflect dataset bias rather than semantic equivalence; human review remains necessary for high-stakes or ambiguous cases.
- Academic benchmarking and reproducible SSL research (academia) The paper provides a comparatively simple, information-theoretically motivated alternative to multi-objective methods such as DINOv2. Researchers can use JEM to study patch-level representation quality, collapse prevention, view construction, scaling laws, and transfer across dense prediction tasks. Potential outputs: open pretrained checkpoints, standardized dense-feature benchmarks, and ablation studies comparing alignment, global-collapse, and structural-collapse regularization. Dependencies: reproducibility depends on implementation details, large-batch training, data curation, augmentation choices, and substantial GPU resources, especially at 1B–7B scale.
- More efficient domain adaptation workflows (enterprise AI, cloud ML) Because JEM reportedly performs strongly without the additional refinement stages used by some competing systems, organizations could use it as a foundation for a simpler transfer-learning pipeline: frozen backbone, lightweight probe, and selective fine-tuning. Dependencies: the reported efficiency advantage is not a complete cost-of-ownership analysis; inference memory, training energy, checkpoint availability, and performance on proprietary data must be measured directly.
- Everyday image organization and accessibility tools (consumer software) JEM-based local models could improve photo search, automatic object or scene labeling, background separation, image cropping, and visual similarity search. Dense features may also support assistive descriptions that identify regions or objects in an image. Dependencies: consumer deployment may require distillation or quantization because a 300M-parameter model or larger is resource-intensive; privacy-preserving on-device inference and bias testing are necessary.
Long-Term Applications
- Foundation models for video understanding and embodied AI (robotics, autonomous driving, video analytics) JEM’s multi-view framework could be extended from spatial image patches to spatiotemporal tokens, enabling self-supervised learning of object persistence, motion-aware segmentation, depth, and scene dynamics. This could support robots that learn from large collections of unlabeled video. Dependencies: temporal view construction, motion-specific collapse modes, long-range correspondence, compute and memory requirements, and evaluation in closed-loop interaction remain unresolved.
- Multimodal and cross-sensor representation learning (healthcare, robotics, remote sensing) The alignment principle could be applied to paired views from RGB, depth, thermal, LiDAR, radar, ultrasound, or text. Shared representations could support sensor substitution, cross-modal retrieval, and robust perception when one sensor is degraded. Dependencies: the modalities must contain sufficiently shared task-relevant information; imperfect synchronization and missing modalities complicate patch correspondence; aligning modalities indiscriminately could suppress modality-specific signals that are operationally important.
- Large-scale 3D and spatial foundation models (construction, mapping, robotics, digital twins) Patch-level structural regularization could be generalized to point clouds, multi-view 3D reconstruction, or neural rendering. A resulting model might provide reusable features for indoor mapping, infrastructure inspection, object pose estimation, and digital-twin construction. Dependencies: 2D patch correspondence does not directly transfer to irregular 3D tokens; geometric invariances, occlusion handling, coordinate systems, and large 3D datasets require new methods.
- Clinical segmentation and decision-support foundation models (healthcare) With suitable medical pretraining, a JEM-like model could become a general-purpose encoder for organ segmentation, tumor delineation, surgical video understanding, and longitudinal disease monitoring. Its dense representations could reduce annotation requirements across related clinical tasks. Dependencies: prospective clinical studies, expert agreement standards, explainability, fairness across populations and institutions, privacy-preserving training, and regulatory clearance. Representation quality on benchmark segmentation datasets alone would not demonstrate clinical utility.
- Low-label learning for underrepresented languages, regions, and domains (public policy, education, global development) Governments, universities, and NGOs could pretrain JEM-like models on locally collected unlabeled imagery for disaster response, infrastructure inventories, biodiversity monitoring, or educational-content indexing where labeled datasets are scarce. Dependencies: data governance, consent, community participation, geographic bias, and the validity of local multi-view augmentations. Cropping or photometric transformations may destroy information important for culturally or geographically specific tasks.
- Adaptive and continual visual learning (industrial AI, edge computing) The alignment and anti-collapse formulation could support continual updating from new unlabeled images while preserving spatially useful features. This might enable inspection systems to adapt to new products, cameras, lighting conditions, or seasonal environments without full relabeling. Dependencies: catastrophic forgetting, stability of the EMA teacher under distribution shifts, monitoring for representation drift, and safeguards against learning corrupted or adversarial data.
- Energy- and compute-efficient visual foundation models (cloud infrastructure, edge AI, sustainability) Since JEM achieves strong dense performance with fewer reported training tokens than the compared DINOv3 setup, future work could investigate whether its simpler objective reduces training cost, data requirements, or refinement stages. Distilled JEM models could provide dense perception on mobile and embedded hardware. Dependencies: the paper does not provide a full energy, wall-clock, or hardware-cost comparison; large ViT models remain expensive, and compression may reduce the spatial fidelity that motivates JEM.
- Policy and standards for self-supervised visual AI (public policy, AI governance) JEM offers a concrete case for evaluating foundation models beyond classification accuracy, emphasizing segmentation, depth, spatial locality, collapse behavior, and scaling stability. These metrics could inform procurement standards for public-sector computer vision and safety-critical perception systems. Dependencies: benchmarks must become more representative of real-world populations and environments; policy use requires reporting dataset provenance, demographic performance, uncertainty, robustness, and environmental cost.
- Human-computer interaction based on spatial visual embeddings (education, accessibility, productivity software) Future applications could use JEM-like patch representations for region-aware image editors, visual question answering interfaces, educational object identification, augmented-reality assistance, and tools that let users search or manipulate image regions by meaning. Dependencies: JEM alone is not a language or reasoning model; these systems require multimodal interfaces, grounding mechanisms, user studies, accessibility validation, and protection against incorrect or misleading visual interpretations.
Glossary
- Anti-collapse term: A component of an objective designed to prevent representations from becoming constant or uninformative. “an anti-collapse term that maintains view-specific information”
- Attention probe: A lightweight evaluation method that uses attention-based readouts from a frozen representation. “we report attention probe accuracy on ImageNet-1k”
- Categorical distribution: A probability distribution over a finite set of discrete classes. “We use classes for the categorical distributions.”
- Class token: A special transformer token intended to summarize an entire input, often used for classification. “our model has no class token”
- Conditional entropy: The uncertainty of one random variable given another. “maintaining targets with high marginal entropy and low conditional entropy ”
- Contrastive learning: A representation-learning approach that brings related examples together and separates unrelated examples. “Contrastive methods \citep{he2020moco,chen2020simple} pull together representations of different views of the same image while pushing apart representations of other images.”
- Cosine similarity: A measure of similarity based on the angle between two vectors. “where denotes the cosine similarity”
- Data processing inequality: An information-theoretic result stating that processing data cannot increase its mutual information with another variable. “By the data processing inequality, we have .”
- Dense prediction: A task requiring an output for many or all spatial locations in an input, such as segmentation. “strong performance on both global and dense probing tasks”
- Discrete representation target: A target representation whose values come from a finite or countable set. “our use of discrete representation targets can be seen as an architectural bottleneck”
- Exponential moving average (EMA): A weighted running average that gives greater weight to recent values, commonly used to update teacher models. “the teacher an EMA of the student”
- Foundation model: A large pretrained model intended to support many downstream tasks. “DINOv3 7B foundation model”
- Global collapse: A failure mode in which representations become constant or lose information across inputs. “preventing global collapse of the targets”
- Gram anchoring: A regularization strategy that preserves relationships among feature vectors through their Gram matrix. “DINOv3 even adds a further term, Gram anchoring”
- High-resolution fine-tuning: Additional training on inputs with greater spatial resolution to improve visual detail. “gram and high-resolution fine-tuning”
- Image-level objective: A learning objective applied to a representation of an entire image rather than to local regions. “Unlike methods with separate global and local objectives”
- Information bottleneck: A principle that preserves information predictive of a target while discarding irrelevant information. “the information bottleneck principle suggests retaining predictive information while compressing superfluous information”
- Information collapse: The loss of informative variation in learned representations or targets. “one preventing information collapse”
- Information-theoretic objective: An optimization objective formulated using quantities from information theory. “we construct an information-theoretic objective decomposing into interpretable terms.”
- Inductive bias: A preference introduced by a model architecture, training procedure, or data construction that influences what it learns. “learning good representations also relies on inductive biases from architecture and view construction.”
- Instance segmentation: A vision task that identifies and delineates each individual object instance. “mask AP for instance segmentation on COCO”
- Joint embedding: A representation space in which inputs from different views or modalities are mapped for comparison. “a novel joint-embedding multi-view method”
- Kullback–Leibler divergence: A measure of how one probability distribution differs from another. “where is the entropy of the target and RI(X_S; X_T)$.”
- Panoptic segmentation: A vision task combining semantic segmentation and instance segmentation. “PQ for panoptic segmentation on COCONut”
- Patch-level objective: A training objective applied to local image patches rather than only to a whole-image representation. “our model has no class token and applies the objective exclusively to patch representations”
- Photometric transformation: An image transformation that changes appearance-related properties such as color, brightness, or contrast. “Student and teacher views are also independently transformed with a standard set of photometric transformations”
- Projection head: A learned mapping that transforms an encoder representation into a prediction or target space. “a projection head maps the representation to a predictive distribution”
- Redundancy reduction: A learning strategy that discourages different representation dimensions or elements from encoding the same information. “Redundancy-reduction methods”
- Representation collapse: A degeneracy in which different inputs receive identical or nearly identical representations. “the representation avoids the collapse to a single, redundant vector”
- Self-clustering: A self-supervised approach that learns cluster assignments and makes them consistent across views. “Self-clustering approaches such as SwAV~\cite{caron2021swav} and DINO~\citep{caron2021emerging} learn cluster assignments and align those across views.”
- Self-distillation: A training method in which a student learns to match predictions produced by a teacher model without external labels. “global DINO self-distillation”
- Semantic segmentation: A task that assigns a semantic class label to each image pixel. “For semantic segmentation, we report linear probe mIoU”
- Spatial locality: The property that nearby spatial positions retain locally meaningful and distinct information. “dense feature maps progressively lose spatial locality.”
- Structural collapse: A failure mode in which different spatial elements encode redundant or overly similar information. “which we term a structural collapse”
- Student–teacher paradigm: A framework in which a student network learns to match targets generated by a teacher network. “the variational approximation which naturally fits the student-teacher paradigm”
- Total correlation: A multivariate measure of redundancy among several random variables. “the total correlation $\TC(T)$ measures the redundancy among elements”
- Variational approximation: An approximation that uses an optimized auxiliary distribution to estimate an otherwise difficult probabilistic quantity. “we use a variational approximation which naturally fits the student-teacher paradigm”
- Variational lower bound: A tractable quantity that provides a lower bound on an objective such as mutual information. “We can maximize this objective using a variational lower bound”
- View invariance: The property that representations remain similar when an input is altered through transformations or different viewpoints. “enforce view invariance while explicitly preventing collapse across dimensions.”
- Vision Transformer (ViT): A transformer architecture that processes an image as a sequence of patch tokens. “Our encoder is a Vision Transformer (ViT)”
- Visual representation: A learned numerical encoding of visual input used by downstream tasks. “Self-supervised learning (SSL) at scale produces powerful visual representations.”