Emergent Multi-View Geometry Through Self-Distillation
Abstract: Over a century ago, Henri Poincaré argued that a motionless observer cannot acquire the notion of space. Yet, most visual representation learning methods operate on individual images, while those that leverage multiple views rely on RGB reconstruction, entangling geometry with appearance. We propose Poincar3, a self-supervised method that learns representations from multiple views through self-distillation instead of RGB reconstruction. We combine masked patch and image-level distillation with a teacher that observes additional views, enabling training from scratch without explicit 3D supervision. Poincar3 outperforms both previous single and multi-view self-supervised approaches such as DINOv3, MuM, and Muskie on correspondence estimation, camera pose estimation, and 3D reconstruction. Using a lightweight Poincaré adapter, we also find that our learned features encode camera motion more accurately than existing self-supervised representations.
Paper Prompts
Sign up for free to create and run prompts on this paper.
Top Community Prompts
Explain it Like I'm 14
1. What is this paper about?
This paper introduces Poincar3, an artificial intelligence system that learns about 3D space by studying many pictures of the same place.
For example, imagine showing a computer several photos of a room taken from different positions. From these pictures, the computer should learn:
- which parts of the pictures show the same object,
- how the camera moved,
- where objects are located in 3D,
- and what the scene would look like from another viewpoint.
Most computer-vision systems learn from one image at a time. But one picture does not contain enough information to fully understand depth and space. Poincar3 instead learns from groups of related images, such as frames from a video.
The important idea is that Poincar3 does not need human-created 3D labels or instructions to rebuild every pixel. It learns by comparing the knowledge of two versions of itself.
2. What questions are the researchers asking?
The researchers mainly want to know:
- Can a computer learn 3D information without being given 3D answers?
- Can it learn from several views of a scene without reconstructing the exact colors of every pixel?
- Do the features learned by the model help with important 3D tasks, such as matching objects across images, estimating camera movement, and building 3D models?
- Which parts of the training method are most useful?
The paper is inspired by an idea from the mathematician Henri Poincaré: an observer may need to experience movement in order to understand space. In a similar way, the model learns about space by seeing a scene from different viewpoints.
3. How did the researchers build and test Poincar3?
Training with a teacher and a student
Poincar3 uses a method called self-distillation. This is like giving a student a slightly more experienced version of the same student as a teacher.
- The student network sees images with some small square regions hidden.
- The teacher network sees complete images.
- The student tries to produce representations similar to the teacher’s representations.
- The teacher is not manually taught. Instead, it is updated slowly using an average of the student’s past versions.
A representation is the computer’s internal summary of an image. It is not simply a copy of the image. It is more like a set of notes describing useful information, such as shapes, positions, and relationships between objects.
Using several views
The images come from the same scene. The student may see some frames, while the teacher sees those frames plus extra frames from the same scene.
This gives the teacher a broader view of the scene. It is similar to asking one student to solve a puzzle using only a few clues, while another student is allowed to look at more clues. The first student then learns from the second student’s better understanding.
Predicting hidden patches
The student’s images are divided into small square pieces called patches. Some patches are hidden, and the student must create useful internal features for them.
Importantly, the student does not have to guess the exact RGB colors of the missing pixels. Instead, it tries to match the teacher’s more meaningful feature predictions.
This matters because exact colors and textures can change between views. For example, an object may look brighter from one angle, but its shape and position remain more stable.
Comparing whole images
Poincar3 also compares a summary of each complete image. This is called the image-level objective.
This helps the model understand the overall scene, not just separate small patches. The researchers found that this part was especially important for learning 3D information.
The model and training data
The main model is a large transformer, a type of neural network that can compare many parts of images and many images with one another.
The model was trained on unlabeled image sequences from sources such as:
- internet videos,
- real-estate videos,
- indoor scans,
- and collections of photographs.
The training was self-supervised, meaning that people did not need to label the images with camera positions, object locations, or 3D shapes.
The researchers then tested Poincar3 on several tasks:
- Correspondence estimation: finding the same object or image patch across different views.
- Camera pose estimation: figuring out how the camera moved.
- Point-cloud estimation: creating a collection of 3D points representing a scene.
- 3D reconstruction: building a model of the scene from several images.
4. What did the researchers discover?
Poincar3 learned useful 3D information without labels
The model was able to match parts of a scene across different images, even though it was never directly told which patches matched.
For example, if a chair appears in several images, Poincar3 can often identify the chair in each image. This ability is called zero-shot correspondence estimation: the model performs the task without being specially trained for it.
It performed better than earlier methods
Poincar3 was compared with systems such as DINOv3, MuM, and Muskie. These are other methods for learning visual features.
Poincar3 generally performed better on:
- matching points across multiple images,
- estimating the camera’s movement,
- estimating depth,
- and reconstructing 3D scenes.
For multi-view matching, the paper reports that Poincar3 reached an accuracy of 94.9% at a relatively generous matching distance on the ScanNet dataset. This was higher than the other self-supervised methods tested.
The teacher’s extra views were very helpful
The researchers tested different versions of their method. Their results showed that performance improved when they:
- used several views instead of one,
- added the whole-image comparison,
- allowed the teacher to see extra views,
- and increased the model’s size and training resources.
In one experiment, the matching score on ScanNet improved from 54.5 with RGB reconstruction to 83.7 with the full Poincar3 method.
It learned about camera motion
The researchers also tested whether the model’s features contained information about camera movement.
They added a small extra network called a Poincaré adapter. This adapter learned to translate changes in the model’s features into changes in the camera’s position and rotation.
Poincar3’s features were better at this than the features from the comparison methods. This suggests that the model’s internal knowledge contains information about how the camera moves through space.
Why avoiding RGB reconstruction helps
Older approaches often train by trying to recreate the exact colors of missing pixels. The problem is that colors and lighting can change between views.
For example, the same wall might look different because:
- the camera angle changed,
- sunlight moved,
- the image was darker,
- or the video was compressed.
Poincar3 focuses more on stable information, such as shape, location, and matching parts. This makes its features more useful for geometry and 3D tasks.
5. Why is this research important?
Poincar3 shows that a computer can develop a useful understanding of 3D space from ordinary collections of images and videos, without needing expensive human-made 3D labels.
This could help improve:
- robots that move around in the real world,
- self-driving vehicles,
- augmented and virtual reality,
- 3D mapping,
- camera tracking,
- and systems that create 3D models from photographs.
The method could also make it easier to train powerful 3D systems because unlabeled videos are much easier to collect than carefully labeled 3D data.
However, the system still has limitations. It requires a large amount of computing power, its training can be sensitive to the exact settings, and it is better at understanding geometry than the meaning of objects. It also currently relies on manually chosen frames from videos.
Overall, the paper’s main message is that seeing a scene from many viewpoints can help an AI learn the structure of 3D space, even without being told the correct 3D answers. Poincar3 demonstrates that this learning can happen through comparisons between a student model and a slowly changing teacher model rather than through exact pixel-by-pixel reconstruction.
Knowledge Gaps
Knowledge gaps, limitations, and open questions
- Generalization beyond curated indoor-scene benchmarks remains unresolved. Most evaluations use ScanNet, ScanNet++, NAVI, MegaDepth, ETH3D, DTU, and RealEstate10K; performance on outdoor, urban, aerial, natural, underwater, low-light, and highly heterogeneous environments is not established.
- Robustness to dynamic scenes is unexplored. The method assumes that image sequences depict the same static scene, but its behavior with moving people, vehicles, deformable objects, changing illumination, and camera-induced motion is not evaluated.
- The effect of inaccurate or weakly related frame sequences is unknown. Internet videos may contain cuts, camera motion unrelated to scene geometry, repeated frames, or frames from different locations; the paper does not quantify how such temporal or scene-consistency errors affect training.
- The handcrafted frame-selection strategy remains a substantial unresolved dependency. The paper does not compare alternative sampling policies or determine which temporal spacing, viewpoint diversity, or overlap criteria produce the most useful geometric representations.
- The source of the apparent geometric emergence is not fully identified. It remains unclear whether performance is primarily caused by multi-view attention, teacher access to additional frames, global distillation, masking, data scale, or the 650M-parameter architecture.
- Interactions among design components are only partially studied. The ablations add components sequentially and do not systematically evaluate their pairwise or higher-order interactions across multiple downstream tasks.
- The learned invariances are not characterized. The paper does not establish how features respond to appearance changes such as lighting, weather, texture replacement, seasonal variation, camera exposure, color shifts, or nonrigid changes—despite motivating the method as less appearance-dependent than RGB reconstruction.
- The trade-off between geometric and semantic information is insufficiently quantified. The paper reports reduced semantic performance but does not analyze which semantic capabilities are lost, whether this trade-off is intrinsic to the objective, or whether multitask objectives can recover semantic performance without harming geometry.
- Performance under severe viewpoint and scale changes is not established. The correspondence experiments use sequences of eight views, but the limits of matching under very small overlap, wide baselines, substantial zoom, occlusion, or strong perspective changes remain unknown.
- Occlusion and visibility handling are not explicitly addressed. The method distills patch predictions across views without an explicit visibility model, leaving its reliability on partially observed or newly revealed regions unclear.
- The representation’s metric and geometric properties remain uncertain. High correspondence accuracy and improved pose probing do not establish whether the features encode metrically accurate depth, calibrated camera motion, global scale, or a consistent world coordinate system.
- The Poincaré adapter evaluation does not demonstrate universally intrinsic structure. The adapter is trained separately for each scene, uses only 20 held-out ScanNet++ scenes, and achieves relatively low average ; it remains unclear whether the structure transfers across scenes, cameras, environments, and motion distributions.
- The dependence on scene-specific adapters limits the practical interpretation of camera-motion decodability. The paper does not test a single adapter trained across scenes or evaluate zero-shot camera-motion decoding without per-scene fitting.
- Long-sequence behavior is insufficiently evaluated. Although register attention is intended to support longer sequences, experiments use at most 24 training views and generally short downstream sequences; memory, accuracy, and stability for hundreds or thousands of frames remain open questions.
- Scaling laws are not established. The paper compares selected model and compute settings but does not determine how representation quality changes with model size, number of views, training data, image resolution, or training duration.
- The claimed compute advantage is not compared under fully standardized conditions. The paper reports substantially less compute than DINOv3, but differences in architecture, input modality, data mixture, training duration, and evaluation protocol make the source of the efficiency advantage unclear.
- Sensitivity to optimization and architectural hyperparameters is not systematically mapped. The paper reports fragility to EMA decay, weight decay, and learning-rate scaling, but does not provide stability ranges, failure modes, or principled methods for selecting these values.
- Collapse prevention mechanisms are not individually analyzed in depth. The relative contributions and interactions of Sinkhorn–Knopp normalization, KoLeo regularization, EMA updating, masking, and the global objective are not fully isolated.
- The method’s dependence on large-scale compute and memory remains a practical limitation. Training requires a roughly 650M-parameter model and multiple H200 GPUs; effectiveness for smaller models, consumer hardware, or resource-constrained applications is not demonstrated.
- The role of unlabeled versus 3D-annotated training data is not fully disentangled. Although additional internet data improves results, the paper does not report controlled experiments that match dataset size, scene diversity, and frame statistics between labeled and unlabeled sources.
- Potential dataset leakage and overlap are not examined. The relationship between training collections and evaluation benchmarks is not documented sufficiently to rule out near-duplicate scenes, videos, or camera trajectories influencing the reported results.
- Comparison fairness across baselines remains uncertain. Some baselines use different architectures, pretraining regimes, input resolutions, or supervision sources, and several are excluded from certain finetuning comparisons; a fully controlled comparison with matched capacity and compute is still needed.
- Uncertainty and failure detection are not evaluated. The model does not report confidence estimates for correspondences, pose predictions, or reconstructed geometry, leaving its reliability in safety-critical or downstream automated systems unclear.
- Robustness to distribution shifts and corruptions is untested. Camera noise, compression artifacts, blur, missing frames, lens distortion, rolling shutter, and unusual image resolutions may substantially affect self-distillation and geometric matching.
- The method’s applicability beyond camera pose and reconstruction is unexplored. Its utility for visual localization, SLAM, navigation, robotic manipulation, 3D tracking, novel-view synthesis, depth completion, and scene change detection remains to be established.
- The relationship between attention-based matching and learned patch features is unresolved. Attention maps outperform nearest-neighbor feature matching, but the paper does not determine whether this reflects genuine geometric correspondence, architectural information flow, or an evaluation-specific artifact.
- Theoretical understanding of multi-view self-distillation is lacking. The paper does not explain why additional teacher views and the image-level objective prevent collapse or induce geometry, nor does it characterize the conditions under which the objective admits degenerate solutions.
- Reproducibility is potentially constrained by incomplete implementation and data details. The paper refers to appendix hyperparameters and a dataset mixture, but the provided text does not specify all sampling, augmentation, masking, projection-head, and evaluation details needed to independently reproduce the results.
Practical Applications
Immediate Applications
- 3D reconstruction from ordinary image sequences — robotics, mapping, and software
- Deploy Poincar3 as a pretrained feature backbone for feed-forward estimation of camera pose, depth, point clouds, and surface normals from unlabeled or weakly labeled image sequences.
- Potential products include mobile 3D-scanning applications, rapid room digitization, visual localization modules, and software for converting handheld videos into approximate spatial models.
- The model is particularly useful where RGB reconstruction is undesirable because lighting, texture, or appearance varies across views.
- Dependencies: Practical deployment still requires task-specific heads or fine-tuning, sufficient multi-view overlap, static or approximately static scenes, and validation on the target environment. The reported model has approximately 650 million parameters, so edge deployment may require distillation, quantization, or cloud inference.
- Zero-shot feature matching and image correspondence — computer vision and geospatial systems
- Use Poincar3 patch features or attention maps to match corresponding image regions across multiple views without training correspondence labels for each new domain.
- Applications include panorama alignment, image mosaicing, structure-from-motion initialization, visual localization, inspection-image registration, and correspondence-assisted 3D reconstruction.
- The reported attention-based matching performance, including high PCK at larger tolerances, suggests a practical workflow in which Poincar3 supplies candidate matches that are subsequently filtered by geometric verification such as RANSAC.
- Dependencies: Performance may decline with severe occlusion, dynamic objects, very large viewpoint changes, motion blur, domain shifts, or scenes lacking reliable covisibility. Safety-critical systems should retain geometric consistency checks rather than treating feature matches as definitive.
- Camera-pose estimation with lightweight adapters — augmented reality, robotics, and mobile devices
- Attach a small pose or prediction head to frozen Poincar3 representations for relative camera-motion estimation.
- This can support AR anchoring, indoor navigation, robot visual odometry, camera relocalization, and stabilization of multi-camera or handheld imaging systems.
- A practical implementation could combine Poincar3 features with inertial measurements and a temporal filter, using the visual model to correct accumulated drift.
- Dependencies: The paper demonstrates decodability of camera motion through scene-specific adapters, not a universally calibrated, production-ready pose estimator. Generalization across buildings, cameras, sensor types, and dynamic environments requires additional evaluation and likely domain adaptation.
- Low-label 3D perception development — industrial computer vision
- Use the released weights and code as initialization for inspection, warehouse perception, construction documentation, cultural-heritage capture, and indoor mapping systems where 3D labels are expensive.
- Organizations can pretrain or adapt the representation using their own unlabeled image sequences, then label only a small set of poses, depths, or correspondences for downstream calibration.
- This is especially actionable for companies with large archives of inspection videos but limited annotated 3D data.
- Dependencies: Training from scratch remains computationally demanding, and unlabeled sequences must depict the same scene or object from multiple views. Data governance, licensing, privacy, and removal of personally identifiable imagery are also required.
- Research and teaching infrastructure for self-supervised 3D vision — academia
- Use Poincar3 as an open baseline for studying how geometric structure emerges from self-distillation, including experiments on attention, correspondence, pose representations, and representation collapse.
- The public implementation enables reproducible comparisons with single-view models, RGB-reconstruction methods, and supervised 3D models.
- It can support coursework or laboratory workflows involving feature extraction, nearest-neighbor matching, pose probing, and fine-tuning on datasets such as ScanNet, MegaDepth, or indoor-scene collections.
- Dependencies: Results depend sensitively on EMA decay, weight decay, learning-rate scaling, batch size, and view sampling. Reproduction also requires substantial GPU resources and careful handling of the paper’s malformed or incomplete source excerpts.
- Video and image-sequence indexing — media and content-management software
- Use dense geometric features to identify recurring physical locations, align frames from different recordings, and organize videos by spatial rather than purely semantic similarity.
- Potential tools include automatic scene-transition analysis, location-aware video search, duplicate-scene detection, and alignment of footage captured by different users or cameras.
- Dependencies: The representation prioritizes geometry over semantics, as acknowledged by the authors. A production system would likely need to combine Poincar3 with a semantic embedding model and account for scene changes, lighting variation, and moving objects.
- Policy and public-sector 3D documentation
- Municipalities, museums, emergency-response organizations, and infrastructure agencies could use ordinary video to create preliminary spatial records of buildings, streets, archaeological sites, or damaged assets.
- Because explicit 3D annotations are not required for representation learning, agencies can exploit existing video archives before commissioning expensive surveys.
- Dependencies: Outputs should be treated as approximate documentation rather than legally authoritative measurements unless independently surveyed. Privacy, consent, copyright, data retention, and model bias across geographic and architectural environments must be addressed.
Long-Term Applications
- Robust visual odometry and autonomous navigation — robotics and autonomous vehicles
- Develop a Poincar3-based visual-odometry or SLAM system that uses geometrically consistent patch features for tracking, relocalization, and map maintenance.
- The model’s multi-view correspondence capability could improve navigation for mobile robots, drones, warehouse vehicles, and service robots, especially when explicit depth sensors are unavailable.
- A likely workflow would combine Poincar3 correspondences, differentiable pose optimization, inertial sensing, loop closure, and uncertainty estimation.
- Dependencies: The paper evaluates static-scene benchmarks and relative pose tasks, not complete long-duration SLAM. Further work is needed for dynamic scenes, temporal drift, scale ambiguity, real-time latency, failure detection, and safety certification.
- General-purpose 3D foundation models — software platforms and robotics
- Extend the method into a modular foundation model that jointly supports correspondence, depth, pose, segmentation, object tracking, novel-view understanding, and scene reconstruction.
- The separation between geometric pretraining and downstream heads could enable one shared backbone for multiple products, reducing the need for task-specific labeled datasets.
- Dependencies: The current model is comparatively geometry-oriented and has weaker semantic performance. Joint semantic-geometric training, multimodal inputs, improved calibration, and evaluation across outdoor, medical, industrial, and highly dynamic domains are required.
- Embodied AI and robot manipulation
- Use learned multi-view features to provide robots with persistent object and surface correspondences while they move around a workspace.
- Potential applications include grasp-point transfer between viewpoints, manipulation under camera motion, rearrangement tasks, and updating 3D maps during interaction.
- Dependencies: Manipulation requires object-level identity, fine-grained geometry, occlusion reasoning, depth and force feedback, and robust behavior under object motion. Poincar3 alone does not establish these capabilities.
- AR/VR scene capture and persistent spatial computing
- Integrate Poincar3 into consumer devices or head-mounted systems for rapid spatial mapping, shared anchors, and cross-device alignment without requiring dense RGB reconstruction.
- Its appearance-robust features could help maintain spatial alignment across changes in illumination, texture, or camera viewpoint.
- Dependencies: Real-time inference, low-power operation, accurate metric scale, privacy-preserving on-device processing, and robust handling of people and moving objects remain unresolved. Existing AR systems also require stringent latency and drift guarantees.
- Infrastructure inspection and digital twins — construction, energy, and manufacturing
- Build systems that repeatedly compare image sequences of bridges, factories, power infrastructure, or construction sites, using stable correspondences to detect geometric changes and update digital twins.
- Longitudinal comparison could support progress monitoring, deformation analysis, maintenance prioritization, and remote inspection.
- Dependencies: Reliable change detection requires precise camera calibration, repeatable acquisition, known uncertainty, and separation of true structural change from lighting, weather, vegetation, or viewpoint effects. Regulatory acceptance would require extensive validation.
- Navigation and mapping in GPS-denied environments — emergency services and defense
- Adapt the representation for underground facilities, disaster zones, mines, and indoor environments where GPS and detailed prior maps are unavailable.
- Unlabeled responder video could be used to construct provisional maps and estimate motion while preserving cross-view geometric consistency.
- Dependencies: These settings include smoke, darkness, debris, dynamic crowds, and severe occlusion—conditions not established by the paper’s benchmarks. Robustness, cybersecurity, human oversight, and mission-specific validation would be essential.
- Scientific and medical 3D imaging from multi-view capture
- Explore use in microscopy, specimen digitization, endoscopy, surgical video, and biological imaging where multiple views exist but dense 3D labels are scarce.
- The self-supervised objective could reduce annotation requirements for reconstructing anatomical or scientific structures.
- Dependencies: Medical and scientific imagery may violate assumptions about static scenes, natural-image appearance, camera motion, and scene overlap. Domain-specific validation, uncertainty quantification, patient privacy, and regulatory approval would be necessary before clinical use.
- Automated temporal view selection and scalable internet-video training
- Replace the paper’s handcrafted frame-selection procedure with a learned sampler that selects informative, geometrically diverse, and covisible frames from long videos.
- This could reduce training cost while improving performance on difficult viewpoints and enable continual pretraining from large-scale video streams.
- Dependencies: Selection must avoid redundant frames, scene cuts, dynamic content, and misleading correspondences. Automated mining also introduces copyright, consent, dataset contamination, and demographic or geographic coverage concerns.
- Geometry-aware generative and simulation systems
- Combine Poincar3 features with generative video, novel-view synthesis, or simulation models to enforce consistent camera motion and scene geometry across generated frames.
- Potential products include more spatially coherent virtual environments, synthetic training data for robots, and interactive digital twins.
- Dependencies: The paper does not demonstrate generation or temporal synthesis. Bridging the representation to generative models requires explicit 3D consistency objectives, controllable camera conditioning, evaluation of hallucinated geometry, and safeguards against visually plausible but metrically incorrect scenes.
Glossary
- Ablation study: An experiment that systematically removes or changes components of a method to measure their individual contributions. “An ablation study of the key components underlying our method.”
- Attention map: A representation of the strength of attention assigned to different input elements by a neural network. “The attention map is an even more powerful correspondence estimator”
- Backbone network: The main feature-extraction network whose outputs are used by later task-specific components. “we parameterize our model by a backbone network ”
- Camera pose estimation: The task of determining a camera’s position and orientation relative to a scene. “Poincar3 outperforms both previous single and multi-view self-supervised approaches such as DINOv3, MuM, and Muskie on correspondence estimation, camera pose estimation, and 3D reconstruction.”
- Contrastive learning: A representation-learning approach that brings related examples closer and separates unrelated examples in feature space. “Later work focused on clustering~\citep{Caron_2018_ECCV,asano2020self,caron2020unsupervised,caron2021dino} and contrastive learning”
- Correspondence estimation: Identifying matching points or image patches across different views of the same scene. “Correspondence estimation is at the heart of multiple-view geometry”
- Cross-entropy loss: A loss function measuring the discrepancy between a target probability distribution and a predicted distribution. “denotes the cross-entropy loss.”
- Dense feature: A feature representation computed for many or all local image regions rather than for the image as a whole. “Our goal is to learn a set of dense patch features”
- Depth head: A task-specific neural-network component that predicts the distance of scene points from the camera. “train only the pose and depth head”
- Distillation: Training one model to reproduce the predictions or representations of another model. “We propose Poincar3, a self-supervised method that learns representations from multiple views through self-distillation”
- Exponential moving average (EMA): A running weighted average that gives greater influence to recent parameter values and is used here to update the teacher network. “the teacher parameters are updated as an exponential moving average (EMA) of the student parameters”
- Feed-forward reconstruction: Directly predicting scene geometry from input images in a single forward pass, without iterative geometric optimization. “where a multi-view transformer directly predicts scene geometry from a sequence of images.”
- Foundation model: A broadly pretrained model intended to support many downstream tasks. “We introduced Poincar3, a multi-view foundation model trained from scratch in a self-supervised manner.”
- Global attention: Attention in which tokens can exchange information across the entire input, rather than only within a local region or frame. “Rotary positional embeddings (RoPE)~\citep{rope:2023} are applied only within the frame-wise attention layers”
- Global representation: A feature vector intended to summarize an entire image rather than an individual patch. “where $\mathcal{H}=\{#1{H}_1,\ldots,#1{H}_M\}$ is the global ([CLS]) representation for each image.”
- Image-level distillation: Distillation that aligns representations or predictions summarizing complete images. “We additionally align the image-level representations”
- Masked image modeling: A self-supervised objective in which parts of an image are hidden and the model must infer information about them. “CroCo~\citep{weinzaepfel2022croco,weinzaepfel2023croco}, MuM~\citep{nordstrom2026mum}, and Muskie~\citep{li2025muskiemultiviewmaskedimage} extend masked image modeling”
- Masked patch prediction: Predicting the representation of an image patch after that patch has been hidden from the student model. “Following iBOT~\citep{zhou2022ibot}, the patch prediction loss is computed only over masked patches”
- Multi-view geometry: The study of spatial structure and camera relationships using multiple images of a scene. “multi-view geometry emerges without labels”
- Multi-view transformer: A transformer architecture designed to exchange information among multiple images or views. “We operate on image sequences from the same scene where a multi-view transformer propagates information between frames.”
- Nearest-neighbor matching: Matching an item to the item with the most similar representation according to a chosen distance or similarity measure. “tracks are produced by either nearest-neighbor matching in feature space”
- Normal consistency: A metric comparing the orientations of surface normals in reconstructed and reference geometry. “normal consistency (NC) by the cosine of the angle between the normals.”
- Photometric augmentation: An image transformation that changes visual properties such as brightness, color, or contrast while preserving scene structure. “the student processes a masked and photometrically augmented subset.”
- Point cloud estimation: Predicting a collection of three-dimensional points representing the geometry of a scene. “Poincar3 outperforms state-of-the-art SSL baselines across multi-view geometric tasks, including pose estimation, point cloud estimation, and image matching.”
- Projection head: A neural-network module that maps learned features into a space used for a training objective. “The backbone outputs are passed through a projection head (MLP + softmax)”
- Poincaré adapter: A lightweight learned mapping intended to transform nonlinear feature changes into a space where camera motion is approximately linear. “The network is called a Poincaré adapter”
- Positional embedding: Information added to token representations to encode their location or ordering. “Rotary positional embeddings (RoPE)~\citep{rope:2023} are applied only within the frame-wise attention layers”
- Pretext task: A self-supervised task constructed from the input data itself to train useful representations without manual labels. “early works used hand-crafted pretext tasks”
- Representational collapse: A failure mode in which a model maps many or all inputs to nearly identical representations. “Sinkhorn-Knopp~\citep{SinkhornKnopp1967} and KoLeo~\citep{sablayrolles2019spreading} regularization are applied to prevent representational collapse.”
- Rotary positional embedding (RoPE): A positional encoding method that represents token positions by rotating components of their feature vectors. “Rotary positional embeddings (RoPE)~\citep{rope:2023}”
- Self-distillation: Training a model using a teacher derived from the model itself rather than from an independently supervised model. “We address this question through self-distillation”
- Self-supervised learning (SSL): Learning representations from unlabeled data by constructing supervisory signals from the data itself. “Self-supervised learning (SSL) has produced increasingly powerful single-image representations”
- Sinkhorn–Knopp regularization: A normalization procedure based on iterative row and column scaling, used here to stabilize learned feature distributions. “We apply Sinkhorn-Knopp and KoLeo regularization to prevent representational collapse.”
- Siamese training: Training shared or related network branches on paired inputs so their outputs encode useful relationships. “by fitting a trainable adapter (an MLP) in a Siamese manner”
- State of the art: The highest-performing known method or benchmark result for a particular task. “We illustrate its state-of-the-art empirical performance”
- Teacher–student framework: A training arrangement in which a student network learns to match outputs from a teacher network. “we adopt the teacher--student self-distillation framework.”
- Transformer decoder: A transformer component that processes encoded representations to produce or refine task-relevant outputs. “followed by a multi-view transformer decoder.”
- Zero-shot correspondence estimation: Estimating correspondences without task-specific training or correspondence labels. “we show that the learned representations exhibit emergent multi-view geometric capabilities, including zero-shot multi-view correspondence estimation”
- : The mathematical group of three-dimensional rotations and translations, representing rigid-body transformations in 3D space. “we evaluate whether the features capture the geometry of .”
- twist: A six-dimensional representation of an infinitesimal 3D rigid motion, comprising rotational and translational components. “The motion between frames and is the twist”






