Papers
Topics
Authors
Recent
Search
2000 character limit reached

Poincar3: Multi-View Representation Learning

Updated 2 October 2026
  • Poincar3 is a multi-view self-supervised approach that learns spatial structures from image sequences, integrating surface identity and cross-view correspondence without explicit ground-truth data on depth, camera poses, or 3D reconstructions.
  • The method emphasizes latent predictions over pixel reconstruction, utilizing cross-view feature agreement as the primary learning signal, keeping geometric context crucial, and ensuring latent representations are consistent.
  • Poincar3 leverages a student-teacher architecture with exponential moving-average teachers that observe richer visual contexts, achieving high correspondence and camera motion accuracy.

Poincar3 is a self-supervised multi-view representation-learning method that learns dense visual features encoding surface identity, cross-view correspondence, and camera motion from unlabeled image sequences. It uses self-distillation rather than RGB reconstruction: a student predicts latent patch and image-level representations produced by an exponential-moving-average teacher that observes additional views. The method is trained from scratch without ground-truth depth, camera poses, correspondence labels, explicit 3D reconstruction targets, or RGB reconstruction losses (Nordström et al., 30 Sep 2026).

1. Problem formulation and design rationale

Given a sequence of images of a common scene,

I={I1,…,IM},\mathcal I=\{I_1,\ldots,I_M\},

Poincar3 produces dense patch representations

Z={Z1,…,ZM},Zi∈RN×C,\mathcal Z=\{Z_1,\ldots,Z_M\},\qquad Z_i\in\mathbb R^{N\times C},

and one global representation per image,

H={H1,…,HM},Hi∈RC.\mathcal H=\{H_1,\ldots,H_M\},\qquad H_i\in\mathbb R^C.

The intended representation assigns similar features to image patches depicting the same 3D point or geometrically meaningful surface region, even under viewpoint, illumination, and appearance changes.

Single-image self-supervised learning methods—including DINO, DINOv2, DINOv3, iBOT, and MAE—can learn semantic and visual invariances from augmentations of one image, but a single image generally does not determine camera motion, depth and scale, cross-view correspondence, or whether a change is caused by viewpoint or appearance. Poincar3 instead uses image sequences so that spatial structure can be learned from changes induced by camera motion.

Prior multi-view masked-image methods such as CroCo, MuM, and Muskie typically use one view to reconstruct RGB content in another. Such objectives provide cross-view supervision but require representations to encode geometry together with texture, lighting, color, and view-dependent effects. Poincar3 transfers latent predictions rather than pixels, making cross-view feature agreement the direct training signal.

The method’s central asymmetry is that the teacher receives a richer visual context than the student. The student is required to infer masked latent representations from visible content and cross-view information, whereas the teacher can use additional unmasked frames to construct targets.

2. Architecture and multi-view processing

Poincar3 uses a large multi-view transformer fθf_\theta with a ViT-L/16 image encoder initialized from scratch. Images are resized to 256×256256\times256 pixels and divided into patches of size $16$. The encoder incorporates RoPE positional embeddings, QK normalization, and LayerScale.

Per-frame tokens are processed by a multi-view transformer decoder inspired by VGGT-Ω\Omega. The decoder has 12 layers, hidden dimension C=1024C=1024, alternating frame-wise and global/inter-frame attention, RoPE within frame-wise attention, and register attention in selected blocks. Register-bottleneck blocks {2,6,9}\{2,6,9\} use 16 register tokens. The complete model contains approximately 650 million trainable parameters.

The backbone outputs patch and global representations,

(Z,H)=fθ(I).(\mathcal Z,\mathcal H)=f_\theta(\mathcal I).

Separate projection heads map these representations to DINO/iBOT-style prototype distributions:

Z={Z1,…,ZM},Zi∈RN×C,\mathcal Z=\{Z_1,\ldots,Z_M\},\qquad Z_i\in\mathbb R^{N\times C},0

The patch and global heads use 65,536 prototypes, hidden dimension 2048, and bottleneck dimension 256.

Poincar3 maintains a student network Z={Z1,…,ZM},Zi∈RN×C,\mathcal Z=\{Z_1,\ldots,Z_M\},\qquad Z_i\in\mathbb R^{N\times C},1 and a teacher network Z={Z1,…,ZM},Zi∈RN×C,\mathcal Z=\{Z_1,\ldots,Z_M\},\qquad Z_i\in\mathbb R^{N\times C},2. The teacher is an exponential-moving-average copy of the student. The student receives Z={Z1,…,ZM},Zi∈RN×C,\mathcal Z=\{Z_1,\ldots,Z_M\},\qquad Z_i\in\mathbb R^{N\times C},3 views, while the teacher receives those same Z={Z1,…,ZM},Zi∈RN×C,\mathcal Z=\{Z_1,\ldots,Z_M\},\qquad Z_i\in\mathbb R^{N\times C},4 views plus Z={Z1,…,ZM},Zi∈RN×C,\mathcal Z=\{Z_1,\ldots,Z_M\},\qquad Z_i\in\mathbb R^{N\times C},5 additional views. Student views are masked; teacher views, including teacher-only views, are unmasked. Losses are computed only on the views seen by both networks.

In the main training description, Z={Z1,…,ZM},Zi∈RN×C,\mathcal Z=\{Z_1,\ldots,Z_M\},\qquad Z_i\in\mathbb R^{N\times C},6 is sampled uniformly from 0 to 12. The appendix specifies that student sequence length is sampled uniformly from 2 to 24, with 0–12 additional teacher frames. The additional teacher frames are not directly supervised; they influence the teacher’s representations through inter-frame attention.

3. Self-distillation objective

Student and teacher images receive independently augmented versions of the sequence. Poincar3 retains full images rather than using the local/global crop strategy common in DINO. For each student image, a binary mask

Z={Z1,…,ZM},Zi∈RN×C,\mathcal Z=\{Z_1,\ldots,Z_M\},\qquad Z_i\in\mathbb R^{N\times C},7

is sampled, where Z={Z1,…,ZM},Zi∈RN×C,\mathcal Z=\{Z_1,\ldots,Z_M\},\qquad Z_i\in\mathbb R^{N\times C},8 denotes a masked patch. The masked image is

Z={Z1,…,ZM},Zi∈RN×C,\mathcal Z=\{Z_1,\ldots,Z_M\},\qquad Z_i\in\mathbb R^{N\times C},9

Masking is sampled independently per frame. The masking probability is H={H1,…,HM},Hi∈RC.\mathcal H=\{H_1,\ldots,H_M\},\qquad H_i\in\mathbb R^C.0, the mask ratio is sampled from H={H1,…,HM},Hi∈RC.\mathcal H=\{H_1,\ldots,H_M\},\qquad H_i\in\mathbb R^C.1 following iBOT, only student images are masked, and teacher images remain unmasked.

Teacher outputs are converted into target distributions using Sinkhorn–Knopp normalization. With teacher temperature H={H1,…,HM},Hi∈RC.\mathcal H=\{H_1,\ldots,H_M\},\qquad H_i\in\mathbb R^C.2 and student temperature H={H1,…,HM},Hi∈RC.\mathcal H=\{H_1,\ldots,H_M\},\qquad H_i\in\mathbb R^C.3,

H={H1,…,HM},Hi∈RC.\mathcal H=\{H_1,\ldots,H_M\},\qquad H_i\in\mathbb R^C.4

The student temperature is fixed at H={H1,…,HM},Hi∈RC.\mathcal H=\{H_1,\ldots,H_M\},\qquad H_i\in\mathbb R^C.5. The teacher temperature is warmed from H={H1,…,HM},Hi∈RC.\mathcal H=\{H_1,\ldots,H_M\},\qquad H_i\in\mathbb R^C.6 to H={H1,…,HM},Hi∈RC.\mathcal H=\{H_1,\ldots,H_M\},\qquad H_i\in\mathbb R^C.7 during the first 30,000 steps. Sinkhorn–Knopp uses three iterations, and the final layer of each head is frozen for the first 2,000 steps. The paper reports that Sinkhorn–Knopp is more resistant to collapse than simple centering in this multi-view setting.

Masked patch distillation

Let H={H1,…,HM},Hi∈RC.\mathcal H=\{H_1,\ldots,H_M\},\qquad H_i\in\mathbb R^C.8 denote the masked patch indices in image H={H1,…,HM},Hi∈RC.\mathcal H=\{H_1,\ldots,H_M\},\qquad H_i\in\mathbb R^C.9. For teacher and student patch distributions fθf_\theta0 and fθf_\theta1, the patch loss is

fθf_\theta2

where

fθf_\theta3

The student therefore predicts the teacher’s latent patch distribution at locations whose pixels have been removed. Since the teacher has processed additional views, its target can incorporate geometric evidence unavailable to the student.

Image-level distillation

For global prototype distributions fθf_\theta4 and fθf_\theta5 on common frames,

fθf_\theta6

The image-level objective is reported as essential for stabilizing the multi-view extension and encouraging a 3D-aware representation. Teacher-only frames affect this loss indirectly through the teacher’s multi-view attention.

Total loss and optimization

Poincar3 combines patch distillation, global distillation, and KoLeo regularization:

fθf_\theta7

with

fθf_\theta8

Only the student receives gradient updates. The teacher is updated by

fθf_\theta9

The appendix reports an EMA schedule from 256×256256\times2560 to 256×256256\times2561 during early training, with 256×256256\times2562 as the main decay value.

4. Training regime and data

Poincar3 is trained for 400,000 steps with AdamW. The principal settings are learning rate 256×256256\times2563, weight decay 256×256256\times2564, and EMA decay approximately 256×256256\times2565. Additional settings include a 10,000-step learning-rate warmup, gradient clipping at norm 256×256256\times2566, bf16 mixed precision, dynamic batching, a budget of 64 frames per GPU, eight H200 GPUs, and approximately three days of training.

The dataset mixture contains approximately 204,376 scenes from internet video and 3D datasets:

  • SpatialVID
  • DL3DV
  • RealEstate10K
  • MegaDepth
  • AerialMD
  • BlendedMVS
  • Hypersim
  • TartanAir v2
  • Map-Free
  • ScanNet++
  • FlyingThings3D
  • ARKitScenes
  • UnrealStereo4K
  • Virtual KITTI 2

The 3D-annotated datasets are included in the mixture but do not provide training targets for the self-distillation objective. The training signal is formed from image sequences, masking, teacher-generated latent targets, cross-view attention, image-level distillation, and appearance augmentation.

The additional teacher views function as context rather than extra supervised examples. They are excluded directly from patch and global losses, but modify teacher representations of the common frames through multi-view attention.

5. Poincaré adapter and geometric probes

The name Poincar3 refers to Henri Poincaré’s observation that spatial understanding requires observing motion. It does not indicate that the method embeds features in a Poincaré ball or performs hyperbolic neural-network operations. No hyperbolic distance, Möbius addition, exponential map, or logarithmic map on a hyperbolic manifold is used in the adapter. The only Lie-group logarithm in the evaluation is the logarithm on 256×256256\times2567.

For camera poses 256×256256\times2568, relative motion is represented by

256×256256\times2569

The six-vector contains three rotational and three translational components. A lightweight MLP $16$0 maps a frame feature into a space in which camera changes are approximately linear:

$16$1

A learned linear map $16$2 is fitted using

$16$3

The adapter is trained separately per scene on 20 held-out ScanNet++ scenes using a chronological 80/20 split, strides $16$4, four candidate feature depths, ten adapter seeds, held-out frame pairs, and $16$5 as the principal metric. The reported clipped mean is

$16$6

together with the fractions of scenes having $16$7 and $16$8.

Method and context Avg. clipped $16$9 Ω\Omega0
DINOv3, single-view 0.046 0.9%
MuM, multi-view 0.081 6.5%
Poincar3, single-view 0.053 2.0%
Poincar3, multi-view 0.098 7.1%

Poincar3’s multi-view features provide the highest reported camera-motion decodability among the listed methods under this protocol. The adapter is a post-hoc probe rather than an intrinsic equivariant Ω\Omega1 representation; it measures decodability after nonlinear feature transformation.

6. Empirical performance, ablations, and limitations

Correspondence estimation

Zero-shot correspondence is evaluated on eight covisible views. Query patches are sampled in the first image and tracked through the other seven. Feature nearest-neighbor matching selects the target patch with maximum cosine similarity, while attention matching selects the target patch with maximum cross-view attention activation.

For nearest-neighbor matching, Poincar3 obtains ScanNet PCK values of 27.9, 46.2, 79.5, and 89.6 at thresholds 5, 10, 25, and 50 pixels, respectively. On NAVI, the corresponding values are 19.2, 33.7, 68.7, and 84.2.

For attention matching, Poincar3 obtains ScanNet values of 27.0, 43.9, 83.7, and 94.9, and NAVI values of 19.3, 34.7, 74.5, and 86.5, at the same thresholds. The reported ScanNet attention result of 94.9 PCK@50 demonstrates correspondence-like behavior emerging without correspondence supervision.

On ScanNet-1500, Poincar3 obtains nearest-neighbor PCK@8/16/32 of 24.3/42.5/58.6 and linear-probe PCK@8/16/32 of 45.3/67.0/79.3.

Camera pose and reconstruction

For feed-forward reconstruction, the pretrained backbone is paired with a camera head predicting translation, rotation quaternion, and field of view, and a depth head predicting per-pixel depth and confidence. Relative pose is evaluated using AUC at Ω\Omega2 and Ω\Omega3. Point clouds are obtained by unprojecting predicted depth, followed by Umeyama similarity alignment and ICP refinement.

With frozen-backbone heads, Poincar3 reports:

Evaluation Poincar3 result
RealEstate10K AUC@3° / @30° 2.4 / 51.9
ScanNet++ AUC@3° / @30° 0.1 / 48.7
MegaDepth AUC@3° / @30° 3.7 / 67.5
ETH3D accuracy / normal consistency 0.82 / 0.65
DTU accuracy / normal consistency 10.60 / 0.59

With full finetuning, the reported RealEstate10K values are 8.3 at Ω\Omega4 and 68.2 at Ω\Omega5; ScanNet++ values are 4.5 and 67.0; MegaDepth values are 7.4 and 74.6; ETH3D accuracy and normal consistency are 0.81 and 0.80; and DTU accuracy and normal consistency are 6.73 and 0.63.

Ablations

The central multi-view PCK@25 ablation is:

Objective or configuration ScanNet NAVI
RGB reconstruction, MuM objective 54.5 46.3
Single-view DINOv2 objective 47.2 35.7
Naive multi-view DINOv2 49.7 35.9
Multi-view iBOT 55.7 46.4
Plus image-level objective 66.7 58.1
Plus teacher sees more views 70.2 61.5
Poincar3 scale 83.7 74.5

The ablations indicate that a naive multi-view extension of DINO is inadequate, masked patch distillation alone is insufficient, image-level distillation substantially improves correspondence, additional teacher views provide further geometric context, and model and compute scale contribute to the final performance. Retaining full images is reported to be important for stable training.

Training only on 3D-labeled data produces ScanNet PCK@50 of 87.5, compared with 94.9 when all available data are used; NAVI improves from 78.5 to 86.5. BYOL-style training initially appeared promising but later suffered feature degradation, while centering, alternative Sinkhorn–Knopp configurations, and VICReg did not resolve all stability problems.

Scope and limitations

Poincar3’s reported advantages concern correspondence, camera pose estimation, point-cloud reconstruction, and camera-motion decodability under the stated protocols. They do not establish that every feature dimension has a unique physical interpretation, that the representation is globally isometric to Ω\Omega6, or that geometry is completely disentangled from appearance. Removing RGB reconstruction makes geometry a more direct source of agreement, but does not prove complete disentanglement.

The paper identifies several limitations: semantic performance is substantially below DINOv3; the model remains computationally expensive; self-distillation is sensitive to EMA, weight decay, temperature, and learning-rate choices; training is more sensitive to low-quality data than RGB reconstruction; frame selection from video is hand-designed; the Poincaré adapter is a post-hoc probe rather than an intrinsic equivariant representation; and geometric evaluation remains benchmark- and protocol-dependent.

Poincar3 therefore represents a multi-view self-distillation framework in which latent consistency across camera motion replaces pixel reconstruction as the principal geometric signal. Its defining mechanism is the combination of masked patch prediction, image-level distillation, cross-view attention, and privileged teacher context from additional views.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Poincar3.