---
title: 'Poincar3: Multi-View Representation Learning'
url: https://www.emergentmind.com/topics/poincar3
type: topic
---

# Poincar3: Multi-View Representation Learning

Poincar3 is a self-supervised multi-view representation-learning method that learns dense visual features encoding surface identity, cross-view correspondence, and camera motion from unlabeled image sequences. It uses self-distillation rather than RGB reconstruction: a student predicts latent patch and image-level representations produced by an exponential-moving-average teacher that observes additional views. The method is trained from scratch without ground-truth depth, camera poses, correspondence labels, explicit 3D reconstruction targets, or RGB reconstruction losses [2609.39227].

## 1. Problem formulation and design rationale

Given a sequence of images of a common scene,

$$
\mathcal I=\{I_1,\ldots,I_M\},
$$

Poincar3 produces dense patch representations

$$
\mathcal Z=\{Z_1,\ldots,Z_M\},\qquad Z_i\in\mathbb R^{N\times C},
$$

and one global representation per image,

$$
\mathcal H=\{H_1,\ldots,H_M\},\qquad H_i\in\mathbb R^C.
$$

The intended representation assigns similar features to image patches depicting the same 3D point or geometrically meaningful surface region, even under viewpoint, illumination, and appearance changes.

Single-image self-supervised learning methods—including DINO, DINOv2, DINOv3, iBOT, and MAE—can learn semantic and visual invariances from augmentations of one image, but a single image generally does not determine camera motion, depth and scale, cross-view correspondence, or whether a change is caused by viewpoint or appearance. Poincar3 instead uses image sequences so that spatial structure can be learned from changes induced by camera motion.

Prior multi-view masked-image methods such as CroCo, MuM, and Muskie typically use one view to reconstruct RGB content in another. Such objectives provide cross-view supervision but require representations to encode geometry together with texture, lighting, color, and view-dependent effects. Poincar3 transfers latent predictions rather than pixels, making cross-view feature agreement the direct training signal.

The method’s central asymmetry is that the teacher receives a richer visual context than the student. The student is required to infer masked latent representations from visible content and cross-view information, whereas the teacher can use additional unmasked frames to construct targets.

## 2. Architecture and multi-view processing

Poincar3 uses a large multi-view transformer $f_\theta$ with a ViT-L/16 image encoder initialized from scratch. Images are resized to $256\times256$ pixels and divided into patches of size $16$. The encoder incorporates RoPE positional embeddings, QK normalization, and LayerScale.

Per-frame tokens are processed by a multi-view transformer decoder inspired by VGGT-$\Omega$. The decoder has 12 layers, hidden dimension $C=1024$, alternating frame-wise and global/inter-frame attention, RoPE within frame-wise attention, and register attention in selected blocks. Register-bottleneck blocks $\{2,6,9\}$ use 16 register tokens. The complete model contains approximately 650 million trainable parameters.

The backbone outputs patch and global representations,

$$
(\mathcal Z,\mathcal H)=f_\theta(\mathcal I).
$$

Separate projection heads map these representations to DINO/iBOT-style prototype distributions:

$$
P_i=g_p(Z_i),\qquad G_i=g_g(H_i).
$$

The patch and global heads use 65,536 prototypes, hidden dimension 2048, and bottleneck dimension 256.

Poincar3 maintains a student network $f_{\theta_s}$ and a teacher network $f_{\theta_t}$. The teacher is an exponential-moving-average copy of the student. The student receives $M$ views, while the teacher receives those same $M$ views plus $T$ additional views. Student views are masked; teacher views, including teacher-only views, are unmasked. Losses are computed only on the views seen by both networks.

In the main training description, $T$ is sampled uniformly from 0 to 12. The appendix specifies that student sequence length is sampled uniformly from 2 to 24, with 0–12 additional teacher frames. The additional teacher frames are not directly supervised; they influence the teacher’s representations through inter-frame attention.

## 3. Self-distillation objective

Student and teacher images receive independently augmented versions of the sequence. Poincar3 retains full images rather than using the local/global crop strategy common in DINO. For each student image, a binary mask

$$
\mathbf m_i\in\{0,1\}^{N}
$$

is sampled, where $\mathbf m_i(u)=1$ denotes a masked patch. The masked image is

$$
\widetilde I_i=(1-\mathbf m_i)\odot I_i.
$$

Masking is sampled independently per frame. The masking probability is $0.5$, the mask ratio is sampled from $[0.1,0.5]$ following iBOT, only student images are masked, and teacher images remain unmasked.

Teacher outputs are converted into target distributions using Sinkhorn–Knopp normalization. With teacher temperature $\tau_t$ and student temperature $\tau_s$,

$$
p^t=\operatorname{SK}\!\left(\frac{z^t}{\tau_t}\right),\qquad
p^s=\operatorname{softmax}\!\left(\frac{z^s}{\tau_s}\right).
$$

The student temperature is fixed at $\tau_s=0.1$. The teacher temperature is warmed from $0.04$ to $0.07$ during the first 30,000 steps. Sinkhorn–Knopp uses three iterations, and the final layer of each head is frozen for the first 2,000 steps. The paper reports that Sinkhorn–Knopp is more resistant to collapse than simple centering in this multi-view setting.

### Masked patch distillation

Let $\mathcal M_i$ denote the masked patch indices in image $i$. For teacher and student patch distributions $P^t_{i,u}$ and $P^s_{i,u}$, the patch loss is

$$
\mathcal L_{\mathrm{patch}}
=
\frac{1}{\sum_{i=1}^{M}|\mathcal M_i|}
\sum_{i=1}^{M}
\sum_{u\in\mathcal M_i}
\operatorname{CE}(P^t_{i,u},P^s_{i,u}),
$$

where

$$
\operatorname{CE}(p,q)=-\sum_k p_k\log q_k.
$$

The student therefore predicts the teacher’s latent patch distribution at locations whose pixels have been removed. Since the teacher has processed additional views, its target can incorporate geometric evidence unavailable to the student.

### Image-level distillation

For global prototype distributions $G_i^t$ and $G_i^s$ on common frames,

$$
\mathcal L_{\mathrm{global}}
=
\frac{1}{M}
\sum_{i=1}^{M}
\operatorname{CE}(G_i^t,G_i^s).
$$

The image-level objective is reported as essential for stabilizing the multi-view extension and encouraging a 3D-aware representation. Teacher-only frames affect this loss indirectly through the teacher’s multi-view attention.

### Total loss and optimization

Poincar3 combines patch distillation, global distillation, and KoLeo regularization:

$$
\mathcal L
=
\mathcal L_{\mathrm{patch}}
+
\alpha\mathcal L_{\mathrm{global}}
+
\beta\mathcal L_{\mathrm{KoLeo}},
$$

with

$$
\alpha=0.5,\qquad \beta=0.1.
$$

Only the student receives gradient updates. The teacher is updated by

$$
\theta_t\leftarrow\lambda\theta_t+(1-\lambda)\theta_s.
$$

The appendix reports an EMA schedule from $0.994$ to $0.999$ during early training, with $0.999$ as the main decay value.

## 4. Training regime and data

Poincar3 is trained for 400,000 steps with AdamW. The principal settings are learning rate $2\times10^{-4}$, weight decay $0.04$, and EMA decay approximately $0.999$. Additional settings include a 10,000-step learning-rate warmup, gradient clipping at norm $1.0$, bf16 mixed precision, dynamic batching, a budget of 64 frames per GPU, eight H200 GPUs, and approximately three days of training.

The dataset mixture contains approximately 204,376 scenes from internet video and 3D datasets:

- SpatialVID
- DL3DV
- RealEstate10K
- MegaDepth
- AerialMD
- BlendedMVS
- Hypersim
- TartanAir v2
- Map-Free
- ScanNet++
- FlyingThings3D
- ARKitScenes
- UnrealStereo4K
- Virtual KITTI 2

The 3D-annotated datasets are included in the mixture but do not provide training targets for the self-distillation objective. The training signal is formed from image sequences, masking, teacher-generated latent targets, cross-view attention, image-level distillation, and appearance augmentation.

The additional teacher views function as context rather than extra supervised examples. They are excluded directly from patch and global losses, but modify teacher representations of the common frames through multi-view attention.

## 5. Poincaré adapter and geometric probes

The name Poincar3 refers to Henri Poincaré’s observation that spatial understanding requires observing motion. It does not indicate that the method embeds features in a Poincaré ball or performs hyperbolic neural-network operations. No hyperbolic distance, Möbius addition, exponential map, or logarithmic map on a hyperbolic manifold is used in the adapter. The only Lie-group logarithm in the evaluation is the logarithm on $\mathrm{SE}(3)$.

For camera poses $P_t\in\mathrm{SE}(3)$, relative motion is represented by

$$
\Delta P_{t,s}
=
\left(
\operatorname{Log}(P_t^{-1}P_{t+s})
\right)^\vee
\in\mathbb R^6.
$$

The six-vector contains three rotational and three translational components. A lightweight MLP $\varphi_\phi$ maps a frame feature into a space in which camera changes are approximately linear:

$$
\varphi_\phi:\mathbb R^C\rightarrow\mathbb R^D.
$$

A learned linear map $W$ is fitted using

$$
\Delta P_{t,s}
\approx
W\left(
\varphi_\phi(H_{t+s})-\varphi_\phi(H_t)
\right).
$$

The adapter is trained separately per scene on 20 held-out ScanNet++ scenes using a chronological 80/20 split, strides $s\in\{2,\ldots,60\}$, four candidate feature depths, ten adapter seeds, held-out frame pairs, and $R^2$ as the principal metric. The reported clipped mean is

$$
\overline{R^2}
=
\operatorname{mean}\bigl(\max(0,R^2)\bigr),
$$

together with the fractions of scenes having $R^2>0$ and $R^2>0.3$.

| Method and context | Avg. clipped $R^2$ | $R^2>0.3$ |
|---|---:|---:|
| DINOv3, single-view | 0.046 | 0.9% |
| MuM, multi-view | 0.081 | 6.5% |
| Poincar3, single-view | 0.053 | 2.0% |
| Poincar3, multi-view | **0.098** | **7.1%** |

Poincar3’s multi-view features provide the highest reported camera-motion decodability among the listed methods under this protocol. The adapter is a post-hoc probe rather than an intrinsic equivariant $\mathrm{SE}(3)$ representation; it measures decodability after nonlinear feature transformation.

## 6. Empirical performance, ablations, and limitations

### Correspondence estimation

Zero-shot correspondence is evaluated on eight covisible views. Query patches are sampled in the first image and tracked through the other seven. Feature nearest-neighbor matching selects the target patch with maximum cosine similarity, while attention matching selects the target patch with maximum cross-view attention activation.

For nearest-neighbor matching, Poincar3 obtains ScanNet PCK values of 27.9, 46.2, 79.5, and 89.6 at thresholds 5, 10, 25, and 50 pixels, respectively. On NAVI, the corresponding values are 19.2, 33.7, 68.7, and 84.2.

For attention matching, Poincar3 obtains ScanNet values of 27.0, 43.9, 83.7, and 94.9, and NAVI values of 19.3, 34.7, 74.5, and 86.5, at the same thresholds. The reported ScanNet attention result of 94.9 PCK@50 demonstrates correspondence-like behavior emerging without correspondence supervision.

On ScanNet-1500, Poincar3 obtains nearest-neighbor PCK@8/16/32 of 24.3/42.5/58.6 and linear-probe PCK@8/16/32 of 45.3/67.0/79.3.

### Camera pose and reconstruction

For feed-forward reconstruction, the pretrained backbone is paired with a camera head predicting translation, rotation quaternion, and field of view, and a depth head predicting per-pixel depth and confidence. Relative pose is evaluated using AUC at $3^\circ$ and $30^\circ$. Point clouds are obtained by unprojecting predicted depth, followed by Umeyama similarity alignment and ICP refinement.

With frozen-backbone heads, Poincar3 reports:

| Evaluation | Poincar3 result |
|---|---:|
| RealEstate10K AUC@3° / @30° | 2.4 / 51.9 |
| ScanNet++ AUC@3° / @30° | 0.1 / 48.7 |
| MegaDepth AUC@3° / @30° | 3.7 / 67.5 |
| ETH3D accuracy / normal consistency | 0.82 / 0.65 |
| DTU accuracy / normal consistency | 10.60 / 0.59 |

With full finetuning, the reported RealEstate10K values are 8.3 at $3^\circ$ and 68.2 at $30^\circ$; ScanNet++ values are 4.5 and 67.0; MegaDepth values are 7.4 and 74.6; ETH3D accuracy and normal consistency are 0.81 and 0.80; and DTU accuracy and normal consistency are 6.73 and 0.63.

### Ablations

The central multi-view PCK@25 ablation is:

| Objective or configuration | ScanNet | NAVI |
|---|---:|---:|
| RGB reconstruction, MuM objective | 54.5 | 46.3 |
| Single-view DINOv2 objective | 47.2 | 35.7 |
| Naive multi-view DINOv2 | 49.7 | 35.9 |
| Multi-view iBOT | 55.7 | 46.4 |
| Plus image-level objective | 66.7 | 58.1 |
| Plus teacher sees more views | 70.2 | 61.5 |
| Poincar3 scale | **83.7** | **74.5** |

The ablations indicate that a naive multi-view extension of DINO is inadequate, masked patch distillation alone is insufficient, image-level distillation substantially improves correspondence, additional teacher views provide further geometric context, and model and compute scale contribute to the final performance. Retaining full images is reported to be important for stable training.

Training only on 3D-labeled data produces ScanNet PCK@50 of 87.5, compared with 94.9 when all available data are used; NAVI improves from 78.5 to 86.5. BYOL-style training initially appeared promising but later suffered feature degradation, while centering, alternative Sinkhorn–Knopp configurations, and VICReg did not resolve all stability problems.

### Scope and limitations

Poincar3’s reported advantages concern correspondence, camera pose estimation, point-cloud reconstruction, and camera-motion decodability under the stated protocols. They do not establish that every feature dimension has a unique physical interpretation, that the representation is globally isometric to $\mathrm{SE}(3)$, or that geometry is completely disentangled from appearance. Removing RGB reconstruction makes geometry a more direct source of agreement, but does not prove complete disentanglement.

The paper identifies several limitations: semantic performance is substantially below DINOv3; the model remains computationally expensive; self-distillation is sensitive to EMA, weight decay, temperature, and learning-rate choices; training is more sensitive to low-quality data than RGB reconstruction; frame selection from video is hand-designed; the Poincaré adapter is a post-hoc probe rather than an intrinsic equivariant representation; and geometric evaluation remains benchmark- and protocol-dependent.

Poincar3 therefore represents a multi-view self-distillation framework in which latent consistency across camera motion replaces pixel reconstruction as the principal geometric signal. Its defining mechanism is the combination of masked patch prediction, image-level distillation, cross-view attention, and privileged teacher context from additional views.

Source: https://www.emergentmind.com/topics/poincar3