Interpretability of Correspondence Structures Under Probing
Ascertain the interpretability of correspondence-related structures encoded within the Visual Geometry Grounded Transformer (VGGT) representation space when assessed via camera-token probing of intermediate layers.
References
How interpretable these structures are in the representation space is not yet clear under this probing scheme.
— On Geometric Understanding and Learned Data Priors in VGGT
(2512.11508 - Bratulić et al., 12 Dec 2025) in Section 4.1 (Does VGGT encode geometry? If so, where?)
We have not run the same test on objects other than people, so we do not claim the behaviour is specific to people: a model that matches any moving surface across frames would pass it, and for our purpose that would serve equally.
— WildHSR: Metric Feed-Forward 4D People-Scene Reconstruction from a 3D Foundation Model
(2609.29106 - Bright et al., 24 Sep 2026) in Appendix, Section “Identity Features and Pelvis Localization,” subsection “What the probe does not show”; see also Conclusion, “Limitations”