Perception failure driving Replica per-scene performance deficit

Identify which perception failure drives the lower per-scene mean of the VOIM open-vocabulary 3D semantic mapping system under Replica’s all-classes scoring protocol.

Background

The paper reports that VOIM achieves a higher pooled mIoU than OVO-SLAM on Replica at matched RGB-D inputs and under fully monocular operation, but OVO-SLAM achieves the higher per-scene mean. The discrepancy is especially pronounced in small office scenes, where a few large-surface classes dominate and individual misassignments have a substantial effect on the score.

Although the paper attributes the broader room-scale limitation to open-vocabulary labeling and notes that the deficit is systematic rather than sampling noise, it explicitly leaves unresolved which particular perception failure causes the lower per-scene mean. Determining this failure could guide improvements to the detector, segmentation masks, region descriptors, or label-fusion components.

References

We have not isolated which perception failure drives the per-scene mean here and we do not claim the mapping stage could repair it.

VOIM: Training-Free Open-Vocabulary 3D Instance Mapping for RGB-D and Monocular SLAM  (2609.00775 - Song et al., 1 Sep 2026) in Section 4.2, “Replica: monocular capability”