Determine the mechanism underlying active-versus-passive observation performance gaps

Determine whether the terminal-success gap between model-directed active camera control and passive five-view observation in VA-Bench is caused by view relevance, cross-view integration, or other factors.

Background

VA-Bench reports that removing camera control reduces task success even when the passive condition supplies five predefined views at every observation point. The active protocol additionally permits task-conditioned viewpoint selection and bounded local refinement, whereas the passive protocol prevents the model from selecting, reordering, or extending the supplied views.

The reported comparison establishes a performance difference but does not identify its cause. Possible explanations include whether active control helps models select task-relevant views, integrate information across views, or benefit from differences in visual-context length. Resolving this issue would clarify which aspect of active evidence acquisition is responsible for the observed improvement.

References

A systematic, double-annotated audit of access-mode reversals remains future work. It should record the earliest trajectory evidence for observation, alignment, integration, abduction, verification, and repair, and test whether removing specific images changes the downstream localization or patch.

— SWE-PolyVision: Benchmarking Cross-Image Abductive Reasoning for Repository-Level Software Engineering  (2609.29754 - Wu et al., 24 Sep 2026) in Section 7, paragraph “Construct validation”

The mechanism behind the gap remains open. It may involve view relevance, cross-view integration, or other factors (Appendix B).

— VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control  (2609.19554 - Zhang et al., 17 Sep 2026) in Section 5.1, Active Evidence Acquisition vs. Passive Multi-View Input; reiterated in Appendix B.1