Identify the causal contribution of OCR and architectural mechanisms to transfer

Establish, through matched interventions, whether OCR supervision causally improves non-text grounding and downstream action transfer, and whether the proposed recurrent-memory, gating, and multi-level visual-injection mechanisms explain the observed architectural transfer differences.

Background

The paper interprets OCR as a possible catalyst because it couples fine-grained visual discrimination with text-region alignment, but the reported task-mixture experiments do not isolate an OCR-only causal effect. Likewise, the discussion proposes mechanisms involving recurrent compression, output gating, feature accessibility, and alignment, while comparisons among complete pretrained foundations confound multiple architectural and training differences.

The authors therefore leave both the causal role of OCR and the explanatory status of the proposed architectural mechanisms unresolved, calling for matched interventions that isolate these factors.

References

OCR-specific causality and the proposed architectural mechanisms remain hypotheses requiring matched interventions.

— GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives  (2609.39601 - Yu et al., 30 Sep 2026) in Section 6, “Limitations and Future Work”