Generalization of mechanism evidence beyond contrastive scorers
Determine whether the residual-geometry and modality-gap explanation for Orthogonal Matching Pursuit’s behavior persists when frame selection uses image-text matching heads, multimodal-language-model attention, or subtitle-aware scorers rather than frozen contrastive image-text embeddings.
References
The selector ordering survives a swap to SigLIP (\S\ref{sec:scorer}), but both encoders share a contrastive image--text objective; BLIP-ITM heads, MLLM-attention scorers, and subtitle-aware selection remain untested, and the modality-gap explanation in \S\ref{sec:mechanism} is a property of this embedding space rather than of pursuit selection in general.
— Select, Compress, Reinvest: A Controlled Study of Visual-Token Allocation in Long-Video MLLMs
(2609.03820 - Khatri, 3 Sep 2026) in Section 6, “Limitations,” paragraph “Mechanism evidence is encoder-specific”