Study mixed extractor–verifier configurations

Investigate mixed extractor–verifier configurations for MIDR’s extract–verify–refine pipeline to determine how pairing different models affects ingestion cost and enrichment quality.

Background

The paper evaluates several multimodal LLMs as complete enrichment backends and observes that verifier calibration can substantially affect the volume and quality of indexed enrichment. In particular, the open-source Qwen3-Omni verifier is described as overly strict relative to its extractor, whereas GPT-5.1 often expands coverage during verification. The authors therefore leave unresolved whether using different models for extraction and verification can provide a favorable cost–quality tradeoff.

References

A systematic study of mixed extractor/verifier configurations is left to future work.

MIDR: Enrichment-Augmented Indexing for Multimodal Document Retrieval  (2609.01316 - Mahata et al., 1 Sep 2026) in Appendix, Section “Enrichment MLLM Sensitivity,” paragraph “Implication”