IRPAPERS: Visual-Document IR Benchmark
- The paper presents IRPAPERS as a benchmark that evaluates page-level scientific retrieval and QA using both image and OCR text views.
- It employs detailed needle-in-the-haystack questions and demonstrates that multimodal fusion outperforms pure text or image-based systems.
- Experiments reveal efficiency–accuracy tradeoffs and complementary strengths between text and visual retrieval methods for fine-grained analysis.
Searching arXiv for the benchmark paper and a few related benchmarks mentioned in the source material. {"query": "\"IRPAPERS\" visual document benchmark scientific retrieval question answering", "max_results": 5} {"query": "\"ViDoRe\" visual document retrieval arXiv", "max_results": 5} {"query": "\"FinanceBench\" arXiv question answering financial filings", "max_results": 5} IRPAPERS is a visual-document benchmark for scientific retrieval and question answering over PDFs at the page level, with both image and OCR text views of each page. It comprises 3,230 pages from 166 information retrieval papers and 180 “needle-in-the-haystack” questions, each associated with one specific gold page. The benchmark was introduced to compare image-based systems, text-based systems, and multimodal hybrids in a setting where many papers share vocabulary and topical structure, so success depends on recovering fine-grained methodological detail rather than broad topical similarity. It evaluates both retrieval and retrieval-augmented generation, and it emphasizes the complementarity of page images and OCR transcriptions rather than treating one modality as a simple surrogate for the other (Shorten et al., 5 Feb 2026).
1. Scope and research agenda
IRPAPERS was designed around a specific tension in modern document intelligence. Traditional scientific search systems usually convert PDFs into text and metadata, whereas multimodal foundation models increasingly operate directly on page images. The benchmark therefore asks how image-based scientific retrieval compares to established text-based methods, whether the two modalities fail in complementary ways, how multimodal fusion behaves, and what efficiency–accuracy tradeoffs arise when multi-vector image retrieval is made approximate through methods such as MUVERA (Shorten et al., 5 Feb 2026).
The benchmark is deliberately narrow in subject matter and dense in semantic overlap. All 166 documents come from the citation graph of the survey “LLMs for Information Retrieval: A Survey,” which means the collection contains many papers on closely related topics such as reranking, query rewriting, dense retrieval, and pseudo-relevance feedback. This design makes the retrieval problem adversarial in a specific way: a retriever cannot rely on coarse lexical or topic cues alone, because many distractor pages are topically plausible. Instead, it must recover details such as which instruction-following model HyDE used, what content types GRF generated, or which metric was reported for a specific benchmark page (Shorten et al., 5 Feb 2026).
This setup distinguishes IRPAPERS from general-purpose visual-document benchmarks. Its primary object is not document understanding in isolation, but page-level retrieval and question answering across a semantically crowded scientific corpus. A plausible implication is that IRPAPERS measures not only multimodal representation quality, but also how well a system can discriminate among near-neighbor scientific claims and implementation details (Shorten et al., 5 Feb 2026).
2. Corpus construction and annotation design
The benchmark corpus contains 3,230 pages extracted from 166 scientific papers. Each page is stored in two forms: a base64-encoded page image and a GPT-4.1 OCR transcription. The image representation averages approximately per page, while the OCR text averages approximately 1,125 output tokens per page and is approximately smaller in storage than the image representation. The OCR prompt transcribes text and tables as Markdown, omits images, and preserves captions, which is consequential because captions often carry enough semantic information for text-based retrieval and QA to remain competitive even on visually rich pages (Shorten et al., 5 Feb 2026).
The question set consists of 180 curated “needle-in-the-haystack” questions. These were generated from 19 “Query Writing” papers. For each non-reference page in those papers, the page image was presented to Claude Sonnet 4.5 with instructions to generate one self-contained question, a short answer, and an explanation of why the page uniquely answered the question and why the answer was specific to that paper rather than to related IR papers more broadly. Each query is therefore grounded in one gold page, making the retrieval task single-relevance at the page level (Shorten et al., 5 Feb 2026).
The page construction choices have direct operational implications. Image preprocessing is deterministic and inexpensive in compute, but storage-heavy. OCR text is compact and easier to index, but incurs substantial preprocessing cost and depends on the chosen OCR model. The source material reports GPT-4.1 OCR costs of approximately for all 3,230 pages, with a throughput of approximately 13 pages per minute at 30k tokens per minute. This operational asymmetry is part of the benchmark’s motivation rather than a peripheral implementation detail (Shorten et al., 5 Feb 2026).
3. Retrieval and question-answering tasks
IRPAPERS defines two tasks: page-level retrieval and retrieval-augmented question answering. In the retrieval task, a system receives a question and must rank the page set so that the gold page appears as high as possible. The primary metric is Recall@, with , defined as
Because each query has one relevant page, the metric is effectively Success@ under a single-relevance protocol (Shorten et al., 5 Feb 2026).
The QA task evaluates retrieval-augmented generation in two modalities. In TextRAG, the retrieved context is the OCR transcription of the top-0 pages. In ImageRAG, the retrieved context is the page images themselves. In both cases the reader is GPT-4.1 used through DSPy, and the prompt is shared across modalities. The benchmark also includes “No Retrieval,” “Hard Negative,” and “Oracle Retrieval” conditions to separate parametric knowledge, near-miss distractors, and perfect retrieval from actual end-to-end system behavior (Shorten et al., 5 Feb 2026).
Answer quality is measured by a GPT-4.1 alignment judge rather than by lexical overlap. For each query, the judge sees the question, the system answer, and the gold answer, returns a Boolean semantic-alignment decision, and the benchmark takes the majority vote over three independent judgments. The resulting alignment score is the mean majority-vote accuracy across all queries. This evaluation protocol matters because the answers are often short, technical, and semantically specific; exact-string metrics would understate correct paraphrases and overstate partial lexical overlap (Shorten et al., 5 Feb 2026).
4. Retrieval architectures and multimodal fusion
IRPAPERS compares sparse, dense, late-interaction, and fused retrieval systems. On the text side, the benchmark uses BM25 over GPT-4.1 OCR transcriptions, Arctic 2.0 dense text embeddings with HNSW search and inner-product similarity, and a hybrid text system that combines BM25 and Arctic 2.0 by Relative Score Fusion. In that fusion scheme, each retriever’s scores are min–max normalized to 1, then combined by equal-weight addition (Shorten et al., 5 Feb 2026).
On the image side, the benchmark emphasizes multi-vector late-interaction retrieval. The main open-source image model is ColModernVBERT, which combines a 150M-parameter ModernBERT text encoder and a 100M-parameter SigLIP-2 vision encoder and produces approximately 1,000 vectors of dimension 128 per page. It is scored with ColBERT-style MaxSim,
2
which preserves token-level or patch-level interaction but is expensive in storage and inference. Larger multi-vector models, ColPali and ColQwen2, are also evaluated to examine scale effects on page retrieval (Shorten et al., 5 Feb 2026).
To study efficiency, the benchmark applies MUVERA. A variable-sized multi-vector page representation is converted into a single fixed-dimensional encoding through SimHash bucketing, per-bucket aggregation, random projection, and repetition. With 3, 4, and 10 repetitions, the resulting fixed-dimensional encoding has 2,560 dimensions, reducing ColModernVBERT’s approximate index size from approximately 5 to approximately 6. Retrieval then proceeds in two stages: approximate HNSW search in fixed-dimensional space, followed by exact MaxSim rescoring on the top-7 candidates (Shorten et al., 5 Feb 2026).
The benchmark’s central systems are multimodal hybrids that combine BM25, Arctic 2.0, and ColModernVBERT. Two fusion rules are tested: Relative Score Fusion and Reciprocal Rank Fusion. Relative Score Fusion performs better in the reported experiments. In the multimodal case, the text-hybrid score is combined with the image score through a parameter 8, and with 9 the effective weighting is 0.25 BM25, 0.25 Arctic 2.0, and 0.5 ColModernVBERT. This design makes the fusion behavior interpretable and exposes whether image evidence contributes beyond what sparse and dense text retrieval already recover (Shorten et al., 5 Feb 2026).
5. Empirical findings
The benchmark reports that open-source hybrid text search achieves 46% Recall@1, 78% Recall@5, and 91% Recall@20, while open-source image retrieval with ColModernVBERT reaches 43%, 78%, and 93%, respectively. Their failure sets are complementary: there are 22 queries where hybrid text is correct at rank 1 and image retrieval is not, and 18 queries where image retrieval is correct at rank 1 and text retrieval is not. Multimodal hybrid search improves on both, reaching 49% Recall@1, 81% Recall@5, and 95% Recall@20 (Shorten et al., 5 Feb 2026).
| System | Recall@1 | Recall@5 / Recall@20 |
|---|---|---|
| Hybrid text (BM25 + Arctic 2.0) | 46% | 78% / 91% |
| ColModernVBERT | 43% | 78% / 93% |
| Multimodal hybrid | 49% | 81% / 95% |
| Cohere Embed v4.0 (images) | 58% | 87% / 97% |
Among larger open-source image models, ColPali reaches 45% Recall@1, 79% Recall@5, and 93% Recall@20, while ColQwen2 reaches 49%, 81%, and 94%. This pattern suggests that scale mainly improves top-rank precision rather than deep recall, which is already high for multiple systems by 0. Among closed-source models, Cohere Embed v4.0 page-image embeddings outperform Voyage 3 Large text embeddings and all tested open-source models, achieving 58% Recall@1, 87% Recall@5, and 97% Recall@20. A closed-source multimodal hybrid of Cohere, Voyage, and BM25 improves deeper retrieval further to 91% Recall@5 and 98% Recall@20, while leaving Recall@1 at 58% (Shorten et al., 5 Feb 2026).
MUVERA exposes a clear efficiency–performance tradeoff. Applied to ColModernVBERT, MUVERA with 1 reduces Recall@1 from 43% to 41% and Recall@20 from 93% to 88%; lower 2 values reduce recall more sharply, with 35% Recall@1 and 66% Recall@20 at 3. The practical significance is that approximate multi-vector retrieval can reduce memory by approximately 4, but exact late interaction still remains important when recall at small 5 matters (Shorten et al., 5 Feb 2026).
Question answering results show a different asymmetry. TextRAG is stronger than ImageRAG: TextRAG achieves an alignment score of 0.62 at 6 and 0.82 at 7, whereas ImageRAG achieves 0.40 at 8 and 0.71 at 9. Oracle single-page retrieval yields 0.74 for text and 0.68 for images, so in both modalities multi-page retrieval outperforms oracle single-page retrieval. This is one of the benchmark’s most consequential findings: the answer often benefits from several related pages rather than from the gold page alone, even under a page-level retrieval objective (Shorten et al., 5 Feb 2026).
6. Modality-specific strengths, limitations, and significance
IRPAPERS shows that text and image representations are not interchangeable. Text retrieval is stronger on exact lexical constraints, named entities, acronyms, reported metrics, and subtle distinctions embedded directly in prose. BM25 contributes materially in such cases because image encoders do not provide a direct analogue of exact token matching. Image retrieval is stronger when OCR loses structure, when layout matters, or when visual organization provides cues that linearized text omits. The benchmark’s analysis therefore argues for multimodal fusion as a genuine modeling gain rather than as a redundant ensemble (Shorten et al., 5 Feb 2026).
The QA analysis sharpens that distinction. Most figures in the benchmark remain text-accessible because captions and surrounding prose already encode their key claims. For that reason, text-based QA generally outperforms image-based QA even on visually rich pages. Yet the benchmark also identifies genuinely visual cases, most notably a t-SNE figure from HyDE. Under oracle retrieval for a set of questions about geometric relations in that figure, image QA reaches 70% accuracy while text QA reaches 30%. This suggests that some scientific visuals contain irreducible geometric information not recoverable from OCR text alone (Shorten et al., 5 Feb 2026).
The benchmark is also explicit about its limitations. It covers only information retrieval papers, operates at page rather than document or section granularity, and uses a modest set of 180 questions. It relies on GPT-4.1 both for OCR and for the QA judge, and some of its best-performing retrieval baselines are closed-source. Nonetheless, its design yields a targeted testbed for scientific multimodal retrieval: page-level, single-relevance, semantically dense, and directly comparable across text, image, and fused systems (Shorten et al., 5 Feb 2026).
Within the emerging landscape of visual-document evaluation, IRPAPERS occupies a distinctive position. It is not a generic document-understanding benchmark, nor a large-scale web-search benchmark, nor a purely text scientific QA benchmark. Its defining contribution is to make page-image retrieval, OCR-based retrieval, multimodal hybrid search, and image-vs-text RAG directly comparable on tightly related scientific PDFs. The principal conclusion is not that image systems replace text systems, but that multimodal scientific search performs best when it can exploit both representations and when retrieval depth is large enough to supply a reader with multiple supporting pages (Shorten et al., 5 Feb 2026).