---
title: 'MULTI3IR: Multi-domain, Multi-modal Info Retrieval'
url: https://www.emergentmind.com/topics/multi3ir
type: topic
---

# MULTI3IR: Multi-domain, Multi-modal Info Retrieval

MULTI3IR, formally “Multi-perspective Multi-domain Multi-modal Information Retrieval,” is a benchmark for evaluating whether information-retrieval systems cover the distinct facets of open-ended queries across heterogeneous subject domains and document modalities. Introduced in 2026, it contains 104,916 Stack Exchange queries, 521,739 annotated perspectives, and 1,012,061 supporting text and image documents. Its central premise is that retrieval quality should account not only for relevance to a query, but also for comprehensive coverage of the query’s implicit perspectives, domains, and modalities [2608.30949].

## 1. Purpose and conceptual foundations

MULTI3IR targets open-ended information needs for which a single query may encode several complementary questions. For example, the query “Why are so few foods blue?” can require biological, psychological, and chemical perspectives. A retriever that returns many documents about only one of these interpretations may achieve conventional relevance while failing to provide comprehensive evidence.

A query is represented as a set of textual perspective descriptions,

$$
P=\{p_1,p_2,\ldots,p_{|P|}\},
$$

where each perspective is a self-contained, search-friendly description of one facet of the query. Every perspective $p_i$ is associated with a supporting-document set,

$$
D_i=\{d_{i1},d_{i2},\ldots\}.
$$

The principal retrieval objective is to retrieve at least one document supporting every perspective:

$$
D^*(k)\cap D_i\neq\varnothing
\qquad\text{for every }p_i\in P.
$$

This formulation differs from conventional top-$k$ retrieval, where documents are typically evaluated against a single query-level relevance judgment. MULTI3IR instead evaluates whether the retrieved set spans the query’s implicit information structure.

The benchmark combines three types of diversity:

- **Perspective diversity**: coverage of distinct facets or sub-questions.
- **Domain diversity**: coverage across subject areas such as science, law, engineering, psychology, or computer science.
- **Modality diversity**: coverage of evidence represented as text and images.

On average, each query has 4.98 perspectives, 3.34 domains, and 1.91 modalities. The benchmark therefore measures a retrieval capability that is not equivalent to ordinary multimodal search: a system may retrieve both text and images while still concentrating on only one interpretation of the query.

## 2. Dataset composition and taxonomy

The benchmark contains:

- **104,916 open-ended queries**;
- **521,739 annotated perspectives**;
- **1,012,061 supporting documents**;
- documents spanning 94 of a possible 100 subject-domain labels;
- text and image evidence.

The source queries were obtained from official Stack Exchange data dumps hosted on Archive.org. Stack Exchange was selected because its question-and-answer structure supplies multiple naturally occurring responses, community voting and answer acceptance provide quality signals, and answer links can serve as sources for perspective-grounded evidence.

The authors collected posts from 77 Stack Exchange sites and organized them into five broad categories:

| Category | Queries | Percentage |
|---|---:|---:|
| Culture & recreation | 34,776 | 33.5% |
| Science | 20,344 | 19.6% |
| Business | 19,132 | 18.4% |
| Technology | 15,756 | 15.2% |
| Life & arts | 13,869 | 13.4% |

A question was retained only if it had more than three answers and received votes from more than three distinct users. Category-level uniform sampling with per-category caps was applied to avoid overrepresenting technology. For each retained post, the query is formed as the concatenation of its title and body:

$$
Q=\text{title}+\text{body}.
$$

The initial filtering and sampling process reduced approximately 29.8 million raw questions to the final benchmark.

MULTI3IR uses a two-level domain taxonomy consisting of 10 top-level classes and 10 subclasses per class, yielding 100 possible fine-grained domains. The classes include computer science, philosophy and psychology, religion, social sciences, language, science, technology, arts and recreation, literature, and history and geography.

The document collection contains 631,377 text documents, representing 62.4% of the collection, and 380,684 image documents, representing 37.6%. The inclusion of both modalities is intended to test whether retrieval systems can identify heterogeneous evidence sources rather than merely process multiple input formats.

## 3. Perspective and document annotation

Perspective descriptions were extracted from the answer sets using GPT-5-mini. Each perspective was required to be self-contained, search-friendly, specific to one facet, and non-redundant. The extraction output included a perspective description together with supporting facts or quotations.

Two automatic filters were applied.

The **uniqueness filter** encoded perspective descriptions with `all-MiniLM-L6-v2`. A candidate perspective $\hat p_j$ was discarded if its maximum cosine similarity with another perspective exceeded 0.8:

$$
\max_{\hat p\in\hat P\setminus\{\hat p_j\}}
\operatorname{cos}(\hat p_j,\hat p)>0.8.
$$

The threshold was selected through manual inspection.

The **faithfulness filter** used Bespoke-MiniCheck-7B to determine whether an extracted perspective was entailed by the answer set. Questions with fewer than four retained perspectives were removed. The final inventory contains 521,739 perspectives, or 4.97 perspectives per query on average; the principal benchmark summary rounds this value to 4.98.

Supporting documents were retrieved separately for each perspective. Text documents were obtained from C4, while images were retrieved through Google Image Search using the Serper API. Initial retrieval used Qwen3-Embedding-4B. The top 10 text and image documents were combined:

$$
\hat D_i=\hat D_i^{\text{text}}\cup\hat D_i^{\text{image}}.
$$

Text retrieval used the instruction:

> “Instruct: Given a web search query, retrieve relevant passages that answer the query Query:{query}”

Documents were encoded without this prefix.

Because a document may appear relevant to several perspectives belonging to the same query, MULTI3IR applies **exclusive-support verification**. Qwen3-VL-30B-A3B-Instruct was asked whether each document supported its assigned perspective and whether it fully supported any sibling perspective. A document was retained only when it fully supported its assigned perspective and did not fully support a sibling perspective.

The document-processing stages were:

| Stage | Samples | Drop rate |
|---|---:|---:|
| Document candidates | 8,254,298 | — |
| Own-perspective support | 1,503,151 | 81.8% |
| Sibling exclusivity | 1,012,061 | 32.7% |

The resulting collection contains 1,012,061 multimodal supporting documents. Exclusive-support filtering makes perspective labels more discriminative, although it also imposes a strict criterion: a document that supports several perspectives may be excluded from the gold set for each of them.

Human verification was performed by qualified Amazon Mechanical Turk annotators from English-speaking countries who had completed more than 10,000 HITs, maintained an approval rate above 95%, and passed qualification HITs concerning perspective meaningfulness and redundancy.

For perspective verification, annotators judged whether each perspective was meaningful, identified its closest alternative perspective, and assessed whether the pair was redundant or paraphrastic. They judged 99.1% of perspectives to be relevant or meaningful and 96.2% to be unique rather than redundant. Agreement for closest-perspective selection ranged from $\kappa=0.692$ in science to $\kappa=0.740$ in culture.

For document-support verification, annotators rated support for the target perspective and its closest sibling as “Fully supports,” “Partially supports,” or “Does not support.” The overlap subset achieved Fleiss’ $\kappa=0.588$.

The human-verified test split contains approximately 1,000 queries, 4,800 perspectives, and 12,500 supporting documents. The large-scale training and validation setup uses 72,000 training queries and 32,000 validation queries. Main test results are evaluated over the full 1.01-million-document gold pool.

## 4. Retrieval task and evaluation metrics

Given a query $Q$, corpus $C$, perspective set $P$, and supporting sets $D_i$, a conventional retriever $f$ returns:

$$
D^*(k)=\operatorname{Top}\text{-}k_{d\in C} f(Q)^\top f(d).
$$

MULTI3IR evaluates whether this result set covers the perspectives associated with $Q$.

The principal metric is **hard coverage**. A perspective counts as covered only when at least one retrieved document belongs to its annotated supporting set:

$$
\operatorname{HC}@k=
\frac{1}{|P|}
\sum_{i=1}^{|P|}
\mathbf{1}
\left[
D^*(k)\cap D_i\neq\varnothing
\right].
$$

Hard coverage measures the fraction of perspectives for which the retrieval set contains at least one gold supporting document. It is stricter than query-level recall because multiple highly relevant documents for one perspective do not compensate for missing another perspective.

The supplied benchmark description identifies hard coverage as the formal coverage criterion. It also describes a soft-coverage evaluation in which GPT-5-mini judges whether retrieved documents support perspectives, with strong agreement against human-majority judgments: Fleiss’ $\kappa=0.840$, compared with human-human agreement of $0.743$. This suggests that soft evaluation is intended to accommodate semantic support that may not be captured by exact membership in a predefined document set.

MULTI3IR’s evaluation philosophy distinguishes several failure modes:

- **Single-perspective bias**: repeated retrieval of documents addressing the same facet.
- **Domain omission**: failure to retrieve evidence from relevant subject areas.
- **Modality omission**: failure to retrieve required image or text evidence.
- **Redundant retrieval**: high within-perspective relevance but low coverage of the complete perspective set.
- **Evidence misassignment**: retrieval of documents that appear related to the query but do not support the intended perspective.

Consequently, standard metrics such as query-level recall or relevance ranking alone are insufficient to characterize performance on the benchmark.

## 5. Relationship to previous benchmarks

MULTI3IR differs from closed-ended retrieval and question-answering benchmarks in the structure of its information needs. HotpotQA, 2WikiMultihopQA, MultimodalQA, WebQA, InfoSeek, M-BEIR, and related datasets generally evaluate retrieval or answering for a relatively specified information need. Their average domain and modality diversity is lower than that of MULTI3IR.

Prior open-ended benchmarks address portions of the problem. AmbigQA considers multiple interpretations of ambiguous questions but primarily evaluates question answering over Wikipedia. PIR evaluates retrieval for an explicitly specified perspective. BeRDS evaluates whether documents retrieved from an original question cover multiple perspectives, but the perspectives remain implicit to the retriever.

MULTI3IR extends this setting in three ways:

1. Perspectives are implicit and must be inferred from the original query.
2. Supporting evidence spans multiple subject domains.
3. Supporting evidence includes both text and images.

The benchmark thus evaluates:

$$
\text{perspective diversity}
+
\text{domain diversity}
+
\text{modality diversity}.
$$

A system may perform well on a multimodal benchmark by retrieving a relevant image or passage without covering distinct interpretations of the query. MULTI3IR treats modality diversity as one dimension of a broader coverage problem rather than as the sole defining characteristic of multimodal retrieval.

## 6. SPIN and perspective-aware retrieval

The benchmark introduces **SPIN**, or “Steering Perspectives by Injecting Noise,” as a parameter- and label-efficient method for converting a frozen retriever into a perspective-aware multi-vector retriever. SPIN learns noise vectors that steer embeddings toward diverse but meaningful semantic directions.

The method is motivated by the observation that existing multimodal retrievers exhibit single-perspective bias on MULTI3IR. A conventional embedding function maps the original query to one dominant representation, after which top-$k$ retrieval can repeatedly favor the same semantic interpretation. SPIN instead uses learned perturbations to produce multiple retrieval directions associated with distinct perspectives.

The supplied benchmark description establishes SPIN’s purpose and reported qualitative behavior but does not provide the complete optimization equations, architecture, training schedule, or numerical results in the available data. A defensible interpretation is that SPIN operates as a lightweight steering mechanism over a frozen retriever rather than as a fully retrained multimodal encoder. Its parameter and label efficiency are intended to reduce the cost of adapting a pre-existing retrieval model to perspective coverage.

The benchmark reports that SPIN substantially improves perspective coverage on MULTI3IR and generalizes well to unseen open-ended information-retrieval benchmarks. These results position SPIN as a method for mitigating semantic concentration while preserving the underlying retriever.

## 7. Significance, limitations, and research directions

MULTI3IR reframes retrieval evaluation from isolated document relevance to set-level evidence coverage. Its key contribution is the explicit annotation of implicit perspectives and the association of those perspectives with exclusive multimodal supporting documents. The benchmark therefore enables evaluation of whether retrieval systems provide breadth across facets rather than merely depth within one interpretation.

Several limitations follow from the construction and evaluation methodology.

First, perspective extraction depends on GPT-5-mini, while faithfulness and document-support verification depend on additional language models. Although human validation reports high agreement for several components, the benchmark remains influenced by model-generated interpretations and filtering decisions.

Second, exclusive-support verification may remove documents that genuinely support multiple perspectives. This increases label specificity but can underrepresent documents whose value is integrative rather than perspective-exclusive.

Third, image evidence is retrieved through Google Image Search using the Serper API, and the benchmark description does not establish that image semantics, provenance, captions, or visual content are uniformly represented across domains.

Fourth, the domain taxonomy contains 100 possible labels, but documents span 94 of them. Coverage may therefore be uneven across the domain space.

Fifth, hard coverage depends on annotated supporting sets. A document can be substantively useful yet fail hard coverage if it is not included in the corresponding gold set. Soft-coverage evaluation partially addresses this issue but introduces dependence on an evaluation judge.

Finally, the benchmark evaluates retrieval coverage rather than the correctness of a generated answer, the sufficiency of an evidence synthesis, or the factual consistency of downstream reasoning. A retrieval system may cover every annotated perspective while returning documents of unequal quality or failing to integrate them coherently.

These limitations motivate several research directions: perspective discovery directly from queries, uncertainty-aware coverage estimation, domain- and language-adaptive multimodal retrieval, calibrated soft-support judgments, evidence diversity optimization, and retrieval-augmented generation evaluated jointly for coverage and factuality. The benchmark also provides a setting for studying whether dense, sparse, multi-vector, or hybrid retrievers can avoid single-perspective concentration.

MULTI3IR’s principal methodological significance is its separation of three properties that are often conflated: relevance to a query, diversity of retrieved evidence, and coverage of the query’s implicit information structure. By making perspectives explicit in the annotations while leaving them implicit at retrieval time, it evaluates whether a system can infer and satisfy multifaceted information needs rather than merely match the most salient semantic interpretation [2608.30949].

Source: https://www.emergentmind.com/topics/multi3ir