No More Free Lunch: Corpus Task Complexity Matters as Corpora Grow
Abstract: Given a large corpus, the questions one might ask can vary -- from "When was the first human heart transplant?" to "What are all the contradictory claims in this literature?" -- but what makes some questions more challenging than others? In this work, we define a notion of Corpus Task Complexity (CTC) that characterizes tasks by how their difficulty grows with corpus size; for instance, a retrieval query only requires a single linear pass over a corpus, while finding contradictions requires checking a quadratically growing set of claim pairs. Observing that prior work has largely only studied tasks whose difficulty grows linearly with corpus size, which we call low CTC tasks, we introduce 10 new tasks belonging to a class of high CTC whose difficulty grows quadratically or more in corpus size. We find that high-CTC tasks not only grow much more challenging on average at longer contexts for LCLMs, they reverse many modeling conclusions drawn solely from low-CTC evaluations. For instance, efficient block-sparse and hybrid attention approaches consistently match full attention performance on low-CTC tasks, but degrade much more on high-CTC tasks. Large-corpus high-CTC reasoning thus remains an open challenge as full attention is too costly to scale, motivating future research on these tasks. We release our code, data, and 22-task suite (CTC-Bench), to facilitate future research in this area.
Paper Prompts
Sign up for free to create and run prompts on this paper.
Top Community Prompts
Explain it Like I'm 14
1. What is this paper about?
This paper studies how difficult it is for LLMs to answer questions about large collections of documents, such as scientific papers, websites, or book passages.
The authors argue that not all questions become difficult in the same way as the collection gets larger. For example:
- Finding one document that answers a question is relatively simple.
- Finding every pair of documents that contradict each other is much harder because many documents must be compared with one another.
To describe this difference, the authors introduce a new idea called Corpus Task Complexity, or CTC.
Their main message is:
Systems that work well on simple large-corpus tasks may fail badly on more complex tasks, especially when the corpus becomes very large.
2. What questions did the researchers ask?
The researchers wanted to understand several things:
- Why are some large-corpus questions harder than others?
- Does difficulty increase faster for certain types of tasks as the number of documents grows?
- Do LLMs perform worse on complex tasks when they read longer inputs?
- Do faster, cheaper attention methods work as well as regular attention on every kind of task?
- Can models trained on shorter texts successfully handle much longer collections?
The authors were especially interested in whether common tests give an overly positive picture of LLMs because they mostly use simpler tasks.
3. What is Corpus Task Complexity?
Imagine you are looking through a pile of documents.
Low-CTC tasks
A low-CTC task usually needs about one check per document. For example:
“Which document says when Ralph Lauren was founded?”
You can look through the documents one by one. If there are documents, the amount of work grows roughly like .
This is called linear growth, written as .
High-CTC tasks
A high-CTC task requires comparing documents with one another. For example:
“Find every pair of documents that makes contradictory claims.”
With 3 documents, there are only a few pairs to compare. With 100 documents, there are:
possible pairs.
The work grows approximately like , or quadratically. This means that doubling the number of documents can require about four times as many comparisons.
Some tasks can be even harder. For example, checking groups of three documents can require work growing like .
| Task type | Example | How the work grows |
|---|---|---|
| Low CTC | Find a document containing an answer | |
| High CTC | Find all contradictory document pairs | |
| Very high CTC | Find groups of three documents with a shared property |
The authors compare this to searching for a friend at school. Finding one person might mean checking students one by one. Finding every pair of students who disagree about something requires comparing many pairs of students.
4. How did the researchers conduct the study?
Building CTC-BENCH
The researchers created a collection of 22 tasks, called CTC-BENCH:
- 12 were low-CTC tasks.
- 10 were high-CTC tasks.
The tasks included:
- Finding relevant passages
- Answering questions using documents
- Finding contradictions
- Finding missing text
- Grouping documents by topic
- Finding unusual or rare topics
- Reordering shuffled text
- Finding groups of documents with matching features
This gave the researchers a way to compare simple and complex tasks fairly.
Testing LLMs
They trained LLMs to solve these tasks using input lengths from about:
- 2,000 tokens
- 4,000 tokens
- 8,000 tokens
- 16,000 tokens
- 32,000 tokens
A token is a small piece of text, such as a word or part of a word. More tokens mean a longer collection of documents.
The researchers mainly used a model called Qwen3.5-4B, but they also tested other model families and sizes to see whether the results were consistent.
Comparing attention methods
LLMs use a process called attention to decide which parts of the input are important.
The study compared:
- Full attention: Every part of the input can directly interact with every other part. This can be powerful but becomes very expensive for long inputs.
- Block-sparse attention: Documents mostly examine themselves, while only a few parts communicate across documents. This is faster and cheaper.
- Hybrid models: These combine full attention with faster methods that use less memory and computation.
The researchers also studied length generalization. This means training a model on shorter inputs and then testing whether it can handle much longer inputs.
5. What did the researchers find?
High-CTC tasks became much harder as documents increased
The main result was that high-CTC tasks lost performance much faster than low-CTC tasks.
From 2,000 to 32,000 tokens:
- Low-CTC performance fell from about 91% to 79%.
- High-CTC performance fell from about 85% to 51%.
So even when models were trained using examples of different lengths, tasks involving many comparisons became much harder as the corpus grew.
Faster attention worked well only on simpler tasks
Block-sparse attention performed almost as well as full attention on low-CTC tasks such as retrieval and question answering.
However, on high-CTC tasks, its performance dropped significantly. At 32,000 tokens, block-sparse attention was about 66% worse than full attention on average for the high-CTC tasks.
This is important because earlier research might suggest that block-sparse attention is a “free lunch”: it is cheaper but seems to work just as well. The new study shows that this is true mainly for simpler tasks.
Hybrid models also struggled with high-CTC tasks
Hybrid models were often similar to full-attention models on low-CTC tasks.
But on high-CTC tasks, hybrid models usually performed worse. This suggests that reducing communication between different parts of the input can remove information needed for comparing many documents.
Training on short inputs did not always prepare models for long inputs
The researchers tested models trained up to 32,000 tokens on inputs as long as 128,000 tokens.
At that length:
- Low-CTC tasks kept between about 28% and 100% of their 32,000-token performance.
- High-CTC tasks kept only about 4% to 14%.
This shows that high-CTC tasks may be especially poor at handling inputs longer than those seen during training.
Bigger models helped, but did not solve everything
The general pattern appeared across different model sizes and model families.
Smaller models usually lost performance more quickly as the corpus grew. However, even larger models still faced serious problems on high-CTC tasks.
The study also found that tasks with more difficult basic operations became even more challenging when their complexity grew from to .
6. Why are these results important?
Many language-model evaluations focus on tasks such as:
- Finding one relevant passage
- Answering a question from a document
- Retrieving a small number of supporting documents
These are mostly low-CTC tasks. They are useful, but they do not test whether a model can carefully compare every part of a large collection.
The paper shows that a model may appear excellent on common tests but still struggle to:
- Find all contradictions in a scientific literature collection
- Discover unusual topics
- Match many questions with their correct documents
- Reconstruct the order of shuffled passages
- Compare every document with many others
In other words, there is no single method that is both extremely cheap and equally powerful for every kind of task. The title, No More Free Lunch, refers to this trade-off.
Full attention can handle more detailed interactions, but its cost grows very quickly. Faster methods reduce the cost, but they may lose important information needed for high-CTC reasoning.
7. Possible impact of the research
The paper suggests that future language-model research should test more than simple retrieval and question answering.
The released CTC-BENCH benchmark can help researchers measure whether new systems can handle difficult large-corpus tasks. It may encourage the development of models that can:
- Compare documents efficiently
- Preserve important connections across a large corpus
- Solve complex tasks without using impossibly large amounts of computing power
- Work reliably on collections much larger than those used during training
This could be useful in areas such as scientific research, law, medicine, journalism, and internet search.
The biggest challenge is finding a system that combines the strengths of both approaches:
- The power of full attention for detailed comparisons
- The low cost of efficient attention for very large collections
The paper concludes that solving high-CTC tasks at large scale remains an open problem. It also warns researchers not to assume that success on simple long-context tasks means a model can truly understand and analyze a large collection of documents.
Knowledge Gaps
Knowledge gaps, limitations, and open questions
The paper leaves the following issues unresolved:
- CTC is defined in terms of oracle-call counts rather than end-to-end computational cost. The framework does not establish how oracle-call complexity maps to actual latency, FLOPs, memory use, energy consumption, communication overhead, or monetary cost for modern LCLMs.
- The claimed optimality of the proposed algorithms is largely unproven. Several tasks are assigned or higher CTC using brute-force candidate algorithms, but the paper does not prove that substantially cheaper algorithms are impossible when indexing, embeddings, approximate search, document structure, or task-specific preprocessing are allowed.
- The assumption that corpora are unstructured is restrictive. It remains unclear how CTC changes when documents have metadata, timestamps, citations, hyperlinks, known topic labels, duplicate structure, or other exploitable organization.
- The relationship between CTC and empirical difficulty is not formally characterized. The experiments show correlations between higher CTC and performance degradation, but do not establish a predictive law connecting asymptotic complexity, corpus size, oracle difficulty, and model accuracy.
- The boundaries between CTC classes are underspecified. Tasks with sublinear indexing, preprocessing, fixed numbers of relevant documents, approximate answers, or data-dependent rather than worst-case complexity are not systematically classified.
- The benchmark’s high-CTC tasks are mostly synthetic or semi-synthetic. Contradictions, shuffled corpora, word-sequence matching, artificial grouping, and constrained triple-search tasks may not represent the noise, ambiguity, redundancy, and evolving schemas of real scientific, legal, enterprise, or web corpora.
- The validity of automatically generated labels is uncertain. The paper does not provide extensive human validation of LLM-generated contradictions, topic assignments, matching pairs, outliers, or triple-based answers, leaving open whether benchmark performance reflects reasoning ability or artifacts in the data-generation process.
- The benchmark has limited scale. Experiments reach 32K tokens for the main comparisons and 128K for length generalization, far below the million-token or larger corpora that motivate the paper. It is unresolved whether the observed trends remain monotonic at realistic corpus scales.
- The evaluation covers only a narrow range of model families and sizes. Most experiments use Qwen3.5-4B, with limited comparisons involving OLMo, Llama, and a few smaller models; the conclusions may differ for frontier-scale models, multimodal models, retrieval-augmented systems, or models specifically trained for long-context reasoning.
- Training and evaluation data are largely in-distribution. Models are fine-tuned separately on 20,000 examples per task with training context lengths matching evaluation lengths, so the findings do not establish how high-CTC tasks behave under zero-shot, few-shot, cross-domain, or realistic distribution-shift conditions.
- The study does not isolate memorization from reasoning. Because some corpora and task templates are derived from established datasets or repeated document sources, it remains unclear how much performance depends on memorized facts, formatting regularities, or source familiarity.
- The comparison of attention architectures is not fully controlled. Full, block-sparse, and hybrid models differ in pretraining, parameterization, training procedures, masks, and continued-pretraining exposure, making it difficult to attribute performance gaps solely to the attention pattern.
- The mask-mixing intervention is not comprehensively ablated. The paper reports a curriculum from high to zero full-attention probability, but does not determine which mixing schedules, probabilities, or training budgets are optimal across CTC classes and corpus sizes.
- The benefits of block-sparse attention are evaluated primarily under a fixed document-level partition. More flexible sparse patterns—such as learned routing, hierarchical attention, retrieval-conditioned sparsity, global memory tokens, or adaptive pair selection—are not tested.
- The paper does not compare against strong non-LCLM pipelines. Classical information retrieval, dense retrieval, clustering, graph algorithms, database joins, approximate nearest-neighbor methods, and agentic or tool-using systems may solve several high-CTC tasks more efficiently, but their accuracy–cost trade-offs are not evaluated.
- Approximate solutions are not addressed. Many high-CTC tasks may not require exhaustive enumeration; approximate contradiction discovery, top- matching, probabilistic outlier detection, and candidate generation followed by verification could offer useful recall–cost trade-offs that the benchmark does not measure.
- Evaluation metrics may inadequately capture set-valued answers. The paper uses measures such as set-F1, pair-F1, and partial credit, but does not analyze how these metrics penalize missed items, reward duplicates, handle uncertainty, or reflect the practical usefulness of incomplete results.
- The task difficulty caused by output length is not separated from reasoning difficulty. High-CTC tasks often require producing many pairs, groups, or identifiers, so performance degradation may partly result from long decoding sequences, formatting errors, or output-token limits rather than inability to perform corpus-wide comparisons.
- The impact of answer sparsity is not systematically controlled. High-CTC tasks differ in the number and distribution of correct pairs or items; the study does not determine whether degradation is driven by CTC itself or by rare positives, dense answer sets, class imbalance, or changing output cardinality.
- The interaction between CTC and corpus redundancy is unexplored. Duplicate, near-duplicate, highly correlated, or clustered documents could reduce the effective search space, but the benchmark does not quantify or manipulate effective corpus size separately from token count.
- The role of document length and token-level complexity is unclear. Experiments vary context length, but do not disentangle the number of documents, document length, number of claims, number of candidate relations, and total token count.
- The generality of the oracle-difficulty analysis is limited. The paper compares a small number of paired tasks and measures oracle difficulty using one model’s performance; it does not establish whether the observed interaction generalizes across oracle definitions, models, domains, or human judgments.
- The paper does not identify the point at which preprocessing becomes worthwhile. It remains open when the cost of building indexes, embeddings, clusters, graphs, or summaries is offset by reduced high-CTC inference cost, particularly under changing corpora and repeated queries.
- Robustness to noisy, adversarial, or conflicting documents is not evaluated. Real corpora may contain misinformation, ambiguous claims, contradictory annotations, adversarial distractors, or distributional shifts that could alter both CTC and architectural rankings.
- The practical quality requirements for high-CTC applications are unspecified. The study does not determine acceptable recall, precision, calibration, abstention, provenance, or human-review costs for applications such as scientific contradiction discovery or legal corpus analysis.
- The benchmark does not test iterative or interactive workflows. Systems that progressively retrieve, cluster, verify, and revise candidate relations may have very different scaling behavior from one-shot LCLM inference, but such workflows are outside the evaluation.
- The theoretical implications for scalable architectures remain open. The paper demonstrates that current sparse and hybrid methods lose accuracy on high-CTC tasks, but does not propose or validate an architecture that achieves near-full-attention quality with subquadratic cost.
- It is unclear whether high-CTC performance can be improved through training alone. The study does not determine whether specialized objectives, synthetic pairwise supervision, curriculum learning, intermediate representations, explicit relational memory, or tool-use training can mitigate the observed scaling failures.
- The effect of corpus ordering is insufficiently explored. Although some tasks shuffle documents, the paper does not systematically test adversarial, topical, temporal, random, or relevance-based ordering and its interaction with positional biases and sparse attention.
- Reproducibility and benchmark stability may be affected by generated data and evolving model versions. The paper does not report how sensitive results are to random seeds, label-generation prompts, corpus construction choices, evaluator models, or future revisions of the underlying datasets.
Practical Applications
Immediate Applications
- Benchmarking and procurement for long-context AI systems (software, enterprise AI, academia)
Organizations can use
CTC-BENCHand the paper’s distinction between low- and high-CTC tasks to evaluate models before deployment. Procurement tests should include not only retrieval and question answering, but also contradiction detection, document matching, grouping, reordering, and outlier discovery at increasing corpus sizes. Potential workflow: evaluate candidate models with both full-attention and efficient-attention configurations at the organization’s expected context length, then select architectures based on task-specific accuracy–cost trade-offs. Dependencies: the benchmark must be extended with domain-specific data; synthetic tasks may not fully represent real-world distributions, terminology, noise, or document structure. - Architecture selection for retrieval-augmented generation (RAG) (software, search, customer support, legal technology) Teams can classify a planned RAG workload by CTC before choosing an inference architecture. Conventional retrieval and fixed-hop question answering are generally low-CTC and may use block-sparse, hybrid, or linear-time components. Workflows requiring cross-document comparison—such as finding all inconsistent policies or matching many questions to many documents—should not assume that efficient attention is lossless. Potential product: a CTC-aware RAG router that directs simple queries to inexpensive sparse or hybrid models and complex corpus-wide analyses to more expensive full-attention or multi-stage pipelines. Dependencies: reliable identification of the task’s true computational structure and careful validation at production corpus sizes.
- Large-scale contradiction and consistency screening (healthcare, legal, finance, compliance) Organizations can apply high-CTC task definitions to identify contradiction-search workloads in clinical literature, contracts, financial disclosures, internal policies, and regulatory documents. A practical near-term system would use retrieval or clustering to generate candidate document pairs, followed by an entailment or contradiction classifier, rather than exhaustively comparing every pair. Potential tools: literature-monitoring dashboards, contract inconsistency detectors, policy-diff systems, and financial-disclosure review assistants. Dependencies: candidate-generation recall is critical; missed pairs may be more damaging than false positives. Human review and domain-specific validation remain necessary, especially in healthcare and legal settings.
- Evaluation of long-context model compression and efficiency techniques (AI infrastructure, cloud computing) Researchers and engineering teams can use high-CTC tests to detect failures hidden by standard retrieval benchmarks. Block-sparse attention, hybrid state-space/attention models, context extension, and length-generalization methods should be evaluated on pairwise and higher-order reasoning before being adopted for corpus-analysis products. Actionable practice: add performance-versus-context curves and full-attention baselines to model release evaluations, rather than reporting only aggregate accuracy at one context length. Dependencies: full-attention baselines may be expensive; reproducible hardware, tokenization, training data, and evaluation protocols are needed for fair comparisons.
- Curriculum and training design using mask-mixing (model training, enterprise fine-tuning) The reported mask-mixing method—occasionally training block-sparse models with full-attention masks—can be tested as a practical intervention for improving efficient models, particularly on aggregation and high-CTC tasks. Potential workflow: begin fine-tuning with a substantial proportion of full-attention batches, gradually anneal that proportion, and evaluate whether the resulting model retains more high-order reasoning ability under block-sparse inference. Dependencies: the paper reports results for particular model families and task settings; gains may vary with model size, mask schedule, data quality, and compute budget.
- Scientific literature analysis and evidence review (academia, healthcare, research policy) Research groups can use the benchmark’s task taxonomy to separate ordinary evidence retrieval from more demanding corpus reasoning. Immediate applications include detecting conflicting claims, identifying unmatched documents between literature versions, grouping abstracts by topic, and locating rare or anomalous findings. Potential product: a literature-review assistant that presents candidate contradictions and clusters together with source passages and confidence scores, rather than producing an unsupported summary. Dependencies: LLM-generated labels and contradiction judgments can contain errors; results should be treated as screening aids and verified by subject-matter experts.
- Quality assurance for document collections (publishing, government records, enterprise knowledge management) The X-Absence and reordering task structures suggest practical checks for missing, duplicated, shuffled, or altered content across document repositories, backups, versions, and data migrations. Low-CTC aligned comparisons can be deployed now, while shuffled or unaligned corpus comparison should use candidate matching and verification stages. Dependencies: document segmentation, version alignment, OCR quality, and access to trusted reference copies determine reliability.
- Daily-life tools for comparing and organizing personal documents (consumer software, education) A document assistant could compare multiple versions of leases, insurance policies, school materials, or user manuals, flag potentially contradictory clauses, group related passages, and identify missing sections. For ordinary retrieval and summarization, current efficient long-context systems may be adequate; corpus-wide comparison should expose uncertainty and avoid presenting automated findings as definitive. Dependencies: privacy-preserving local processing, secure storage, accurate OCR, and safeguards against false legal or medical interpretations.
Long-Term Applications
- Scalable contradiction engines for entire scientific and regulatory corpora (healthcare, law, science policy) A mature system could continuously compare claims across millions of papers, clinical guidelines, patents, court decisions, or regulations and construct a graph of agreement, contradiction, qualification, and temporal change. This would support evidence synthesis, policy harmonization, and research-gap discovery. Why long-term: exhaustive pairwise reasoning has quadratic or higher CTC, while full attention becomes computationally impractical at large corpus sizes. Progress likely requires hierarchical indexing, learned candidate generation with recall guarantees, parallel pairwise reasoning, and architectures that preserve cross-document interactions. Dependencies: trustworthy contradiction definitions, temporal and source reliability modeling, multilingual support, and expert adjudication.
- High-CTC scientific discovery and anomaly detection (biotechnology, materials science, climate research) Systems could identify rare topics, outlier experimental results, unusual combinations of observations, and groups of documents satisfying multi-document constraints. Such tools might reveal overlooked research directions or inconsistent experimental findings. Why long-term: outlier detection with a growing number of latent categories and higher-order grouping becomes substantially harder as the corpus expands. Dependencies: well-structured metadata, robust representations, control of publication bias, and mechanisms for distinguishing genuine novelty from noise or poor-quality reporting.
- Corpus-scale legal and compliance reasoning (legal technology, finance, public administration) Future systems could compare every relevant contract clause, statute, regulatory interpretation, and corporate disclosure to identify conflicts, obligations, duplicated provisions, and missing evidence. This could support due diligence, regulatory change management, and automated compliance audits. Why long-term: legal relations are often pairwise or multi-document, and errors have high consequences; efficient attention methods that perform well on retrieval may fail on comprehensive comparison. Dependencies: jurisdiction-specific ontologies, citation and provenance tracking, explainability, audit logs, and mandatory human approval for consequential decisions.
- Corpus-aware autonomous research agents (AI agents, robotics, knowledge work) Research agents could move beyond retrieving a few passages to constructing and testing relationships across a complete corpus: matching questions to all relevant documents, reconciling competing evidence, ordering events, and generating structured knowledge graphs. Why long-term: the paper notes that autoregressive agentic enumeration can require on the order of outputs, making naive multi-agent workflows too slow and expensive. Dependencies: efficient parallel inference, compact intermediate representations, reliable planning, and protection against cascading errors in automatically generated comparisons.
- New attention architectures with adaptive CTC computation (AI research, hardware, cloud infrastructure) The findings motivate models that allocate computation according to task complexity: linear or block-sparse processing for independent retrieval, denser cross-document interaction for pairwise tasks, and specialized mechanisms for triple or higher-order relations. Potential innovation: an adaptive model that estimates whether a query is , , or higher and dynamically activates global attention, pairwise comparison modules, or external indexes. Dependencies: theoretical guarantees, efficient hardware kernels, training objectives that reward global consistency, and methods for avoiding approximate-search recall loss.
- CTC-aware data structures and hybrid indexing systems (databases, search, enterprise software) Database-like systems could combine inverted indexes, dense retrieval, clustering, locality-sensitive hashing, graph search, and targeted language-model verification. Rather than asking an LLM to compare all document pairs, the system would narrow the candidate space while tracking uncertainty and recall. Why long-term: high-CTC tasks require more than simply extending the context window; they require algorithmic decomposition and potentially task-specific indexes. Dependencies: assumptions about corpus structure, stable embeddings, domain drift monitoring, and empirical guarantees that pruning does not remove important relationships.
- Policy standards for evaluating AI systems on corpus-scale reasoning (government, standards bodies, academia) Evaluation standards could require reporting performance by CTC class, corpus size, model scale, attention mechanism, inference cost, and length generalization. This would discourage claims based solely on low-CTC retrieval benchmarks and provide regulators with more realistic evidence about model capabilities. Dependencies: benchmark validity, representative real-world datasets, protection of sensitive corpora, and agreement on metrics for sets of contradictions, clusters, matches, and higher-order outputs.
- Large-scale educational and archival knowledge systems (education, libraries, digital humanities) Future systems could compare textbooks, lecture materials, historical records, and archival collections to identify conflicting accounts, missing passages, evolving terminology, and relationships among sources. Students could receive evidence maps rather than single-answer summaries, while librarians could use automated corpus quality checks. Dependencies: provenance preservation, intellectual-property permissions, transparent uncertainty estimates, and careful handling of contested or culturally sensitive interpretations.
- High-assurance decision-support systems in healthcare and finance (healthcare, insurance, finance) At sufficient maturity, high-CTC models could continuously reconcile patient guidelines, medical evidence, claims records, risk reports, or market disclosures. They could surface inconsistencies for professional review rather than autonomously making decisions. Why long-term: the paper demonstrates that high-CTC performance degrades substantially with context length even under in-distribution training, and smaller models degrade faster under efficient attention. Dependencies: validated domain models, privacy and security controls, regulatory approval, calibrated uncertainty, bias testing, and clear limits on automation.
Glossary
- Agentic scaffold: A system architecture in which an AI model coordinates tools or subtasks to solve a problem. “One path is to develop agentic scaffolds that combine efficient tools like dense retrievers”
- Asymptotic growth: The rate at which a quantity increases as the input size becomes large, usually expressed with Big-O notation. “CTC is the asymptotic growth in oracle calls needed to solve tasks as a function of corpus size N.”
- Block-sparse attention: An attention mechanism that restricts interactions to selected blocks rather than allowing every token to attend to every other token. “block-sparse attention: pre-filling each document independently so that they attend to tokens within the same document only”
- Corpus reasoning task: A task that requires extracting, comparing, aggregating, or otherwise reasoning over a collection of documents. “We refer to these collectively as corpus reasoning tasks”
- Corpus Task Complexity (CTC): A measure of task difficulty based on how the required number of oracle operations scales with corpus size. “We begin by defining corpus task complexity (CTC), which characterizes a task by how the number of operations required to solve it scales with corpus size.”
- Cross-document interaction: Reasoning that requires comparing or relating information across multiple documents. “tasks requiring extensive cross-document interaction”
- Curriculum: A training strategy that gradually changes the difficulty or composition of training examples or conditions. “we anneal on a curriculum from p = 0.8 to p = 0 during training”
- Dense retriever: An information-retrieval model that represents queries and documents as dense vectors and retrieves items according to vector similarity. “agentic scaffolds that combine efficient tools like dense retrievers”
- Distributional question: A question about the frequency, distribution, or statistical pattern of labels or properties in a corpus. “OOLONG[Bertsch et al., 2025] involves distributional or counting questions about class labels of documents.”
- End-to-end: Describing a system that processes input through the complete pipeline in one integrated operation. “architectures capable of ingesting full corpora end-to-end”
- Full attention: An attention mechanism in which every input token can attend to every other input token. “In a regular Transformer model (full attention), pre-filling a corpus has quadratic computational complexity.”
- Gated delta network (GDN): A recurrent sequence-modeling architecture that uses gated state updates based on a delta rule to reduce computational cost. “adjacent approaches like gated delta net (GDN)[Yang et al., 2024] have also enabled attention to reduce computational complexity.”
- Hybrid architecture: A model architecture that combines different sequence-processing mechanisms, such as recurrent layers and full-attention layers. “hybrid architectures that alternate linear and full attention exhibit larger performance gaps on high-CTC tasks”
- In-distribution training: Training and evaluation in which the data conditions, such as task type and context lengths, are matched. “We train Qwen3.5-4B individually for all 22 CTC-BENCH tasks, ranging from 2k to 32k context length in training and evaluation.”
- Information retrieval (IR): The computational task of finding documents or passages relevant to a query. “Information retrieval is arguably the best-studied corpus reasoning task.”
- Length generalization: The ability of a model trained on shorter inputs to perform effectively on longer inputs. “length generalization—fine-tuning a long-context model on shorter contexts than the intended context length at test time”
- Linear attention: An attention formulation whose computational cost grows linearly rather than quadratically with sequence length. “these methods replace all-to-all token interactions with a recurrent state update, giving linear computation and constant memory.”
- Long-context LLM (LCLM): A LLM designed to process unusually long token sequences or contexts. “Long-context LLMs (LCLMs) trained to process large inputs”
- Mask-mixing: A training method that alternates between full-attention and block-sparse attention masks. “we additionally propose a new mask-mixing technique”
- Multi-hop question answering: Question answering that requires retrieving and combining information through multiple reasoning or retrieval steps. “HotpotQA [Yang et al., 2018], a multi-hop task.”
- Oracle operation: A basic information-processing query, often made to a language-model judge, used as an assumed unit of computation in the complexity analysis. “CTC is the asymptotic growth in oracle calls needed to solve tasks as a function of corpus size N.”
- Pairwise comparison: An operation that evaluates relationships between every possible pair of items. “exhaustively checking all possible pairs will generally require quadratic oracle calls in corpus size”
- Quadratic complexity: Computational complexity whose cost grows proportionally to the square of the input size, expressed as . “finding contradictions requires checking a quadratically growing set of claim pairs.”
- Recurrent state update: A sequential operation that updates a model’s hidden state as it processes each new token or input. “these methods replace all-to-all token interactions with a recurrent state update”
- Retrieval-augmented generation (RAG): A method that retrieves relevant documents and supplies them to a generative LLM to produce an answer. “Accelerating retrieval-augmented generation with precomputed kv caches for chunked text.”
- Sliding window attention: An attention mechanism in which each token attends only to nearby tokens within a fixed-size window. “since Olmo3 by default uses sliding window attention”
- Softmax attention: The standard Transformer attention mechanism that converts attention scores into normalized weights using the softmax function. “with softmax attention enabling direct retrieval and reasoning across parts of the conditioned corpus.”
- State-space model (SSM): A sequence model that represents evolving inputs through a latent dynamical state, often enabling efficient long-sequence processing. “alternative architectures such as state-space-models (SSMs)”
- Synthetic benchmark: An evaluation dataset or task generated artificially to control its properties and test specific capabilities. “We often include synthetic / semi-synthetic tasks for controllability and cleaner analysis”
- Token interaction: The computational relationship in which one token attends to or incorporates information from another token. “these methods replace all-to-all token interactions with a recurrent state update”
- Transformer: A neural-network architecture based primarily on attention mechanisms for processing sequences. “In a regular Transformer model (full attention), pre-filling a corpus has quadratic computational complexity.”
- Unstructured corpus: A collection of documents without a predefined relational or tabular organization. “Define a corpus C to be an unstructured set of N documents”
- Zero-shot evaluation: Evaluation in which a model performs a task without task-specific examples in the prompt or additional task training. “BEIR [Thakur et al., 2021], a heterogenous benchmark for zero-shot evaluation of information retrieval models.”