Papers
Topics
Authors
Recent
Search
2000 character limit reached

No More Free Lunch: Corpus Task Complexity Matters as Corpora Grow

Published 24 Sep 2026 in cs.CL and cs.AI | (2609.29245v1)

Abstract: Given a large corpus, the questions one might ask can vary -- from "When was the first human heart transplant?" to "What are all the contradictory claims in this literature?" -- but what makes some questions more challenging than others? In this work, we define a notion of Corpus Task Complexity (CTC) that characterizes tasks by how their difficulty grows with corpus size; for instance, a retrieval query only requires a single linear pass over a corpus, while finding contradictions requires checking a quadratically growing set of claim pairs. Observing that prior work has largely only studied tasks whose difficulty grows linearly with corpus size, which we call low CTC tasks, we introduce 10 new tasks belonging to a class of high CTC whose difficulty grows quadratically or more in corpus size. We find that high-CTC tasks not only grow much more challenging on average at longer contexts for LCLMs, they reverse many modeling conclusions drawn solely from low-CTC evaluations. For instance, efficient block-sparse and hybrid attention approaches consistently match full attention performance on low-CTC tasks, but degrade much more on high-CTC tasks. Large-corpus high-CTC reasoning thus remains an open challenge as full attention is too costly to scale, motivating future research on these tasks. We release our code, data, and 22-task suite (CTC-Bench), to facilitate future research in this area.

Summary

  • The paper introduces Corpus Task Complexity (CTC), a measure defining how computational requirements scale with corpus size, using oracle calls to characterize task complexity with asymptotic scaling analysis.
  • Experimental results show that high-CTC tasks, such as contradiction detection and query-document matching, degrade more rapidly with increasing context length than low-CTC tasks like retrieval, highlighting the importance of task complexity in model performance.
  • The study finds that block-sparse attention, which is effective for low-CTC tasks, fails to maintain performance on high-CTC tasks, indicating that task-specific attention mechanisms are crucial for handling complex corpus interactions.

Conceptual contribution

“No More Free Lunch: Corpus Task Complexity Matters as Corpora Grow” (2609.29245) introduces Corpus Task Complexity (CTC), a complexity measure for corpus reasoning tasks based on how the number of required inference operations scales with corpus size. The paper’s central claim is that long-context evaluations have systematically overrepresented tasks whose computational requirements grow linearly with the number of documents, thereby obscuring a qualitatively different class of tasks requiring higher-order interactions among corpus elements.

The distinction is operationalized through oracle calls. Given a corpus of NN documents, an oracle can perform a constant-size operation, such as determining whether a document is relevant to a query or whether two claims contradict one another. The CTC of a task is then characterized by the asymptotic number of oracle calls required by the best available algorithm over an unstructured corpus. A conventional retrieval task can be solved with one pass over the documents and is therefore classified as OT(N)O_T(N). By contrast, identifying all contradictory claim pairs requires considering document pairs and is classified as OT(N2)O_T(N^2). Finding triples satisfying a joint constraint provides an OT(N3)O_T(N^3) example.

This definition concerns the scaling of the task-level computation, not the intrinsic difficulty of each elementary operation. Two tasks may both be OT(N)O_T(N) while differing substantially in the semantic difficulty of their per-document judgments. Conversely, two OT(N2)O_T(N^2) tasks may differ in the difficulty of the pairwise relation being evaluated. The paper explicitly treats these factors as distinct, and later experiments show that they interact: higher base-operation difficulty amplifies the performance penalty associated with higher CTC.

The CTC framework is intended as a structural property of corpus tasks rather than a complete theory of model difficulty. The authors acknowledge that proving an optimal algorithm is generally difficult, so their empirical classifications use plausible best-known algorithms for unstructured corpora. This leaves open whether some tasks classified as quadratic admit practical subquadratic algorithms under realistic assumptions about document structure, metadata, semantic locality, or indexability.

CTC-BENCH and task taxonomy

The paper evaluates the framework through CTC-BENCH, a 22-task suite comprising 12 low-CTC tasks drawn from common long-context evaluations and 10 newly constructed higher-CTC tasks. The benchmark spans retrieval, aggregation, outlier detection, cross-corpus comparison, clustering, pair discovery, reordering, contradiction detection, and higher-order combinatorial search.

CTC class Representative tasks Required operation pattern
OT(N)O_T(N) NQ, HotpotQA, SciFact, FiQA, MS MARCO, OOLONG One corpus pass or a constant number of passes
OT(NM)O_T(NM) Wikipedia outlier detection, OpenAlex grouping Compare documents against a growing set of categories or groups
OT(N2)O_T(N^2) Contradiction, X-Absence, QDmatch, Strmatch, Reorder Search over document or cross-corpus pairs
OT(N3)O_T(N^3) Textgroups Search over document triples satisfying a joint constraint

The low-CTC category includes standard retrieval tasks such as NQ, HotpotQA, SciFact, FiQA, MS MARCO, and OBLIQ. The authors also classify multi-hop HotpotQA as OT(N)O_T(N)0 because the task requires only a fixed number of retrieval passes, despite requiring composition across two documents. Similarly, OOLONG and fixed-label outlier detection remain linear because each item can be classified independently before a lightweight aggregation step.

The benchmark’s higher-CTC tasks are designed to isolate the cost of cross-document interaction. Contradiction requires finding all contradictory claim pairs in a PubMed-derived corpus. X-Absence compares two shuffled, nearly identical corpora and identifies unmatched elements, removing the positional alignment that makes ordinary absence detection linear. Query-document matching constructs separate query and document pools and asks the model to identify a sparse set of relevant pairs among OT(N)O_T(N)1 possible combinations. Strmatch is a synthetic control in which pairwise matching depends only on shared contiguous word sequences, allowing the effect of comparison count to be studied independently of semantic reasoning.

Other tasks test growing category structure. Wikipedia outlier detection has a fixed-OT(N)O_T(N)2 condition, which remains OT(N)O_T(N)3, and a scale-OT(N)O_T(N)4 condition in which the number of source topics grows with corpus size, yielding OT(N)O_T(N)5. OpenAlex grouping similarly requires assigning documents to a growing number of latent topic groups. Reorder requires recovering the original sequence of shuffled text segments; absent exploitable positional structure, the proposed oracle formulation requires pairwise comparisons. Textgroups extends the suite beyond quadratic complexity by asking for all document triples whose lexical feature counts sum to a target.

This construction is an important methodological contribution because it distinguishes corpus length from corpus interaction structure. A long retrieval context may contain many documents but still require only independent document-query judgments. A shorter corpus can be computationally more demanding if its answer depends on examining a large set of pairwise or higher-order relations. The benchmark therefore tests whether models can maintain relational coverage as the number of candidate interactions increases, rather than merely whether they can locate a salient item in a long sequence.

Experimental design

The principal experiments fine-tune Qwen3.5-4B separately on the 22 tasks, using context lengths from 2K through 32K tokens. The models are trained in-domain to reduce the confounding effects of task unfamiliarity and distribution shift. Except for task-specific deviations, each model receives 20,000 examples over one epoch, distributed evenly across the five context budgets. Full attention, block-sparse attention, and mask-mixed block-sparse attention are compared under matched training conditions.

The setup is deliberately designed to test architectural scaling rather than zero-shot competence. Nevertheless, the in-domain design imposes an important interpretation constraint: the reported degradation reflects limitations in representing and computing the required corpus interactions, but not necessarily the performance of a model trained with substantially more data, more sophisticated curricula, or task-specific decomposition strategies.

Evaluation uses task-appropriate metrics, including set-F1, pair-F1, gold-document F1, Kendall’s tau, MRR, and partial-credit aggregation scores. The benchmark includes both real and synthetic or semi-synthetic data. This increases experimental control, particularly for exact CTC assignments, but also means that some tasks may not capture the distributional complexity of naturally occurring corpus analysis.

Performance scaling with corpus size

The central empirical result is that high-CTC tasks deteriorate substantially faster than low-CTC tasks as context length increases, even when models are trained at the evaluated context lengths. Under full attention, the average low-CTC score falls from 0.910 at 2K tokens to 0.788 at 32K, a 13% relative decrease. In contrast, the average high-CTC score falls from 0.851 to 0.511, a 40% relative decrease. Ten low-CTC tasks are reported to degrade by less than 0.1, whereas seven of the ten high-CTC tasks degrade by more than 0.1.

These results support the paper’s primary claim that context length alone is an inadequate predictor of long-context difficulty. The relevant variable is the number and structure of corpus-level interactions that must be resolved. A model can preserve retrieval performance while failing to maintain reliable coverage over a quadratic candidate space. The implication is direct: evaluations restricted to linear-scan tasks can substantially overestimate the scalability of long-context models for corpus analysis.

The contrast is also visible in matched task pairs. NIAH-contradiction is constructed as a linear-time control for contradiction search: the query directly specifies one claim, and the model must retrieve its contradictory document. The corresponding contradiction task requires finding all contradictory pairs among the corpus. Similarly, ordinary HotpotQA retrieval is compared with QDmatch constructed from HotpotQA questions and documents. These pairings isolate the effect of all-pairs search from the underlying semantic operation. The high-CTC versions degrade more rapidly, indicating that the deterioration cannot be attributed solely to harder document content.

The paper further finds that CTC interacts with oracle difficulty. Among matched OT(N)O_T(N)6 and OT(N)O_T(N)7 tasks, the relative performance gap grows more quickly when the underlying operation is semantically harder. Thus, CTC is not merely an additive computational burden. As the candidate interaction space expands, errors in the elementary relation judgment become increasingly consequential, and the model must sustain precision over a larger number of potential comparisons.

Block-sparse attention and the failure of the apparent free lunch

The strongest architectural result concerns block-sparse attention. The evaluated block-sparse scheme processes documents independently and allows only designated query and answer tokens to attend globally. This reduces prefill computation relative to full attention by suppressing token-level interactions across documents. Such a design is well matched to retrieval tasks in which documents can be independently scored against a query.

On low-CTC tasks, block-sparse attention is effectively lossless relative to full attention across the evaluated context lengths. This reproduces the favorable conclusions of prior long-context work conducted primarily on retrieval, question answering, and related linear-time settings. The paper reports that mask-mixed training, in which batches probabilistically use full attention before annealing toward the block-sparse mask, improves block-sparse performance, particularly on aggregation and high-CTC tasks. Averaged across the 10-task comparison subset, the reported mask-mixing gains are larger for high-CTC tasks than for low-CTC tasks.

The result reverses on high-CTC tasks. Relative to full attention, block-sparse degradation averages approximately 26.9% at 2K tokens, 33.0% at 4K, 43.7% at 8K, 57.4% at 16K, and 65.9% at 32K. The degradation is therefore not a fixed penalty incurred at all context lengths; it increases as the corpus grows. On tasks such as contradiction, QDmatch, X-Absence, and scale-OT(N)O_T(N)8 outlier detection, suppressing direct cross-document interactions removes precisely the computational pathway required to compare candidate elements.

This finding qualifies the broad claim that block sparsity provides a “free lunch” for corpus reasoning. It provides a favorable computation-performance tradeoff for tasks whose solution decomposes into independent document processing followed by limited global aggregation. It does not preserve performance when the answer depends on dense or extensive relational interaction among documents. The implication is that attention sparsity must be evaluated against the interaction graph induced by the task, not merely against the input format of a corpus.

Mask mixing improves this tradeoff but does not eliminate it. The method appears to distill some full-attention behavior into representations usable under block-sparse inference, yet evaluation still applies the sparse mask exclusively. The remaining high-CTC gap indicates that training-time exposure to global interactions cannot fully compensate for their absence at inference time under the tested model and data regime.

Hybrid architectures

The paper extends the comparison to hybrid architectures that combine full-attention layers with recurrent or state-space-like components. It compares a full-attention version of OLMo-3-7B with OLMo-3-7B-Hybrid, which uses a 3:1 mixture of gated delta-net and full-attention layers. The full-attention baseline is adapted from the model’s default sliding-window configuration through continued pretraining.

The two architectures perform similarly on low-CTC tasks, with the hybrid sometimes performing better at 32K. On high-CTC tasks, however, the hybrid exhibits consistent performance gaps relative to full attention. The reported comparison therefore mirrors the block-sparse result: architectures that preserve sufficient information for independent retrieval and aggregation may fail when the task requires repeated, distributed comparison across the corpus.

The result should be interpreted cautiously because the compared OLMo models are generally weaker than the Qwen3.5 models used elsewhere, and the full-attention baseline requires architectural adaptation and continued pretraining. The paper accordingly treats the result as evidence for a robust trend rather than as a definitive estimate of the optimal attention-to-recurrent layer ratio. The specific question left open is whether a hybrid architecture with selectively placed or content-adaptive full-attention layers can retain high-CTC performance while reducing the cost of indiscriminate all-to-all attention.

Length generalization

The experiments also challenge the assumption that long-context capability learned at shorter lengths transfers uniformly across task types. Models trained through 32K tokens are evaluated at 64K and 128K using full attention. Performance is normalized by the corresponding 32K score.

At 128K, low-CTC tasks retain between 28% and 100% of their 32K performance, whereas high-CTC tasks retain only 4% to 14%; every reported high-CTC point falls below every low-CTC point in the comparison. This is a particularly strong separation because the models are evaluated within their native context capacity, rather than being subjected to an external context-extension method.

The implication is that length generalization is task-structural. Training on shorter contexts may teach a model to perform a local or fixed-number retrieval procedure that transfers to larger corpora. It does not necessarily teach the model how to maintain reliable coverage over an interaction space whose size grows quadratically or faster. Consequently, context-length training schedules should be stratified by CTC rather than evaluated solely by token length.

Robustness across model families and scales

The full-versus-block-sparse pattern appears across Qwen3.5, OLMo-3-7B, and Llama-3.2-3B. On HotpotQA, a representative low-CTC task, block-sparse attention causes little degradation. On contradiction, a representative high-CTC task, the degradation is substantial and increases with context length across model families.

Within the Qwen3.5 family, smaller models degrade faster on larger corpora. Additional full-attention experiments show that some high-CTC tasks have minimum scale requirements: Qwen3.5-0.8B and 2B models can fail to learn tasks such as QDmatch or reordering even before the longest context lengths, while the 4B model achieves substantially higher performance. This suggests that CTC does not only govern asymptotic degradation; it also changes the minimum representational and algorithmic capacity needed to acquire the task.

The scale findings qualify any interpretation of CTC as a model-independent difficulty law. The benchmark identifies a structural source of difficulty, but observed performance depends on model size, pretraining, instruction tuning, data construction, and the quality of the elementary operation. CTC predicts a pressure on computation and coverage; it does not determine the absolute score.

Limitations and open questions

The empirical evidence is concentrated at 2K–32K tokens for the main in-domain experiments and extends to 128K only for length generalization. The paper does not directly test million-token corpora, despite arguing that high-CTC penalties may become more pronounced at that scale. The asymptotic interpretation is therefore supported by finite-range trends rather than by measurements over a sufficiently broad scaling regime.

CTC-BENCH is diverse but not exhaustive. Several high-CTC tasks are synthetic or semi-synthetic, and some use simplifying assumptions that may make the required relation easier than in natural corpora. In particular, contradiction pairs include one LLM-generated claim, and the authors note that subtle generation artifacts may remain. X-Absence includes exact copies, making identity matching easier than semantic cross-corpus alignment. Strmatch deliberately removes semantic complexity, while Textgroups uses controlled lexical features. These designs are useful for isolating computational structure, but performance on them should not be equated with performance on unconstrained scientific, legal, or web corpora.

The CTC classification itself depends on the assumed oracle model and corpus representation. Indexing, metadata, embeddings, hierarchical clustering, approximate nearest-neighbor search, or domain-specific structure could reduce the effective cost of some nominally quadratic tasks. The framework correctly distinguishes online and offline costs, but the experiments do not evaluate retrieval indexes, agentic decomposition, recursive processing, or specialized pair-generation algorithms as alternatives to end-to-end LCLM inference.

The architectural comparisons also have scope limitations. Block-sparse attention is implemented through a particular document-level mask, and hybrid results depend on specific OLMo configurations. The models are fine-tuned independently per task, with one training epoch and a fixed data budget. It remains unresolved whether more extensive training, intermediate supervision over pairwise relations, iterative refinement, or learned sparse interaction patterns can close the high-CTC gap without restoring full attention.

Finally, the benchmark evaluates exact or near-exact output structures, such as exhaustive pair lists and complete group assignments. This is appropriate for measuring coverage over candidate interactions, but real applications may permit ranking, calibrated abstention, approximate discovery, or human-in-the-loop verification. Whether high-CTC tasks remain as difficult under these alternative utility functions is an open empirical question.

Conclusion

The paper establishes CTC as a useful organizing principle for long-context corpus reasoning. Its main empirical result is that tasks requiring quadratic or higher-order corpus interactions degrade substantially faster with context length than conventional linear-scan tasks, even under in-domain training and full attention. Block-sparse and hybrid architectures that appear competitive on retrieval and aggregation can therefore incur severe losses on high-CTC tasks, while short-context length generalization is markedly less reliable for the same class.

CTC-BENCH provides a concrete basis for evaluating these distinctions. The paper’s principal methodological conclusion is that long-context evaluations should report not only context length and task accuracy, but also the scaling structure of the required corpus interactions. The specific unresolved problem is how to construct inference procedures that achieve reliable high-CTC reasoning at corpus scales where unrestricted full attention is computationally infeasible.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

Explain it Like I'm 14

1. What is this paper about?

This paper studies how difficult it is for LLMs to answer questions about large collections of documents, such as scientific papers, websites, or book passages.

The authors argue that not all questions become difficult in the same way as the collection gets larger. For example:

  • Finding one document that answers a question is relatively simple.
  • Finding every pair of documents that contradict each other is much harder because many documents must be compared with one another.

To describe this difference, the authors introduce a new idea called Corpus Task Complexity, or CTC.

Their main message is:

Systems that work well on simple large-corpus tasks may fail badly on more complex tasks, especially when the corpus becomes very large.

2. What questions did the researchers ask?

The researchers wanted to understand several things:

  1. Why are some large-corpus questions harder than others?
  2. Does difficulty increase faster for certain types of tasks as the number of documents grows?
  3. Do LLMs perform worse on complex tasks when they read longer inputs?
  4. Do faster, cheaper attention methods work as well as regular attention on every kind of task?
  5. Can models trained on shorter texts successfully handle much longer collections?

The authors were especially interested in whether common tests give an overly positive picture of LLMs because they mostly use simpler tasks.

3. What is Corpus Task Complexity?

Imagine you are looking through a pile of documents.

Low-CTC tasks

A low-CTC task usually needs about one check per document. For example:

“Which document says when Ralph Lauren was founded?”

You can look through the documents one by one. If there are NN documents, the amount of work grows roughly like NN.

This is called linear growth, written as O(N)O(N).

High-CTC tasks

A high-CTC task requires comparing documents with one another. For example:

“Find every pair of documents that makes contradictory claims.”

With 3 documents, there are only a few pairs to compare. With 100 documents, there are:

100×992=4,950\frac{100 \times 99}{2} = 4{,}950

possible pairs.

The work grows approximately like N2N^2, or quadratically. This means that doubling the number of documents can require about four times as many comparisons.

Some tasks can be even harder. For example, checking groups of three documents can require work growing like N3N^3.

Task type Example How the work grows
Low CTC Find a document containing an answer O(N)O(N)
High CTC Find all contradictory document pairs O(N2)O(N^2)
Very high CTC Find groups of three documents with a shared property O(N3)O(N^3)

The authors compare this to searching for a friend at school. Finding one person might mean checking students one by one. Finding every pair of students who disagree about something requires comparing many pairs of students.

4. How did the researchers conduct the study?

Building CTC-BENCH

The researchers created a collection of 22 tasks, called CTC-BENCH:

  • 12 were low-CTC tasks.
  • 10 were high-CTC tasks.

The tasks included:

  • Finding relevant passages
  • Answering questions using documents
  • Finding contradictions
  • Finding missing text
  • Grouping documents by topic
  • Finding unusual or rare topics
  • Reordering shuffled text
  • Finding groups of documents with matching features

This gave the researchers a way to compare simple and complex tasks fairly.

Testing LLMs

They trained LLMs to solve these tasks using input lengths from about:

  • 2,000 tokens
  • 4,000 tokens
  • 8,000 tokens
  • 16,000 tokens
  • 32,000 tokens

A token is a small piece of text, such as a word or part of a word. More tokens mean a longer collection of documents.

The researchers mainly used a model called Qwen3.5-4B, but they also tested other model families and sizes to see whether the results were consistent.

Comparing attention methods

LLMs use a process called attention to decide which parts of the input are important.

The study compared:

  • Full attention: Every part of the input can directly interact with every other part. This can be powerful but becomes very expensive for long inputs.
  • Block-sparse attention: Documents mostly examine themselves, while only a few parts communicate across documents. This is faster and cheaper.
  • Hybrid models: These combine full attention with faster methods that use less memory and computation.

The researchers also studied length generalization. This means training a model on shorter inputs and then testing whether it can handle much longer inputs.

5. What did the researchers find?

High-CTC tasks became much harder as documents increased

The main result was that high-CTC tasks lost performance much faster than low-CTC tasks.

From 2,000 to 32,000 tokens:

  • Low-CTC performance fell from about 91% to 79%.
  • High-CTC performance fell from about 85% to 51%.

So even when models were trained using examples of different lengths, tasks involving many comparisons became much harder as the corpus grew.

Faster attention worked well only on simpler tasks

Block-sparse attention performed almost as well as full attention on low-CTC tasks such as retrieval and question answering.

However, on high-CTC tasks, its performance dropped significantly. At 32,000 tokens, block-sparse attention was about 66% worse than full attention on average for the high-CTC tasks.

This is important because earlier research might suggest that block-sparse attention is a “free lunch”: it is cheaper but seems to work just as well. The new study shows that this is true mainly for simpler tasks.

Hybrid models also struggled with high-CTC tasks

Hybrid models were often similar to full-attention models on low-CTC tasks.

But on high-CTC tasks, hybrid models usually performed worse. This suggests that reducing communication between different parts of the input can remove information needed for comparing many documents.

Training on short inputs did not always prepare models for long inputs

The researchers tested models trained up to 32,000 tokens on inputs as long as 128,000 tokens.

At that length:

  • Low-CTC tasks kept between about 28% and 100% of their 32,000-token performance.
  • High-CTC tasks kept only about 4% to 14%.

This shows that high-CTC tasks may be especially poor at handling inputs longer than those seen during training.

Bigger models helped, but did not solve everything

The general pattern appeared across different model sizes and model families.

Smaller models usually lost performance more quickly as the corpus grew. However, even larger models still faced serious problems on high-CTC tasks.

The study also found that tasks with more difficult basic operations became even more challenging when their complexity grew from O(N)O(N) to O(N2)O(N^2).

6. Why are these results important?

Many language-model evaluations focus on tasks such as:

  • Finding one relevant passage
  • Answering a question from a document
  • Retrieving a small number of supporting documents

These are mostly low-CTC tasks. They are useful, but they do not test whether a model can carefully compare every part of a large collection.

The paper shows that a model may appear excellent on common tests but still struggle to:

  • Find all contradictions in a scientific literature collection
  • Discover unusual topics
  • Match many questions with their correct documents
  • Reconstruct the order of shuffled passages
  • Compare every document with many others

In other words, there is no single method that is both extremely cheap and equally powerful for every kind of task. The title, No More Free Lunch, refers to this trade-off.

Full attention can handle more detailed interactions, but its cost grows very quickly. Faster methods reduce the cost, but they may lose important information needed for high-CTC reasoning.

7. Possible impact of the research

The paper suggests that future language-model research should test more than simple retrieval and question answering.

The released CTC-BENCH benchmark can help researchers measure whether new systems can handle difficult large-corpus tasks. It may encourage the development of models that can:

  • Compare documents efficiently
  • Preserve important connections across a large corpus
  • Solve complex tasks without using impossibly large amounts of computing power
  • Work reliably on collections much larger than those used during training

This could be useful in areas such as scientific research, law, medicine, journalism, and internet search.

The biggest challenge is finding a system that combines the strengths of both approaches:

  • The power of full attention for detailed comparisons
  • The low cost of efficient attention for very large collections

The paper concludes that solving high-CTC tasks at large scale remains an open problem. It also warns researchers not to assume that success on simple long-context tasks means a model can truly understand and analyze a large collection of documents.

Knowledge Gaps

Knowledge gaps, limitations, and open questions

The paper leaves the following issues unresolved:

  • CTC is defined in terms of oracle-call counts rather than end-to-end computational cost. The framework does not establish how oracle-call complexity maps to actual latency, FLOPs, memory use, energy consumption, communication overhead, or monetary cost for modern LCLMs.
  • The claimed optimality of the proposed algorithms is largely unproven. Several tasks are assigned O(N2)O(N^2) or higher CTC using brute-force candidate algorithms, but the paper does not prove that substantially cheaper algorithms are impossible when indexing, embeddings, approximate search, document structure, or task-specific preprocessing are allowed.
  • The assumption that corpora are unstructured is restrictive. It remains unclear how CTC changes when documents have metadata, timestamps, citations, hyperlinks, known topic labels, duplicate structure, or other exploitable organization.
  • The relationship between CTC and empirical difficulty is not formally characterized. The experiments show correlations between higher CTC and performance degradation, but do not establish a predictive law connecting asymptotic complexity, corpus size, oracle difficulty, and model accuracy.
  • The boundaries between CTC classes are underspecified. Tasks with sublinear indexing, O(Nlog⁡N)O(N\log N) preprocessing, fixed numbers of relevant documents, approximate answers, or data-dependent rather than worst-case complexity are not systematically classified.
  • The benchmark’s high-CTC tasks are mostly synthetic or semi-synthetic. Contradictions, shuffled corpora, word-sequence matching, artificial grouping, and constrained triple-search tasks may not represent the noise, ambiguity, redundancy, and evolving schemas of real scientific, legal, enterprise, or web corpora.
  • The validity of automatically generated labels is uncertain. The paper does not provide extensive human validation of LLM-generated contradictions, topic assignments, matching pairs, outliers, or triple-based answers, leaving open whether benchmark performance reflects reasoning ability or artifacts in the data-generation process.
  • The benchmark has limited scale. Experiments reach 32K tokens for the main comparisons and 128K for length generalization, far below the million-token or larger corpora that motivate the paper. It is unresolved whether the observed trends remain monotonic at realistic corpus scales.
  • The evaluation covers only a narrow range of model families and sizes. Most experiments use Qwen3.5-4B, with limited comparisons involving OLMo, Llama, and a few smaller models; the conclusions may differ for frontier-scale models, multimodal models, retrieval-augmented systems, or models specifically trained for long-context reasoning.
  • Training and evaluation data are largely in-distribution. Models are fine-tuned separately on 20,000 examples per task with training context lengths matching evaluation lengths, so the findings do not establish how high-CTC tasks behave under zero-shot, few-shot, cross-domain, or realistic distribution-shift conditions.
  • The study does not isolate memorization from reasoning. Because some corpora and task templates are derived from established datasets or repeated document sources, it remains unclear how much performance depends on memorized facts, formatting regularities, or source familiarity.
  • The comparison of attention architectures is not fully controlled. Full, block-sparse, and hybrid models differ in pretraining, parameterization, training procedures, masks, and continued-pretraining exposure, making it difficult to attribute performance gaps solely to the attention pattern.
  • The mask-mixing intervention is not comprehensively ablated. The paper reports a curriculum from high to zero full-attention probability, but does not determine which mixing schedules, probabilities, or training budgets are optimal across CTC classes and corpus sizes.
  • The benefits of block-sparse attention are evaluated primarily under a fixed document-level partition. More flexible sparse patterns—such as learned routing, hierarchical attention, retrieval-conditioned sparsity, global memory tokens, or adaptive pair selection—are not tested.
  • The paper does not compare against strong non-LCLM pipelines. Classical information retrieval, dense retrieval, clustering, graph algorithms, database joins, approximate nearest-neighbor methods, and agentic or tool-using systems may solve several high-CTC tasks more efficiently, but their accuracy–cost trade-offs are not evaluated.
  • Approximate solutions are not addressed. Many high-CTC tasks may not require exhaustive enumeration; approximate contradiction discovery, top-kk matching, probabilistic outlier detection, and candidate generation followed by verification could offer useful recall–cost trade-offs that the benchmark does not measure.
  • Evaluation metrics may inadequately capture set-valued answers. The paper uses measures such as set-F1, pair-F1, and partial credit, but does not analyze how these metrics penalize missed items, reward duplicates, handle uncertainty, or reflect the practical usefulness of incomplete results.
  • The task difficulty caused by output length is not separated from reasoning difficulty. High-CTC tasks often require producing many pairs, groups, or identifiers, so performance degradation may partly result from long decoding sequences, formatting errors, or output-token limits rather than inability to perform corpus-wide comparisons.
  • The impact of answer sparsity is not systematically controlled. High-CTC tasks differ in the number and distribution of correct pairs or items; the study does not determine whether degradation is driven by CTC itself or by rare positives, dense answer sets, class imbalance, or changing output cardinality.
  • The interaction between CTC and corpus redundancy is unexplored. Duplicate, near-duplicate, highly correlated, or clustered documents could reduce the effective search space, but the benchmark does not quantify or manipulate effective corpus size separately from token count.
  • The role of document length and token-level complexity is unclear. Experiments vary context length, but do not disentangle the number of documents, document length, number of claims, number of candidate relations, and total token count.
  • The generality of the oracle-difficulty analysis is limited. The paper compares a small number of paired tasks and measures oracle difficulty using one model’s performance; it does not establish whether the observed interaction generalizes across oracle definitions, models, domains, or human judgments.
  • The paper does not identify the point at which preprocessing becomes worthwhile. It remains open when the cost of building indexes, embeddings, clusters, graphs, or summaries is offset by reduced high-CTC inference cost, particularly under changing corpora and repeated queries.
  • Robustness to noisy, adversarial, or conflicting documents is not evaluated. Real corpora may contain misinformation, ambiguous claims, contradictory annotations, adversarial distractors, or distributional shifts that could alter both CTC and architectural rankings.
  • The practical quality requirements for high-CTC applications are unspecified. The study does not determine acceptable recall, precision, calibration, abstention, provenance, or human-review costs for applications such as scientific contradiction discovery or legal corpus analysis.
  • The benchmark does not test iterative or interactive workflows. Systems that progressively retrieve, cluster, verify, and revise candidate relations may have very different scaling behavior from one-shot LCLM inference, but such workflows are outside the evaluation.
  • The theoretical implications for scalable architectures remain open. The paper demonstrates that current sparse and hybrid methods lose accuracy on high-CTC tasks, but does not propose or validate an architecture that achieves near-full-attention quality with subquadratic cost.
  • It is unclear whether high-CTC performance can be improved through training alone. The study does not determine whether specialized objectives, synthetic pairwise supervision, curriculum learning, intermediate representations, explicit relational memory, or tool-use training can mitigate the observed scaling failures.
  • The effect of corpus ordering is insufficiently explored. Although some tasks shuffle documents, the paper does not systematically test adversarial, topical, temporal, random, or relevance-based ordering and its interaction with positional biases and sparse attention.
  • Reproducibility and benchmark stability may be affected by generated data and evolving model versions. The paper does not report how sensitive results are to random seeds, label-generation prompts, corpus construction choices, evaluator models, or future revisions of the underlying datasets.

Practical Applications

Immediate Applications

  • Benchmarking and procurement for long-context AI systems (software, enterprise AI, academia) Organizations can use CTC-BENCH and the paper’s distinction between low- and high-CTC tasks to evaluate models before deployment. Procurement tests should include not only retrieval and question answering, but also contradiction detection, document matching, grouping, reordering, and outlier discovery at increasing corpus sizes. Potential workflow: evaluate candidate models with both full-attention and efficient-attention configurations at the organization’s expected context length, then select architectures based on task-specific accuracy–cost trade-offs. Dependencies: the benchmark must be extended with domain-specific data; synthetic tasks may not fully represent real-world distributions, terminology, noise, or document structure.
  • Architecture selection for retrieval-augmented generation (RAG) (software, search, customer support, legal technology) Teams can classify a planned RAG workload by CTC before choosing an inference architecture. Conventional retrieval and fixed-hop question answering are generally low-CTC and may use block-sparse, hybrid, or linear-time components. Workflows requiring cross-document comparison—such as finding all inconsistent policies or matching many questions to many documents—should not assume that efficient attention is lossless. Potential product: a CTC-aware RAG router that directs simple queries to inexpensive sparse or hybrid models and complex corpus-wide analyses to more expensive full-attention or multi-stage pipelines. Dependencies: reliable identification of the task’s true computational structure and careful validation at production corpus sizes.
  • Large-scale contradiction and consistency screening (healthcare, legal, finance, compliance) Organizations can apply high-CTC task definitions to identify contradiction-search workloads in clinical literature, contracts, financial disclosures, internal policies, and regulatory documents. A practical near-term system would use retrieval or clustering to generate candidate document pairs, followed by an entailment or contradiction classifier, rather than exhaustively comparing every pair. Potential tools: literature-monitoring dashboards, contract inconsistency detectors, policy-diff systems, and financial-disclosure review assistants. Dependencies: candidate-generation recall is critical; missed pairs may be more damaging than false positives. Human review and domain-specific validation remain necessary, especially in healthcare and legal settings.
  • Evaluation of long-context model compression and efficiency techniques (AI infrastructure, cloud computing) Researchers and engineering teams can use high-CTC tests to detect failures hidden by standard retrieval benchmarks. Block-sparse attention, hybrid state-space/attention models, context extension, and length-generalization methods should be evaluated on pairwise and higher-order reasoning before being adopted for corpus-analysis products. Actionable practice: add performance-versus-context curves and full-attention baselines to model release evaluations, rather than reporting only aggregate accuracy at one context length. Dependencies: full-attention baselines may be expensive; reproducible hardware, tokenization, training data, and evaluation protocols are needed for fair comparisons.
  • Curriculum and training design using mask-mixing (model training, enterprise fine-tuning) The reported mask-mixing method—occasionally training block-sparse models with full-attention masks—can be tested as a practical intervention for improving efficient models, particularly on aggregation and high-CTC tasks. Potential workflow: begin fine-tuning with a substantial proportion of full-attention batches, gradually anneal that proportion, and evaluate whether the resulting model retains more high-order reasoning ability under block-sparse inference. Dependencies: the paper reports results for particular model families and task settings; gains may vary with model size, mask schedule, data quality, and compute budget.
  • Scientific literature analysis and evidence review (academia, healthcare, research policy) Research groups can use the benchmark’s task taxonomy to separate ordinary evidence retrieval from more demanding corpus reasoning. Immediate applications include detecting conflicting claims, identifying unmatched documents between literature versions, grouping abstracts by topic, and locating rare or anomalous findings. Potential product: a literature-review assistant that presents candidate contradictions and clusters together with source passages and confidence scores, rather than producing an unsupported summary. Dependencies: LLM-generated labels and contradiction judgments can contain errors; results should be treated as screening aids and verified by subject-matter experts.
  • Quality assurance for document collections (publishing, government records, enterprise knowledge management) The X-Absence and reordering task structures suggest practical checks for missing, duplicated, shuffled, or altered content across document repositories, backups, versions, and data migrations. Low-CTC aligned comparisons can be deployed now, while shuffled or unaligned corpus comparison should use candidate matching and verification stages. Dependencies: document segmentation, version alignment, OCR quality, and access to trusted reference copies determine reliability.
  • Daily-life tools for comparing and organizing personal documents (consumer software, education) A document assistant could compare multiple versions of leases, insurance policies, school materials, or user manuals, flag potentially contradictory clauses, group related passages, and identify missing sections. For ordinary retrieval and summarization, current efficient long-context systems may be adequate; corpus-wide comparison should expose uncertainty and avoid presenting automated findings as definitive. Dependencies: privacy-preserving local processing, secure storage, accurate OCR, and safeguards against false legal or medical interpretations.

Long-Term Applications

  • Scalable contradiction engines for entire scientific and regulatory corpora (healthcare, law, science policy) A mature system could continuously compare claims across millions of papers, clinical guidelines, patents, court decisions, or regulations and construct a graph of agreement, contradiction, qualification, and temporal change. This would support evidence synthesis, policy harmonization, and research-gap discovery. Why long-term: exhaustive pairwise reasoning has quadratic or higher CTC, while full attention becomes computationally impractical at large corpus sizes. Progress likely requires hierarchical indexing, learned candidate generation with recall guarantees, parallel pairwise reasoning, and architectures that preserve cross-document interactions. Dependencies: trustworthy contradiction definitions, temporal and source reliability modeling, multilingual support, and expert adjudication.
  • High-CTC scientific discovery and anomaly detection (biotechnology, materials science, climate research) Systems could identify rare topics, outlier experimental results, unusual combinations of observations, and groups of documents satisfying multi-document constraints. Such tools might reveal overlooked research directions or inconsistent experimental findings. Why long-term: outlier detection with a growing number of latent categories and higher-order grouping becomes substantially harder as the corpus expands. Dependencies: well-structured metadata, robust representations, control of publication bias, and mechanisms for distinguishing genuine novelty from noise or poor-quality reporting.
  • Corpus-scale legal and compliance reasoning (legal technology, finance, public administration) Future systems could compare every relevant contract clause, statute, regulatory interpretation, and corporate disclosure to identify conflicts, obligations, duplicated provisions, and missing evidence. This could support due diligence, regulatory change management, and automated compliance audits. Why long-term: legal relations are often pairwise or multi-document, and errors have high consequences; efficient attention methods that perform well on retrieval may fail on comprehensive comparison. Dependencies: jurisdiction-specific ontologies, citation and provenance tracking, explainability, audit logs, and mandatory human approval for consequential decisions.
  • Corpus-aware autonomous research agents (AI agents, robotics, knowledge work) Research agents could move beyond retrieving a few passages to constructing and testing relationships across a complete corpus: matching questions to all relevant documents, reconciling competing evidence, ordering events, and generating structured knowledge graphs. Why long-term: the paper notes that autoregressive agentic enumeration can require on the order of N2N^2 outputs, making naive multi-agent workflows too slow and expensive. Dependencies: efficient parallel inference, compact intermediate representations, reliable planning, and protection against cascading errors in automatically generated comparisons.
  • New attention architectures with adaptive CTC computation (AI research, hardware, cloud infrastructure) The findings motivate models that allocate computation according to task complexity: linear or block-sparse processing for independent retrieval, denser cross-document interaction for pairwise tasks, and specialized mechanisms for triple or higher-order relations. Potential innovation: an adaptive model that estimates whether a query is O(N)O(N), O(N2)O(N^2), or higher and dynamically activates global attention, pairwise comparison modules, or external indexes. Dependencies: theoretical guarantees, efficient hardware kernels, training objectives that reward global consistency, and methods for avoiding approximate-search recall loss.
  • CTC-aware data structures and hybrid indexing systems (databases, search, enterprise software) Database-like systems could combine inverted indexes, dense retrieval, clustering, locality-sensitive hashing, graph search, and targeted language-model verification. Rather than asking an LLM to compare all document pairs, the system would narrow the candidate space while tracking uncertainty and recall. Why long-term: high-CTC tasks require more than simply extending the context window; they require algorithmic decomposition and potentially task-specific indexes. Dependencies: assumptions about corpus structure, stable embeddings, domain drift monitoring, and empirical guarantees that pruning does not remove important relationships.
  • Policy standards for evaluating AI systems on corpus-scale reasoning (government, standards bodies, academia) Evaluation standards could require reporting performance by CTC class, corpus size, model scale, attention mechanism, inference cost, and length generalization. This would discourage claims based solely on low-CTC retrieval benchmarks and provide regulators with more realistic evidence about model capabilities. Dependencies: benchmark validity, representative real-world datasets, protection of sensitive corpora, and agreement on metrics for sets of contradictions, clusters, matches, and higher-order outputs.
  • Large-scale educational and archival knowledge systems (education, libraries, digital humanities) Future systems could compare textbooks, lecture materials, historical records, and archival collections to identify conflicting accounts, missing passages, evolving terminology, and relationships among sources. Students could receive evidence maps rather than single-answer summaries, while librarians could use automated corpus quality checks. Dependencies: provenance preservation, intellectual-property permissions, transparent uncertainty estimates, and careful handling of contested or culturally sensitive interpretations.
  • High-assurance decision-support systems in healthcare and finance (healthcare, insurance, finance) At sufficient maturity, high-CTC models could continuously reconcile patient guidelines, medical evidence, claims records, risk reports, or market disclosures. They could surface inconsistencies for professional review rather than autonomously making decisions. Why long-term: the paper demonstrates that high-CTC performance degrades substantially with context length even under in-distribution training, and smaller models degrade faster under efficient attention. Dependencies: validated domain models, privacy and security controls, regulatory approval, calibrated uncertainty, bias testing, and clear limits on automation.

Glossary

  • Agentic scaffold: A system architecture in which an AI model coordinates tools or subtasks to solve a problem. “One path is to develop agentic scaffolds that combine efficient tools like dense retrievers”
  • Asymptotic growth: The rate at which a quantity increases as the input size becomes large, usually expressed with Big-O notation. “CTC is the asymptotic growth in oracle calls needed to solve tasks as a function of corpus size N.”
  • Block-sparse attention: An attention mechanism that restricts interactions to selected blocks rather than allowing every token to attend to every other token. “block-sparse attention: pre-filling each document independently so that they attend to tokens within the same document only”
  • Corpus reasoning task: A task that requires extracting, comparing, aggregating, or otherwise reasoning over a collection of documents. “We refer to these collectively as corpus reasoning tasks”
  • Corpus Task Complexity (CTC): A measure of task difficulty based on how the required number of oracle operations scales with corpus size. “We begin by defining corpus task complexity (CTC), which characterizes a task by how the number of operations required to solve it scales with corpus size.”
  • Cross-document interaction: Reasoning that requires comparing or relating information across multiple documents. “tasks requiring extensive cross-document interaction”
  • Curriculum: A training strategy that gradually changes the difficulty or composition of training examples or conditions. “we anneal on a curriculum from p = 0.8 to p = 0 during training”
  • Dense retriever: An information-retrieval model that represents queries and documents as dense vectors and retrieves items according to vector similarity. “agentic scaffolds that combine efficient tools like dense retrievers”
  • Distributional question: A question about the frequency, distribution, or statistical pattern of labels or properties in a corpus. “OOLONG[Bertsch et al., 2025] involves distributional or counting questions about class labels of documents.”
  • End-to-end: Describing a system that processes input through the complete pipeline in one integrated operation. “architectures capable of ingesting full corpora end-to-end”
  • Full attention: An attention mechanism in which every input token can attend to every other input token. “In a regular Transformer model (full attention), pre-filling a corpus has quadratic computational complexity.”
  • Gated delta network (GDN): A recurrent sequence-modeling architecture that uses gated state updates based on a delta rule to reduce computational cost. “adjacent approaches like gated delta net (GDN)[Yang et al., 2024] have also enabled attention to reduce computational complexity.”
  • Hybrid architecture: A model architecture that combines different sequence-processing mechanisms, such as recurrent layers and full-attention layers. “hybrid architectures that alternate linear and full attention exhibit larger performance gaps on high-CTC tasks”
  • In-distribution training: Training and evaluation in which the data conditions, such as task type and context lengths, are matched. “We train Qwen3.5-4B individually for all 22 CTC-BENCH tasks, ranging from 2k to 32k context length in training and evaluation.”
  • Information retrieval (IR): The computational task of finding documents or passages relevant to a query. “Information retrieval is arguably the best-studied corpus reasoning task.”
  • Length generalization: The ability of a model trained on shorter inputs to perform effectively on longer inputs. “length generalization—fine-tuning a long-context model on shorter contexts than the intended context length at test time”
  • Linear attention: An attention formulation whose computational cost grows linearly rather than quadratically with sequence length. “these methods replace all-to-all token interactions with a recurrent state update, giving linear computation and constant memory.”
  • Long-context LLM (LCLM): A LLM designed to process unusually long token sequences or contexts. “Long-context LLMs (LCLMs) trained to process large inputs”
  • Mask-mixing: A training method that alternates between full-attention and block-sparse attention masks. “we additionally propose a new mask-mixing technique”
  • Multi-hop question answering: Question answering that requires retrieving and combining information through multiple reasoning or retrieval steps. “HotpotQA [Yang et al., 2018], a multi-hop task.”
  • Oracle operation: A basic information-processing query, often made to a language-model judge, used as an assumed unit of computation in the complexity analysis. “CTC is the asymptotic growth in oracle calls needed to solve tasks as a function of corpus size N.”
  • Pairwise comparison: An operation that evaluates relationships between every possible pair of items. “exhaustively checking all possible pairs will generally require quadratic oracle calls in corpus size”
  • Quadratic complexity: Computational complexity whose cost grows proportionally to the square of the input size, expressed as O(N2)O(N^2). “finding contradictions requires checking a quadratically growing set of claim pairs.”
  • Recurrent state update: A sequential operation that updates a model’s hidden state as it processes each new token or input. “these methods replace all-to-all token interactions with a recurrent state update”
  • Retrieval-augmented generation (RAG): A method that retrieves relevant documents and supplies them to a generative LLM to produce an answer. “Accelerating retrieval-augmented generation with precomputed kv caches for chunked text.”
  • Sliding window attention: An attention mechanism in which each token attends only to nearby tokens within a fixed-size window. “since Olmo3 by default uses sliding window attention”
  • Softmax attention: The standard Transformer attention mechanism that converts attention scores into normalized weights using the softmax function. “with softmax attention enabling direct retrieval and reasoning across parts of the conditioned corpus.”
  • State-space model (SSM): A sequence model that represents evolving inputs through a latent dynamical state, often enabling efficient long-sequence processing. “alternative architectures such as state-space-models (SSMs)”
  • Synthetic benchmark: An evaluation dataset or task generated artificially to control its properties and test specific capabilities. “We often include synthetic / semi-synthetic tasks for controllability and cleaner analysis”
  • Token interaction: The computational relationship in which one token attends to or incorporates information from another token. “these methods replace all-to-all token interactions with a recurrent state update”
  • Transformer: A neural-network architecture based primarily on attention mechanisms for processing sequences. “In a regular Transformer model (full attention), pre-filling a corpus has quadratic computational complexity.”
  • Unstructured corpus: A collection of documents without a predefined relational or tabular organization. “Define a corpus C to be an unstructured set of N documents”
  • Zero-shot evaluation: Evaluation in which a model performs a task without task-specific examples in the prompt or additional task training. “BEIR [Thakur et al., 2021], a heterogenous benchmark for zero-shot evaluation of information retrieval models.”

Tweets

Sign up for free to view the 3 tweets with 195 likes about this paper.