Papers
Topics
Authors
Recent
Search
2000 character limit reached

Corpus Task Complexity (CTC)

Updated 29 September 2026
  • Corpus Task Complexity (CTC) comprises two concepts: interactional CTC, characterized by the multidimensional demands of visual problem-solving and spontaneous dialogue, and computational CTC, defined by the asymptotic growth of operations needed as corpus size increases
  • The Spot the Difference Corpus exemplifies interactional CTC through pairs of participants who collaboratively identify differences in similar images, involving descriptions, location negotiations, and role adaptations which are analyzed for task success and interactional complexity
  • Computational CTC describes the scaling behavior of tasks, distinguishing low-CTC tasks that require linear operations from high-CTC tasks that require exponential or higher-order operations, impacting both training and evaluation processes in corpus-level reasoning

Corpus Task Complexity (CTC) denotes two related but distinct research concepts. In the Spot the Difference Corpus, CTC concerns the multidimensional demands imposed by a controlled visual problem-solving task and spontaneous multimodal dialogue, including perceptual, search, interactional, linguistic, temporal, and cognitive demands (Lopes et al., 2018). In “No More Free Lunch: Corpus Task Complexity Matters as Corpora Grow,” CTC is defined asymptotically as the growth in the minimum number of corpus-level oracle operations required to solve a task as corpus size NN increases (Singhal et al., 24 Sep 2026). The first conception characterizes the complexity of an interactional task; the second characterizes the computational structure of reasoning over a corpus. Together, they distinguish task difficulty from any single surface property such as visual entropy, token length, or semantic difficulty.

1. Conceptual foundations

The interactional conception of CTC treats task complexity as a multidimensional construct. A task may be difficult because participants must search a large visual space, identify subtle differences, formulate referring expressions, establish common ground, coordinate turns, resolve uncertainty, or complete actions under time pressure. The Spot the Difference Corpus is designed to expose these dimensions by combining controlled visual materials with spontaneous speech, gaze, mouse actions, video, system events, and detailed annotations (Lopes et al., 2018).

The computational conception defines a corpus as an unstructured set of NN documents,

C={d1,…,dN},C=\{d_1,\ldots,d_N\},

where each did_i is a token sequence. A corpus task consists of corpus–question–answer triples (Ci,qi,ai)(C_i,q_i,a_i). For a task-solving algorithm ATA_T, let fAT(Ci,qi)f_{A_T}(C_i,q_i) denote the number of oracle calls required to solve an instance. The task is classified as T∈OT(g(N))T\in O_T(g(N)) when there is a constant α\alpha such that

fAT(Ci,qi)≤αg(N)f_{A_T}(C_i,q_i)\leq \alpha g(N)

for all task samples, with NN0 intended to be the lowest-call-count algorithm. CTC is therefore the asymptotic growth of the minimum number of oracle operations required as corpus size increases.

This definition abstracts away the cost of an individual oracle call. A semantic contradiction judgment and an exact string-match judgment may both require NN1 comparisons even though the contradiction oracle is intrinsically more difficult. The latter property is termed oracle difficulty. CTC and oracle difficulty can interact: increasing the number of required operations and increasing the difficulty of each operation jointly enlarge the burden on a model.

The two uses of CTC are not interchangeable. Interactional CTC describes the sources and manifestations of demand within a task episode. Asymptotic CTC describes how corpus-level computational requirements scale. A plausible implication is that the first can provide explanatory variables for observed behavior, whereas the second provides a structural classification of the task’s required computation.

2. The Spot the Difference Corpus as an interactional CTC resource

The Spot the Difference Corpus contains 54 dialogues involving pairs of participants who collaborate to identify all differences between two similar pictures. Participants sat in separate rooms and had no visual access to one another’s scenes, so they had to describe objects, establish locations, compare corresponding regions, and resolve discrepancies through speech. When a possible difference was identified, the participants entered a locating subdialogue and used the mouse to click on the relevant area. A difference counted as found only when the click satisfied the system’s location criterion (Lopes et al., 2018).

Two roles were assigned at the beginning of each interaction. The Instruction Giver was expected to lead the discussion by describing and locating objects, while the Instruction Follower was expected to follow instructions and request clarification. Roles were randomly assigned after the first three-scene set and then retained. However, roles could become reversed during interaction, demonstrating that nominal role assignment does not necessarily correspond to actual participation or cognitive responsibility.

Each dialogue contained three scenes initially designated easy, average, and difficult according to the number of differences. Scene and set order were randomized, and participants first completed a training scene. Condition A paired participants with a different partner for each set, whereas Condition B used the same partner for all three sets. The design therefore supports analysis of task learning and partner adaptation. The expected partner-specific improvement in Condition B was not clearly supported in the preliminary analysis; the linear approximation instead suggested greater improvement in Condition A, although this was presented as a descriptive comparison.

Scenes had a 200-second time limit, but participants could terminate a scene earlier by agreeing to reveal the solution. Audio remained open after the solution was displayed, permitting discussion of missed differences and strategy refinement. The corpus consequently includes problem-solving episodes, post-solution reflection, and between-scene strategy discussion. These segments are relevant to time-on-task, learning, adaptation, and interaction history.

The corpus was produced by 36 participants recruited at KTH in 2016. The participants included 14 women and 22 men, had an average age of 34.3 years, and represented 18 mother tongues. Only one participant was a native English speaker; most non-native speakers reported fluency and used English at work. Participants also completed a short personality questionnaire containing eight Big Five Inventory items used to position them on an introversion–extroversion dimension. The paper does not report analyses of personality, proficiency, education, or individual ability, although these metadata can serve as covariates in later CTC studies.

3. Operational measures and multimodal annotations

The primary task-success measure is the relative number of differences found:

NN2

Success was examined across time and accumulated task experience. The mean duration per scene was 188.3 seconds, close to the 200-second limit. The average number of turns per dialogue was 121.4, with a standard deviation of 44.4. Four dialogues were excluded from data analysis because of technical problems. The corpus does not provide a formal survival analysis, time-to-first-difference measure, time-to-last-difference measure, or temporal growth model of success, although the available data support such analyses.

The materials include beach, church, city, farm, house, jungle, sea, sheep, and winter scenes. Each scene had an XML description specifying object arrangement, object coordinates, object radius, object color, difference status, visibility, and possible referring expressions. These representations support variables beyond image entropy, including object count, grouped and individual objects, object density, spatial relations, difference location, and potentially difference salience.

The corpus contains several synchronized or synchronizable modalities:

  • Audio: head-mounted microphones produced separate mono files for each participant; additional microphones produced stereo set-level recordings.
  • Eye tracking: raw fixation data and object-level fixations were recorded, with object fixations linked to XML identifiers and represented in Praat tiers.
  • Mouse actions: click coordinates were stored as point tiers in Praat.
  • Video: two GoPro cameras recorded each set, supporting facial, visual, and uncertainty analyses.
  • System events: IrisTK logs recorded mouse clicks, eye-tracking information, and other interaction events.

Streams were synchronized at the beginning of each set with a beep because IrisTK timestamps originated from individual computers. The release includes general, set-level, and scene-level JSON descriptions, participant metadata, media files, offsets, logs, Praat annotations, and ELAN uncertainty annotations. The public repository is https://github.com/zedavid/SpotTheDifferenceData.

Speech was processed using IBM Watson automatic speech recognition and manual transcription. The ASR output was considered reasonably accurate for content and useful for estimating utterance boundaries and studying turn-taking. Manual transcription included disfluencies following the scheme of Moniz et al.; filled pauses were separately annotated, including English and first-language filled pauses. Agreement between separate filled-pause annotation and transcription-based coding was 0.61, characterized as moderate.

Each dialogue was manually segmented into four topics:

  1. Describing Scene (DS): describing the overall scene and spatial relations.
  2. Describing Object (DO): describing characteristics of a particular object.
  3. Locating Difference (LD): negotiating the exact location of a difference.
  4. End of Dialogue (EOD): negotiating whether to reveal the solution.

Speech overlaps were automatically extracted. Overlap involving a floor change was annotated as an interruption when the interrupted speaker did not complete the sentence. Collaborative interruptions completed the interrupted speaker’s sentence, whereas competitive interruptions produced material unrelated to the interrupted sentence. Uncertainty was annotated from video using a binary perceptual definition: a participant was uncertain when they appeared not to understand the partner or did not know what to say.

These annotations enable analysis of material, task, interactional, linguistic, and cognitive complexity without reducing them to one scalar measure.

4. Visual complexity and interactional behavior

The initial difficulty classification based on the number of differences did not adequately predict observed performance. Some scenes classified as average produced lower success than scenes classified as difficult. Beach and sea scenes appeared more complex than all scenes in the average group and the sheep scene in the difficult group.

The authors therefore computed Shannon entropy over the histogram of RGB components. If NN3 denotes an RGB-component value sampled according to the image histogram, entropy is

NN4

where NN5 is a possible RGB value or histogram bin and NN6 is the proportion of sampled pixels assigned to that value or bin. The paper does not specify the logarithm base, so the entropy units cannot be determined from the text alone.

With the exception of the house scene, scenes with the highest success rates—jungle, farm, and sheep—also had the lowest entropy values. The sea scene had an entropy value of 3.99, while the median entropy across scenes was 3.57. RGB entropy alone was insufficient to represent task demands: the house scene had low entropy but contained 21 XML-defined objects, compared with a median of 15, whereas sea contained only nine objects but had relatively high entropy. The authors therefore proposed combining RGB entropy with object count and potentially distinguishing grouped from individual objects.

The reported entropy analysis is exploratory and scene-level. The paper does not report a fitted multivariate complexity formula, regression coefficient, correlation coefficient, confidence interval, or statistical test directly linking entropy to success. It therefore does not establish that color variation causes dialogue difficulty.

The principal complexity-related interactional finding was a tendency for more complex scenes to contain more competitive interruptions. The number of interruptions, particularly competitive interruptions, broadly followed the inverse pattern of task success: scenes with higher relative success tended to have fewer interruptions. Scenes with the lowest number of collaborative interruptions were generally those with the highest success rates, whereas overlaps as a whole showed no equally clear trend. Uncertainty did not follow the same pattern as interruption.

Several reported chi-square tests connected uncertainty with overlap and interruption behavior. The interruption–uncertainty interaction was significant for Instruction Givers at NN7, for Instruction Followers at NN8, and for both participants combined at NN9. Overlap–uncertainty differences were significant for Instruction Followers at C={d1,…,dN},C=\{d_1,\ldots,d_N\},0 and for both participants combined at C={d1,…,dN},C=\{d_1,\ldots,d_N\},1, but not for Instruction Givers. The supplied text does not report the associated chi-square statistics, degrees of freedom, effect sizes, comparison-specific sample sizes, or model specifications.

A multimodal prediction experiment used facial features extracted with OpenFace and variance in gaze movement to predict uncertainty with artificial neural networks. On an evenly distributed subset, the best accuracy was 62%, compared with a 54% majority baseline. This result concerns prediction of annotated uncertainty rather than a direct test of scene complexity.

5. Asymptotic CTC and corpus-scale reasoning

The asymptotic formulation distinguishes low-CTC tasks from high-CTC tasks according to how required corpus-level work grows with C={d1,…,dN},C=\{d_1,\ldots,d_N\},2.

Low-CTC tasks generally require one or a constant number of corpus passes. Conventional retrieval is typically C={d1,…,dN},C=\{d_1,\ldots,d_N\},3 because every document may need to be examined. Multi-hop retrieval can remain low CTC when the number of retrieval passes is constant: two passes still yield C={d1,…,dN},C=\{d_1,\ldots,d_N\},4. Other low-CTC tasks include classifying every document and aggregating fixed-label statistics, retrieving a fixed number of relevant passages, comparing positionally aligned copies of two corpora, and finding outliers when possible categories are fixed and known.

High-CTC tasks require pairwise, groupwise, or higher-order interactions. If any pair of documents may contradict another, exhaustive contradiction identification requires

C={d1,…,dN},C=\{d_1,\ldots,d_N\},5

comparisons. Tasks involving C={d1,…,dN},C=\{d_1,\ldots,d_N\},6 documents and C={d1,…,dN},C=\{d_1,\ldots,d_N\},7 latent groups may require C={d1,…,dN},C=\{d_1,\ldots,d_N\},8 work; when C={d1,…,dN},C=\{d_1,\ldots,d_N\},9 grows proportionally with did_i0, this becomes quadratic. Tasks requiring combinations of three documents can require did_i1 work.

The CTC-BENCH suite contains 22 tasks: 12 low-CTC tasks already common in long-context evaluations and 10 newly introduced high-CTC tasks. The low-CTC group includes NIAH-contra, SciFact, FiQA, MS MARCO, OBLIQ, NQ, HotpotQA, MS MARCO rerank, OOLONG, Outlier (Amazon), Outlier (Wikipedia, fixed did_i2), and Absence.

The high-CTC group includes:

  • Contradiction: finding all contradictory claim pairs, nominally did_i3.
  • X-Absence: identifying passages appearing in only one of two shuffled near-identical corpora, did_i4.
  • QDmatch: identifying relevant question–document pairs, did_i5.
  • Outlier (Wikipedia, scale did_i6): identifying chunks from the least common source article, did_i7.
  • Grouping (OpenAlex): partitioning scientific abstracts into topical groups, with did_i8 behavior as the number of groups grows.
  • Strmatch: finding pairs of strings sharing a contiguous sequence of at least did_i9 words, (Ci,qi,ai)(C_i,q_i,a_i)0.
  • Reorder: recovering the original order of shuffled book segments, (Ci,qi,ai)(C_i,q_i,a_i)1.
  • Textgroups: finding triples whose hidden feature values sum to a target, (Ci,qi,ai)(C_i,q_i,a_i)2.

The corresponding metrics include gold-document F1, set-F1, pair-F1, MRR@10, partial-credit scoring, Kendall’s (Ci,qi,ai)(C_i,q_i,a_i)3, and group-F1, depending on the task.

6. Scaling behavior and architectural consequences

Experiments fine-tuned Qwen3.5-4B separately on all 22 tasks at context lengths of 2K, 4K, 8K, 16K, and 32K tokens. Training generally used 20,000 examples per task, evenly split across the five lengths, one epoch, a learning rate of (Ci,qi,ai)(C_i,q_i,a_i)4, full-parameter fine-tuning, greedy decoding, and disjoint training and evaluation examples. The experiments compared full attention, block-sparse attention, mask-mixed block-sparse attention, and hybrid attention architectures, including OLMo-3-7B and OLMo-3-7B-Hybrid.

Under full attention, average low-CTC performance fell from 0.910 at 2K to 0.788 at 32K, a 13% relative drop. Average high-CTC performance fell from 0.851 to 0.511, a 40% relative drop. Seven of the ten high-CTC tasks lost more than 0.1, whereas ten of the twelve low-CTC tasks degraded by less than 0.1. These results indicate that increasing corpus size is substantially more damaging when the evidence structure requires many cross-document interactions, even when models are trained in-distribution at every tested length.

Block-sparse attention often matched full attention on low-CTC tasks, but degraded on high-CTC tasks because independent document processing removes direct document–document interactions. The average high-CTC degradation relative to full attention increased from (Ci,qi,ai)(C_i,q_i,a_i)5 at 2K to (Ci,qi,ai)(C_i,q_i,a_i)6 at 32K. The corresponding low-CTC gaps were approximately (Ci,qi,ai)(C_i,q_i,a_i)7 to (Ci,qi,ai)(C_i,q_i,a_i)8.

Mask mixing, which used full-attention masks with probability (Ci,qi,ai)(C_i,q_i,a_i)9 annealed from ATA_T0 to ATA_T1 during training while evaluating with block-sparse masks, improved block-sparse models. Reported gains included ATA_T2 for Contradiction, ATA_T3 for X-Absence, ATA_T4 for Outlier at scale ATA_T5, and ATA_T6 for QDmatch using NQ. The improvement did not remove the inference-time limitation of block-sparse attention.

The same pattern appeared for hybrid architectures. OLMo-3-7B-Hybrid and full-attention OLMo-3-7B were often comparable on low-CTC tasks, but the hybrid consistently underperformed full attention on high-CTC tasks. Length generalization also differed: at 128K tokens, models trained only through 32K retained between 28% and 100% of their low-CTC performance, but only 4% to 14% of their high-CTC performance.

These results qualify conclusions drawn from retrieval-dominated long-context evaluations. Strong performance on needle retrieval or fixed-passage QA does not establish the ability to enumerate relations distributed across a corpus. Efficient attention can preserve low-CTC performance while discarding interactions required by contradiction detection, cross-corpus matching, grouping, and reordering.

7. Framework construction, limitations, and research directions

A comprehensive CTC framework should separate material, task, interactional, linguistic, cognitive, and computational dimensions. For the Spot the Difference Corpus, candidate independent variables include RGB entropy, object count, grouped and individual objects, number of differences, object density, spatial proximity, difference size, contrast, visibility, salience, scene order, partner familiarity, interaction condition, participant role, language background, age, gender, personality, English proficiency, prior success, strategy discussion, elapsed time, and remaining time.

Potential dependent variables include proportion of differences found, time to first and final differences, timeout or voluntary solution reveal, turn count and duration, participation equality, overlap and interruption rates, repair and clarification frequency, filled pauses, uncertainty duration, fixation measures, gaze-to-object alignment, click accuracy, lexical alignment, referring-expression efficiency, role symmetry or reversal, and topic-specific behavior across DS, DO, LD, and EOD.

The interactional formulation can be represented analytically as

ATA_T7

provided that the terms are treated as analytically distinct sources of demand rather than literal additive psychological quantities. Mixed-effects models with scenes and dyads as random effects could separate scene-level and dialogue-level predictors while controlling for participant ability, role, condition, scene order, and elapsed time.

The asymptotic formulation suggests a complementary benchmark methodology. Evaluations should report corpus size ATA_T8, the scaling of required comparisons, whether the number of latent categories ATA_T9 grows, whether the task requires pairs, triples, or higher-order combinations, performance across corpus lengths, and the performance of efficient attention relative to full attention.

Several limitations affect both formulations. RGB entropy measures color variation rather than task demands, ignores spatial organization and object semantics, does not represent difference properties, is sensitive to image rendering, and is static at the scene level. Its calculation is also underspecified with respect to channel combination, binning, and image aggregation. The corpus contains relatively few distinct scene materials, all participants saw scenes only once, and scene-level variables can be confounded with object layout, difference salience, elapsed time, participant ability, partner adaptation, and post-solution discussion.

The asymptotic framework likewise has qualifications. The optimal algorithm is often unknown, and priors, metadata, indexing, or corpus structure may reduce practical cost. The experiments include 22 tasks, some synthetic or semi-synthetic; standard in-domain experiments reach 32K tokens, while length-generalization tests reach 128K; million-token settings are not directly tested. The contradiction dataset contains LLM-generated contradictory claims and may contain artifacts, and full attention itself is not scalable to very large corpora.

The combined perspective suggests that future CTC research should distinguish the complexity of the underlying materials, the demands of task execution, the interactional and cognitive consequences of those demands, and the asymptotic computation required to solve the task. RGB entropy can function as one reproducible visual predictor, while corpus-scale CTC can classify whether reasoning requires linear passes, pairwise interactions, groupwise comparisons, or higher-order combinations. The central unresolved problem is to develop systems that preserve the cross-document interactions required by high-CTC reasoning without incurring the full computational cost of exhaustive all-to-all comparison.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Corpus Task Complexity (CTC).