Papers
Topics
Authors
Recent
Search
2000 character limit reached

Temporal Hallucination Index (THI)

Updated 5 July 2026
  • Temporal Hallucination Index (THI) is a metric that quantifies temporal instability in sequential decision tasks by focusing on observable delay and timeout events.
  • It operationalizes temporal reliability in paradigms like the tumbling-E task by analyzing reaction times, timeout rates, and staircase convergence towards perceptual thresholds.
  • THI provides a human baseline for comparing dynamic decision-making performance, highlighting the shift from static accuracy metrics to measures of sequential stability.

Searching arXiv for papers on “Temporal Hallucination Index” and closely related temporal hallucination metrics. Temporal Hallucination Index (THI) denotes a temporal-reliability construct introduced in the context of a computerized dynamic tumbling-E task, where sequential perceptual decisions are treated as temporally extended processes rather than as static threshold outcomes. In that framing, THI is intended to capture delay, timeout, drift, persistence, unstable convergence, and broader sequential reliability problems that ordinary accuracy metrics can miss. The defining use case is a human-only baseline for future comparison with artificial agents, but the same term also sits near a broader family of temporal-hallucination evaluations in role-play, video understanding, hidden-state analysis, and sequential detection. In the tumbling-E manuscript, however, only a reduced, observable THI was computed because the exported Correct column was described as non-informative, which prevented reliable estimation of correctness-dependent components such as drift and persistence (Sandhu et al., 20 Jun 2026).

1. Conceptual scope and intended target of THI

In the tumbling-E study, THI is not a synonym for final-task accuracy. It is introduced to quantify temporal instabilities that arise when perceptual decisions unfold across a sequence of trials. The central claim is that a participant or system can be accurate on the final answer while still being slow, delayed, unstable, drifting, persistent, or prone to timeout, and that these behaviors are obscured by static accuracy summaries (Sandhu et al., 20 Jun 2026).

The paper therefore positions THI as a measure of temporal reliability, not merely correctness. The behaviors explicitly linked to THI in that formulation are delay, timeout, drift, persistence, unstable convergence, processing drift, flip-flopping, sudden timeouts, and broader sequential inconsistency. This suggests that the term is meant to index instability in the trajectory of decision-making rather than only the terminal state of that trajectory. Because the paper could not recover correctness-dependent components, the implemented form is narrower than the conceptual one.

A common misconception is to treat THI as a generic hallucination score across all AI settings. The available literature does not support that simplification. Several adjacent papers study temporal hallucination, but they do not define THI by name; instead they use task-specific quantities such as Temporal Hallucination Rate (THR), timestamp prediction accuracy, event-order accuracy, benchmark-specific temporal subtasks, spectral hidden-state signals, or expected detection delay (Sadeq et al., 2024).

2. Operationalization in the dynamic computerized tumbling-E paradigm

The experimental substrate for THI in the tumbling-E work is a computerized dynamic task in which a single E optotype appears on each trial in one of four orientations—up, right, down, or left—and the participant either selects the perceived direction or times out. Unlike a static chart, the task uses an adaptive staircase: early trials use large, easy optotypes; later trials become smaller and harder; and the sequence moves toward the participant’s perceptual threshold. The approximate viewing distance was 30 cm, and the paper emphasizes that the exported Arcmin field was used directly rather than recomputing from pixels (Sandhu et al., 20 Jun 2026).

The basic temporal quantities are defined directly. Reaction time is given by

RT=elapsed time in ms between stimulus presentation and recorded orientation response.RT = \text{elapsed time in ms between stimulus presentation and recorded orientation response}.

A timeout occurs when no valid orientation choice is made within the response budget. The response budget is 3 seconds, so delay is operationalized as

RT>3000 ms,RT > 3000\ \text{ms},

and the exported RT field contains "timeout" for nonresponses (Sandhu et al., 20 Jun 2026).

Within this design, the staircase itself is part of the temporal-reliability interpretation. The task does not only ask whether a direction was eventually selected; it also preserves response latency, timeout behavior, and the path by which stimulus size adapts across trials. A plausible implication is that THI in this setting is inseparable from the dynamics of threshold approach.

3. Observable THI, component rates, and reported human values

Because the full THI decomposition was unavailable, the paper defines an observable THI using directly measurable instability events. The provided formula appears as

Observable THI=(Delay rate+Timeout rate)/\text{Observable THI} = (\text{Delay rate} + \text{Timeout rate}) /

but the denominator is visibly truncated in the supplied text, so the full expression is not recoverable from the excerpt. The manuscript nevertheless makes clear that this observable THI is a narrow index based on delay and timeout components rather than the full correctness-based THI (Sandhu et al., 20 Jun 2026).

The reported human baseline values are as follows:

Quantity Definition in the paper Reported value
Timeout rate 76 timeouts out of 1,154 valid trials 6.6%
Delay rate above 3 s 3 responses exceeding 3,000 ms 0.28% of non-timeout responses
Delay proportion in figure text Delay-over-3-seconds proportion 0.003\approx 0.003
Observable THI Delay and timeout based reduced index 0.034

The dataset included 1,154 valid trials from 21 human identifiers across 77 sessions. There were 1,078 non-timeout responses and 76 timeouts. Non-timeout reaction times were centered near 1.5 seconds, with mean 1546 ms, SD 350 ms, median 1506 ms, and IQR 1306–1713 ms; only 3 responses exceeded 3,000 ms (Sandhu et al., 20 Jun 2026).

The paper interprets the observable THI value of $0.034$ as close to the expected human reference value of approximately $0.03$, and contrasts it with a benchmark AI example value of $0.23$. Within the reported human data, timeouts contributed most of the observable instability, whereas delayed responses were extremely rare. This suggests that, in the implemented form of THI, timeout burden was the dominant source of temporal instability in the human baseline.

4. Staircase behavior, convergence, and the meaning of temporal reliability

The tumbling-E paper treats staircase behavior as evidence about sequential stability. Across 1,077 within-session transitions, 89.2% moved to a smaller next stimulus, 10.2% were unchanged, and 0.6% moved to a larger next stimulus. The authors interpret a smaller-next transition as consistent with the staircase increasing difficulty (Sandhu et al., 20 Jun 2026).

Convergence is summarized through arcminute trajectories. Mean arcminutes declined from 29.42 at trial 0 to 5.04 at trial 19, and the paper presents this as evidence that the staircase converged near a 20/20-level stimulus. Using the standard optotype geometry stated in the manuscript, 5 arcminutes corresponds approximately to 20/20 if the exported Arcmin is total optotype height. On that basis, 5.04 arcmin at trial 19 is interpreted as approximately 20/20.2, and the mean over trials 17–19 of 5.07 arcmin as approximately 20/20.3. The manuscript explicitly frames these as task-equivalent interpretations rather than clinical acuity diagnoses (Sandhu et al., 20 Jun 2026).

This aspect is important because THI is linked not only to isolated events such as delay or timeout but also to whether behavior remains smooth and convergent as difficulty increases. The paper explicitly contrasts static accuracy with temporal-reliability questions such as how long it takes to respond, whether the participant or system stalls, whether it times out, whether the staircase progresses smoothly, and whether behavior remains stable as difficulty increases. In that sense, unstable convergence is part of the conceptual THI even though the observable implementation in this study is reduced.

5. Relation to adjacent temporal-hallucination metrics

Outside the tumbling-E setting, several papers address temporality in hallucination assessment without defining THI itself. In fictional character role-play, the closest explicit analog is Temporal Hallucination Rate (THR), defined as “the number of atomic facts associated with temporal hallucination for every 100 responses.” There, temporal hallucination means stating facts from future story events or outside the relevant temporal window of the story. The RoleFact method reduced GPT-3.5 THR in scene-grounded interviews from 26.5 to 14.7, which the paper states is a 44.5% relative reduction; however, that work explicitly does not define a metric called THI (Sadeq et al., 2024).

In multimodal video understanding, one paper frames temporal hallucination as event-level errors about when events occur or how events are ordered, and evaluates models with timestamp prediction and order prediction rather than a single scalar index. Another video-LLM paper distinguishes spatial hallucination from temporal hallucination, defines the latter as outputs that contradict event order or causal relations, and evaluates mitigation with benchmarks such as VidHalluc, VideoHallucer, and EventHallusion rather than with THI (Sun et al., 2024). A later training-free decoding method, SEASON, remains in this same category: it is directly about temporal hallucination in VideoLLMs, but it does not define or cite THI (Wu et al., 4 Dec 2025).

Detector-oriented work also produces THI-adjacent quantities. A quickest-change-detection formulation defines onset

θ=min{t:yt=1}\theta = \min\{t : y_t = 1\}

and evaluates streaming monitors with average run length and expected detection delay,

ARL(τ)=E[τ],EDD(τ)=Eθ[(τθ)+].ARL(\tau) = E_{\infty}[\tau], \qquad EDD(\tau) = E_{\theta}\big[(\tau-\theta)^+\big].

That paper argues that classifier metrics such as token-level AUC conceal the operational quantity that matters in streaming use: the number of tokens that pass between hallucination onset and alarm. It does not use the THI term, but its EDDEDD is an explicit temporal-responsiveness measure under a false-alarm constraint (Itkin, 10 Jun 2026).

A different line of work analyzes hidden-layer temporal signals during autoregressive generation, applies FFT, and uses the strongest non-DC frequency component as part of a spectral hallucination detector. That paper also does not introduce THI by name, but it can be read as a temporal-signal approach to hallucination scoring rather than a static hidden-state analysis (Li et al., 16 Sep 2025).

Taken together, these papers show that the phrase “temporal hallucination” is used across multiple domains, but the specific construct called THI is currently tied most directly to the tumbling-E temporal-reliability benchmark rather than to a universally standardized scalar metric.

6. Limitations, caveats, and interpretive boundaries

The primary limitation of the tumbling-E THI formulation is that the study could not compute the full index. The authors state that the exported Correct column was non-informative; as a result, drift and persistence were excluded, and only the observable THI derived from delay and timeout could be used (Sandhu et al., 20 Jun 2026). This means that the manuscript establishes a temporal-reliability baseline, but not a full correctness-based THI decomposition.

A second boundary concerns scope. The paper presents a human-only baseline for future comparison with artificial agents, and its human results show fast median RT near 1.5 s, few delays over 3 s, a 6.6% timeout rate, and smooth staircase convergence toward approximately 20/20-level optotypes. These findings support a low temporal-instability baseline, but they do not by themselves define a universal calibration for other tasks, modalities, or model classes. This suggests that cross-domain uses of the term should be made cautiously.

A third caveat is terminological. In adjacent literature, temporality is measured in heterogeneous ways: fact-level temporal hallucination per 100 responses in role-play, event timestamp and order accuracy in video MLLMs, benchmark-specific temporal subtasks in VideoLLM evaluation, token-delay to alarm in sequential detectors, or spectral anomaly signatures in hidden-state trajectories (Sadeq et al., 2024). These are not interchangeable with the tumbling-E observable THI, even when they target related failure modes.

The broader implication is that THI currently names a specific attempt to move beyond ordinary accuracy and ask whether a participant or system is timely, whether it stalls or times out, whether behavior remains smooth and convergent over repeated trials, and whether performance stays stable as difficulty increases. In its present documented form, THI is best understood as a temporally resolved reliability construct whose implemented version is reduced, observable, and explicitly incomplete.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Temporal Hallucination Index (THI).