Test supersession in temporal clinical reasoning evaluation

Establish whether a clinical reasoning evaluation for longitudinal records can determine when later documentation supersedes, rather than merely supplements, earlier clinical information, extending beyond the temporal-order and trend-recognition capabilities measured by TIMER-Eval.

Background

TIMER-Eval is described as the principal existing benchmark for temporal reasoning in longitudinal clinical records. It evaluates whether model outputs respect temporal boundaries, identify trends correctly, and preserve chronological order.

The review explicitly identifies a remaining gap: existing temporal evaluation does not test whether a later clinical record legitimately supersedes an earlier one. This distinction is important because longitudinal records contain updates, revisions, and information that can change the interpretation of prior documentation rather than simply add to it.

References

Temporal reasoning is directly measured mainly by TIMER-Eval, with the supersession problem still open \citep{cui2025}.

— A rubric landscape for evaluating clinical reasoning in large language models: what exists, what is missing, and what needs to be combined  (2610.01938 - Jiang et al., 1 Oct 2026) in Section “What is covered, what is partially covered, and what is missing”