Compare dialogue and visit-organized event-table substrates under controlled conditions

Establish through a controlled comparison whether a visit-organized structured-event table preserves or changes longitudinal clinical reasoning performance relative to the dialogue substrate when both representations encode the same source events.

Background

The paper includes a small pilot comparing the generated dialogue with a bare structured-event table built from the same source events. The table retained the events but did not explicitly represent the narrative relations targeted by T3, T4, and T5, preventing equivalent evaluation of those tasks.

The authors therefore leave unresolved whether the observed contribution of the dialogue layer is specific to the particular bare table representation or would persist against a more appropriate visit-organized structured-event table.

References

This delimits what the dialogue substrate adds; it is not a validation of real-note realism, and a controlled comparison against a visit-organized table is left to future work.

ClinTraceBench: Source-Verifiable Longitudinal Clinical Reasoning over EHR-Derived Dialogues  (2609.01111 - Wang et al., 1 Sep 2026) in Appendix, Section “Substrate pilot: dialogue versus a bare event table”