Establish reliable decision improvements from report-stage reaction simulation

Determine whether the report-stage reaction signal reliably improves the final decision label for models that already perform near ceiling.

Background

The paper evaluates Augur’s synthetic reaction layer through a matched with/without-simulation ablation. Although the mean effect is positive, it falls below the preregistered threshold for declaring that simulation helps, and the benefit varies substantially by solver: the weakest title-only solver improves considerably, whereas a near-ceiling solver regresses slightly. Consequently, the authors explicitly leave unresolved whether the reaction signal can reliably improve final decision labels for already capable models.

References

The matched with/without-simulation arm that prior versions of this work still lacked is now supplied in \Cref{sec:simablation}, and its verdict is a pre-registered null ($+5.5pp$, below the $+6pp$ bar) with a coherent heterogeneity underneath it---so we can now say the report-stage reaction signal is net-neutral-to-helpful and faithful, but we still cannot claim it reliably lifts the final label for a model already near ceiling.

— Augur: A Synthetic Decision Lab for Rehearsing Reactions to Product and Policy Changes  (2609.29952 - Khedar et al., 24 Sep 2026) in Section 6, “Limitations and Evidence Gaps”

Second, the fidelity study of \Cref{sec:fidelity} validates the substance of the synthetic reaction (concern recall $67$--$90\%$), but a separate analysis feeding the model the real contemporaneous reaction in place of the synthetic one shows a calibration gap: the report stage is tuned against the synthetic distribution, so a synthetic-to-real handoff is not yet drop-in, and closing that gap is future work.

— Augur: A Synthetic Decision Lab for Rehearsing Reactions to Product and Policy Changes  (2609.29952 - Khedar et al., 24 Sep 2026) in Section 6, “Limitations and Evidence Gaps”