Determine whether calibration and review budgets are stable across Jev releases
Determine whether the calibration performance and human-review budgets reported for Jev are properties of the Jev model and coding schema generally or only of the audited Jev release.
References
The fourth is a repeated audit across model versions, with the same reference set and the same metric code. It is the only way to learn whether the calibration and the review budgets reported here are properties of the model or of one release, and the released code is designed to make it routine.
— Calibrated Decisions at Scale: Converting Police Crash Narratives into Probabilistic Crash Variables with a System One Model (Jev)
(2609.24052 - Rafe et al., 21 Sep 2026) in Section 6, final paragraph before the appendices