Determine whether calibration and review budgets are stable across Jev releases

Determine whether the calibration performance and human-review budgets reported for Jev are properties of the Jev model and coding schema generally or only of the audited Jev release.

Background

The paper emphasizes that Jev is a closed commercial model whose behavior may change without notice. Although the study records a pinned model identifier and per-call identifiers, it cannot prevent vendor-side changes or guarantee reproducibility if a release is withdrawn. A repeated audit using the same reference set and metric code is therefore needed to establish whether the reported calibration and review-budget results persist across model versions.

References

The fourth is a repeated audit across model versions, with the same reference set and the same metric code. It is the only way to learn whether the calibration and the review budgets reported here are properties of the model or of one release, and the released code is designed to make it routine.

— Calibrated Decisions at Scale: Converting Police Crash Narratives into Probabilistic Crash Variables with a System One Model (Jev)  (2609.24052 - Rafe et al., 21 Sep 2026) in Section 6, final paragraph before the appendices