Quantify discovery-stage sensitivity and the resulting undercount

Estimate the sensitivity of the model-based discovery process, including the number and types of errors that none of the twelve discovery passes proposes, in order to quantify the total amount by which the verified-failure counts understate errors in notes produced by the three commercial ambient scribes.

Background

The census relies on twelve model-based discovery passes before adversarial verification. Errors that no discovery pass proposes never reach the verification panel, and the discovering model also applies an importance filter that removes candidates before verification. The paper therefore treats all reported counts as undercounts but does not estimate the discovery process’s sensitivity or the magnitude of the unseen error population.

References

Discovery has the matching limit, and it is unmeasured: discovery is model-based, its sensitivity for errors no pass proposed is not estimated, and the discovering model's own importance filter dropped 7,780 of the 13,678 candidates before any skeptic saw them, a bin nothing in this paper audits.

One note in three: a verified census of three deployed AI scribes, and the instrument that counted it  (2608.31017 - Fox et al., 31 Aug 2026) in Section 6, “Limitations”