Quantify discovery-stage sensitivity and the resulting undercount
Estimate the sensitivity of the model-based discovery process, including the number and types of errors that none of the twelve discovery passes proposes, in order to quantify the total amount by which the verified-failure counts understate errors in notes produced by the three commercial ambient scribes.
References
Discovery has the matching limit, and it is unmeasured: discovery is model-based, its sensitivity for errors no pass proposed is not estimated, and the discovering model's own importance filter dropped 7,780 of the 13,678 candidates before any skeptic saw them, a bin nothing in this paper audits.
— One note in three: a verified census of three deployed AI scribes, and the instrument that counted it
(2608.31017 - Fox et al., 31 Aug 2026) in Section 6, “Limitations”