Establish the prevalence of benchmark claims that fail reliability checks
Establish what fraction of published comparative claims on small-cohort benchmarks with few independent units, concentrated ground truth, pooled ratio metrics, and a single inherited partition would fail the four-axis reliability protocol.
References
We have not shown what fraction of published claims on such benchmarks would fail, and no single case study could.
Whether they are commonly satisfied is an empirical question we have not settled at scale — auditing a public archive against criteria of this kind is itself a piece of work [45] — but they are satisfied by every public whole-slide quality-control resource we are aware of, and the first three are satisfied by a large part of medical image segmentation, where exhaustive annotation is expensive [27] for the same reasons.