Establish the prevalence of benchmark claims that fail reliability checks

Establish what fraction of published comparative claims on small-cohort benchmarks with few independent units, concentrated ground truth, pooled ratio metrics, and a single inherited partition would fail the four-axis reliability protocol.

Background

The paper demonstrates the reliability protocol on one diffusion-based artifact detector and one whole-slide-image benchmark, finding that several comparative claims do not survive uncertainty from test-set sampling, training stochasticity, partition composition, or undocumented preprocessing. The authors caution that a single case study cannot determine how representative this pattern is.

A cross-benchmark audit applying the protocol to multiple methods and resources is therefore needed to estimate the frequency with which published claims are unsupported by the evaluation design.

References

We have not shown what fraction of published claims on such benchmarks would fail, and no single case study could.

Reliable Benchmarking of Artifact Detection in Computational Pathology: A Reproducibility and Uncertainty Analysis  (2608.30835 - Moutselos et al., 31 Aug 2026) in Section 4.3, Generality and limits of the protocol

Whether they are commonly satisfied is an empirical question we have not settled at scale — auditing a public archive against criteria of this kind is itself a piece of work [45] — but they are satisfied by every public whole-slide quality-control resource we are aware of, and the first three are satisfied by a large part of medical image segmentation, where exhaustive annotation is expensive [27] for the same reasons.

Reliable Benchmarking of Artifact Detection in Computational Pathology: A Reproducibility and Uncertainty Analysis  (2608.30835 - Moutselos et al., 31 Aug 2026) in Section 4.3, Generality and limits of the protocol