Effect of larger execution budgets on flaky-test detection
Determine whether flaky-test detection improves materially when the execution budget exceeds approximately 170 reruns, the threshold associated with 95% confidence in prior work.
References
Three questions our data could not settle remain open: whether detection improves materially as the execution budget grows past the $\sim$170-rerun threshold of \citeauthor{gruber2021empirical}; whether stress-ng configurations tuned for Python's execution model, rather than inherited from Java and Android, recover the concurrency cases the default configuration misses; and whether running each test within its suite, rather than in isolation, restores part of the flakiness that did not reproduce here.
— Evaluating Shaker for Flaky Test Detection in Python Projects
(2609.25528 - Leal et al., 22 Sep 2026) in Section Conclusions