Effect of larger execution budgets on flaky-test detection

Determine whether flaky-test detection improves materially when the execution budget exceeds approximately 170 reruns, the threshold associated with 95% confidence in prior work.

Background

The study evaluates Shaker and plain re-execution using 100 executions per test, while prior work estimated that approximately 170 reruns are needed to achieve 95% confidence that a passing test is not flaky. The authors therefore leave unresolved whether substantially larger execution budgets would improve detection, particularly given that many ground-truth flaky tests never exhibited a verdict change during the study.

References

Three questions our data could not settle remain open: whether detection improves materially as the execution budget grows past the $\sim$170-rerun threshold of \citeauthor{gruber2021empirical}; whether stress-ng configurations tuned for Python's execution model, rather than inherited from Java and Android, recover the concurrency cases the default configuration misses; and whether running each test within its suite, rather than in isolation, restores part of the flakiness that did not reproduce here.

— Evaluating Shaker for Flaky Test Detection in Python Projects  (2609.25528 - Leal et al., 22 Sep 2026) in Section Conclusions