Finite-mean stabilization for deterministic tie-breaking with tied suboptimal means

Establish finite-mean control of the time required for the unguarded deterministic-tie-breaking Bernoulli β-EB-TCI algorithm to enter the stabilized regime when multiple suboptimal arms have equal means, and thereby derive the sharp expected stopping-time bound for that setting.

Background

The paper proves a sharp expected stopping-time bound for the guarded Bernoulli β-EB-TCI algorithm and for an unguarded tie-breaking variant under pairwise-distinct arm means, assuming a finite-mean sufficient-exploration property. For the original unguarded rule with fixed deterministic tie-breaking, the authors obtain only a high-probability stopping bound when suboptimal arms may share the same mean.

The unresolved difficulty is to show that the time needed for the empirical leader to become permanently correct and for the best-arm sampling fraction to approach β has sufficiently light tails to possess a finite mean. The paper proves that one-step comparisons of penalized challenger indices are insufficient by themselves, but explicitly leaves open whether a more global, path-dependent analysis can establish finite-mean exploration and the corresponding expected sample-complexity guarantee.

References

For the unguarded deterministic-tie-breaking rule, the analogous expected bound would require finite-mean control of the time to enter the stabilized regime when several suboptimal arms share the same mean. We do not resolve that problem here.

— Sharp Non-Asymptotic Analysis of the Penalized Challenger in $β$-EB-TCI for Bernoulli Bandits  (2610.01951 - Nguyen et al., 1 Oct 2026) in Section 6, paragraph “The expected bound for the unguarded algorithm remains unresolved”