Extensions beyond the fixed Bernoulli setting

Develop finite-confidence sharp non-asymptotic analyses of the penalized challenger in β-EB-TCI beyond Bernoulli rewards, for online or adaptively chosen β, and for finite-mean stabilization with tied suboptimal means.

Background

The analysis is explicitly restricted to Bernoulli rewards and a fixed leader-sampling probability β. The guarded algorithm provides the expected stopping-time result for Bernoulli instances with a unique best arm, while the unguarded deterministic rule remains unresolved in the presence of tied suboptimal means.

The conclusion identifies broader directions that are not settled by the paper: extending the proof beyond Bernoulli reward distributions, allowing β to be estimated or adapted online, and proving finite-mean stabilization when suboptimal arms have equal means.

References

The proof is Bernoulli-specific with fixed β; extensions beyond Bernoulli rewards, online β adaptation, and finite-mean stabilization with tied suboptimal means remain open.

— Sharp Non-Asymptotic Analysis of the Penalized Challenger in $β$-EB-TCI for Bernoulli Bandits  (2610.01951 - Nguyen et al., 1 Oct 2026) in Section 7, Section 7 conclusion