Direct LCB-based timing-bandit algorithm

Determine how a direct lower-confidence-bound algorithm can efficiently exploit the consecutive feedback structure of the Timing Bandit problem while preserving favorable regret guarantees.

Background

The paper introduces OE-BCAE, which combines lower-confidence-bound selection with the elimination-based BCAE framework. The authors note that a purely LCB-based design would be more direct but that its interaction with consecutive feedback is not understood.

The unresolved issue is both algorithmic and theoretical: it is unclear how to design a direct LCB variant that uses consecutive observations efficiently and achieves a regret bound comparable to the paper’s structure-aware methods.

References

A direct LCB-based design remains open (see \textsf{SA-LCB} in Sec.~\ref{sec:simulation}), as it is unclear how LCB variants can efficiently utilize the consecutive feedback structure; see e.g..

— Learning When to Update: A Near-Optimal Timing Bandit Approach  (2609.37932 - Lin et al., 29 Sep 2026) in Footnote in Section 5.1, immediately following Theorem 5.1