Sharp bandit rate and the remaining logarithmic factor

Close the remaining $\sqrt{\log K}$ factor in the worst-case bandit regret rate obtained from the exact estimated-RCGF formulation, improving the $O(\sqrt{KT\log K})$ bound to the sharp dependence where appropriate.

Background

The paper derives an exact conditional expression for the estimated per-round RCGF in terms of the true losses and observation geometry. Under the worst-case choice of sampling law, this yields the standard small-scale exponential-weights rate with an additional logK\sqrt{\log K} factor. The authors explicitly leave closing that factor unresolved.

References

The worst-case slice $\mu_t=p_t$ caps the per-round value at $K/2$, recovering the $O(\sqrt{KT\log K})$ exponential-weights rate as the small-scale relaxation; closing the remaining $\sqrt{\log K}$ factor remains open (Section~\ref{sec:discussion}).

The concentration game: Bayesian updating, regret, and information  (2608.18061 - Balsubramani, 18 Aug 2026) in Section 7.2, An identity for the estimated RCGF