Sharp bandit rate and the remaining logarithmic factor
Close the remaining $\sqrt{\log K}$ factor in the worst-case bandit regret rate obtained from the exact estimated-RCGF formulation, improving the $O(\sqrt{KT\log K})$ bound to the sharp dependence where appropriate.
References
The worst-case slice $\mu_t=p_t$ caps the per-round value at $K/2$, recovering the $O(\sqrt{KT\log K})$ exponential-weights rate as the small-scale relaxation; closing the remaining $\sqrt{\log K}$ factor remains open (Section~\ref{sec:discussion}).
— The concentration game: Bayesian updating, regret, and information
(2608.18061 - Balsubramani, 18 Aug 2026) in Section 7.2, An identity for the estimated RCGF