Regret theory for the fresh-intercept recursion

Establish regret guarantees for the fresh-intercept recursion used by Odds-Ratio Thompson Sampling in batched binary-reward bandits.

Background

The paper evaluates Odds-Ratio Thompson Sampling empirically in stationary, level-varying, and contrast-varying environments, but does not provide a theoretical regret analysis for its defining update: carrying the joint contrast posterior while fitting and marginalizing a fresh batch-specific intercept. The authors explicitly identify this missing theory as an open problem.

References

Five bounds remain. The evidence is from online services, so clinical time trends are the obvious next test, at batch sizes where Table~\ref{tab:fidelity} may not favour the Gaussian state; the simulations all use one forty-batch loop with one disturbance shape each; sparse adaptive allocations and larger arm sets are untested; stopping and arm-dropping rules are not evaluated; and the regret theory of the fresh-intercept recursion is open.

— Odds-Ratio Thompson Sampling: A Specification and Design Guide for Contrast-Based Multi-Armed Bandits  (2609.19709 - Kim, 17 Sep 2026) in Discussion, final paragraph (“Five bounds remain”)