Close the logarithmic slack in quantum multi-armed bandit regret

Close the logarithmic slack between the \(\Omega(K\log(T/K))\) minimax regret lower bound and the \(O(K\log T)\) upper bound for quantum multi-armed bandits in the quantum reward-oracle model.

Background

The paper proves an Ω(Klog(T/K))\Omega(K\log(T/K)) minimax regret lower bound for quantum multi-armed bandits and notes that this nearly matches the previously known O(KlogT)O(K\log T) upper bound. The remaining gap concerns the precise logarithmic dependence on the horizon and the factor inside the logarithm. Resolving it would establish the sharp regret scale for quantum multi-armed bandits in this oracle-access model.

References

Several questions remain open. For multi-armed bandits, it remains to close the logarithmic slack between our \Omega(K\log(T/K)) lower bound and the O(K\log T) upper bound.

Quantum Multi-Armed Bandits and Linear Bandits: Lower Bounds and Algorithms  (2608.14319 - Liu et al., 14 Aug 2026) in Section 7, Conclusion