Determine the optimal dimension dependence for general-action quantum linear bandits

Determine the optimal dependence on the dimension \(d\) of regret for quantum linear bandits with general, potentially infinite, action sets, between the known \(\Omega(d\log(T/d))\) lower bound and the \(O(d^{3/2}\,\mathrm{polylog}\,T)\) upper bound.

Background

For finite action sets, the paper obtains a nearly linear-in-dd regret upper bound when the number of actions is polynomial in dd. For general action sets, the cited and derived upper bounds are O(d2 polylog T)O(d^2\,\mathrm{polylog}\,T) and O(d3/2 polylog T)O(d^{3/2}\,\mathrm{polylog}\,T), while the paper's lower-bound construction gives Ω(dlog⁡(T/d))\Omega(d\log(T/d)). The optimal dimension dependence in the general-action setting therefore remains undetermined.

References

For general action sets, the optimal dimension dependence is unresolved. The upper bounds are O(d2\,\mathrm{polylog}\,T) from \citet{wan2023quantum} and O(d{3/2}\,\mathrm{polylog}\,T) from a linear-kernel specialization of the analysis of \citet{hikima2024quantum} (see Appendix~\ref{app:hikima-linear}), while our lower bound is \Omega(d\log(T/d)).

— Quantum Multi-Armed Bandits and Linear Bandits: Lower Bounds and Algorithms  (2608.14319 - Liu et al., 14 Aug 2026) in Section 7, Conclusion

Several directions remain open. First, whether the $\sqrt{d}$ gap between the $O(d{3/2}\sqrt{T})$ upper bound and the $\Omega(d\sqrt{T})$ minimax lower bound can be closed for posterior-sampling algorithms without additional structural assumptions is an important open problem.

— Posterior Tempering Explains Variance Inflation in Linear and Generalized Linear Thompson Sampling  (2609.01999 - Jaiswal et al., 2 Sep 2026) in Section Conclusion, subsection "Future directions"