Determine the optimal dimension dependence for general-action quantum linear bandits

Determine the optimal dependence on the dimension \(d\) of regret for quantum linear bandits with general, potentially infinite, action sets, between the known \(\Omega(d\log(T/d))\) lower bound and the \(O(d^{3/2}\,\mathrm{polylog}\,T)\) upper bound.

Background

For finite action sets, the paper obtains a nearly linear-in-dd regret upper bound when the number of actions is polynomial in dd. For general action sets, the cited and derived upper bounds are O(d2polylogT)O(d^2\,\mathrm{polylog}\,T) and O(d3/2polylogT)O(d^{3/2}\,\mathrm{polylog}\,T), while the paper's lower-bound construction gives Ω(dlog(T/d))\Omega(d\log(T/d)). The optimal dimension dependence in the general-action setting therefore remains undetermined.

References

For general action sets, the optimal dimension dependence is unresolved. The upper bounds are O(d2\,\mathrm{polylog}\,T) from \citet{wan2023quantum} and O(d{3/2}\,\mathrm{polylog}\,T) from a linear-kernel specialization of the analysis of \citet{hikima2024quantum} (see Appendix~\ref{app:hikima-linear}), while our lower bound is \Omega(d\log(T/d)).

Quantum Multi-Armed Bandits and Linear Bandits: Lower Bounds and Algorithms  (2608.14319 - Liu et al., 14 Aug 2026) in Section 7, Conclusion