Papers
Topics
Authors
Recent
Search
2000 character limit reached

On the Power of Adaptivity for ε\varepsilon-Best Arm Identification in Linear Bandits

Published 15 May 2026 in cs.LG | (2605.15663v1)

Abstract: We study the minimax sample complexity of ε\varepsilon-best arm identification in linear bandits. Given a compact action set X\mathcal{X} that spans R<sup>d\mathbb{R}<sup>d and an unknown reward vector θR<sup>dθ\in\mathbb{R}<sup>d, the goal is to output an arm x^X\widehat{x}\in\mathcal{X} such that x^,θmaxxXx,θε\langle \widehat{x},θ\rangle \ge \max_{x\in\mathcal{X}} \langle x,θ\rangle - \varepsilon with probability at least $1-δ$, using as few samples as possible. First, we present a non-adaptive fixed-design method with sample complexity O!(dlog(1/δ)ε<sup>2+w(X)<sup>2ε<sup>2)\mathcal{O}!\left(\frac{d\log(1/δ)}{\varepsilon<sup>2}+\frac{w(\mathcal{X})<sup>2}{\varepsilon<sup>2}\right), where w(X)w(\mathcal{X}) is a Gaussian width term dependent on X\mathcal{X}, and we prove a matching lower bound Ω!(dlog(1/δ)ε<sup>2+w(X)<sup>2ε<sup>2)Ω!\left(\frac{d\log(1/δ)}{\varepsilon<sup>2}+\frac{w(\mathcal{X})<sup>2}{\varepsilon<sup>2}\right) for all non-adaptive fixed-design methods. We then turn to adaptive sampling. We raise an important structural question: beyond the canonical basis, are there structured action sets for which adaptivity yields only logarithmic-factor improvements over the optimal non-adaptive rate? We answer in the affirmative for several natural action sets, namely the hypercube, the 2\ell_2 ball, mm-sets, and multi-task multi-armed bandits. Finally, we provide the first construction of an action set X\mathcal{X} for which adaptivity yields a polynomial-factor improvement over every non-adaptive algorithm. A key ingredient behind this separation is an 2\ell_2-norm estimation subroutine: we design an adaptive algorithm that uses O!(dlog(1/δ)ε<sup>2)\mathcal{O}!\left(\frac{d\log(1/δ)}{\varepsilon<sup>2}\right) samples from the unit 2\ell_2 ball in R<sup>d\mathbb{R}<sup>d and outputs an estimate r^\widehat r satisfying r^θ2ε|\widehat r-|θ|_2|\le \varepsilon with probability at least $1-δ$, where θθ is the unknown reward vector.

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.