On the Power of Adaptivity for -Best Arm Identification in Linear Bandits
Abstract: We study the minimax sample complexity of -best arm identification in linear bandits. Given a compact action set that spans and an unknown reward vector , the goal is to output an arm such that with probability at least $1-δ$, using as few samples as possible. First, we present a non-adaptive fixed-design method with sample complexity , where is a Gaussian width term dependent on , and we prove a matching lower bound for all non-adaptive fixed-design methods. We then turn to adaptive sampling. We raise an important structural question: beyond the canonical basis, are there structured action sets for which adaptivity yields only logarithmic-factor improvements over the optimal non-adaptive rate? We answer in the affirmative for several natural action sets, namely the hypercube, the ball, -sets, and multi-task multi-armed bandits. Finally, we provide the first construction of an action set for which adaptivity yields a polynomial-factor improvement over every non-adaptive algorithm. A key ingredient behind this separation is an -norm estimation subroutine: we design an adaptive algorithm that uses samples from the unit ball in and outputs an estimate satisfying with probability at least $1-δ$, where is the unknown reward vector.
Paper Prompts
Sign up for free to create and run prompts on this paper.