Papers
Topics
Authors
Recent
Search
2000 character limit reached

Variance-Aware Linear UCB with Deep Representation for Neural Contextual Bandits

Published 8 Nov 2024 in cs.LG and stat.ML | (2411.05979v2)

Abstract: By leveraging the representation power of deep neural networks, neural upper confidence bound (UCB) algorithms have shown success in contextual bandits. To further balance the exploration and exploitation, we propose Neural-σ<sup>2\sigma<sup>2-LinearUCB, a variance-aware algorithm that utilizes σ<sup>2t\sigma<sup>2_t, i.e., an upper bound of the reward noise variance at round tt, to enhance the uncertainty quantification quality of the UCB, resulting in a regret performance improvement. We provide an oracle version for our algorithm characterized by an oracle variance upper bound σ<sup>2t\sigma<sup>2_t and a practical version with a novel estimation for this variance bound. Theoretically, we provide rigorous regret analysis for both versions and prove that our oracle algorithm achieves a better regret guarantee than other neural-UCB algorithms in the neural contextual bandits setting. Empirically, our practical method enjoys a similar computational efficiency, while outperforming state-of-the-art techniques by having a better calibration and lower regret across multiple standard settings, including on the synthetic, UCI, MNIST, and CIFAR-10 datasets.

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.