Papers
Topics
Authors
Recent
Search
2000 character limit reached

LinearAPT: An Adaptive Algorithm for the Fixed-Budget Thresholding Linear Bandit Problem

Published 10 Mar 2024 in cs.LG and stat.ML | (2403.06230v1)

Abstract: In this study, we delve into the Thresholding Linear Bandit (TLB) problem, a nuanced domain within stochastic Multi-Armed Bandit (MAB) problems, focusing on maximizing decision accuracy against a linearly defined threshold under resource constraints. We present LinearAPT, a novel algorithm designed for the fixed budget setting of TLB, providing an efficient solution to optimize sequential decision-making. This algorithm not only offers a theoretical upper bound for estimated loss but also showcases robust performance on both synthetic and real-world datasets. Our contributions highlight the adaptability, simplicity, and computational efficiency of LinearAPT, making it a valuable addition to the toolkit for addressing complex sequential decision-making challenges.

Authors (3)
Definition Search Book Streamline Icon: https://streamlinehq.com
References (16)
  1. Combinatorial pure exploration of multi-armed bandits. In Z. Ghahramani, M. Welling, C. Cortes, N. Lawrence, and K.Q. Weinberger (eds.), Advances in Neural Information Processing Systems, volume 27. Curran Associates, Inc., 2014. URL https://proceedings.neurips.cc/paper_files/paper/2014/file/e56954b4f6347e897f954495eab16a88-Paper.pdf.
  2. The influence of shape constraints on the thresholding bandit problem. In Jacob Abernethy and Shivani Agarwal (eds.), Proceedings of Thirty Third Conference on Learning Theory, volume 125 of Proceedings of Machine Learning Research, pp.  1228–1275. PMLR, 09–12 Jul 2020. URL https://proceedings.mlr.press/v125/cheshire20a.html.
  3. Problem dependent view on structured thresholding bandit problems. In Marina Meila and Tong Zhang (eds.), Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pp.  1846–1854. PMLR, 18–24 Jul 2021. URL https://proceedings.mlr.press/v139/cheshire21a.html.
  4. Gamification of pure exploration for linear bandits. In Hal Daumé III and Aarti Singh (eds.), Proceedings of the 37th International Conference on Machine Learning, volume 119 of Proceedings of Machine Learning Research, pp.  2432–2442. PMLR, 13–18 Jul 2020. URL https://proceedings.mlr.press/v119/degenne20a.html.
  5. Doubly robust policy evaluation and learning. arXiv preprint arXiv:1103.4601, 2011.
  6. Thresholding bandit for dose-ranging: The impact of monotonicity, 2018.
  7. Active learning for level set estimation. In Proceedings of the Twenty-Third International Joint Conference on Artificial Intelligence, IJCAI ’13, pp.  1344–1350. AAAI Press, 2013. ISBN 9781577356332.
  8. Thresholding graph bandits with grapl. In Silvia Chiappa and Roberto Calandra (eds.), Proceedings of the Twenty Third International Conference on Artificial Intelligence and Statistics, volume 108 of Proceedings of Machine Learning Research, pp.  2476–2485. PMLR, 26–28 Aug 2020. URL https://proceedings.mlr.press/v108/lejeune20a.html.
  9. An optimal algorithm for the thresholding bandit problem. In Maria Florina Balcan and Kilian Q. Weinberger (eds.), Proceedings of The 33rd International Conference on Machine Learning, volume 48 of Proceedings of Machine Learning Research, pp.  1690–1698, New York, New York, USA, 20–22 Jun 2016. PMLR. URL https://proceedings.mlr.press/v48/locatelli16.html.
  10. Top-m identification for linear bandits. In Arindam Banerjee and Kenji Fukumizu (eds.), Proceedings of The 24th International Conference on Artificial Intelligence and Statistics, volume 130 of Proceedings of Machine Learning Research, pp.  1108–1116. PMLR, 13–15 Apr 2021. URL https://proceedings.mlr.press/v130/reda21a.html.
  11. Best-arm identification in linear bandits. CoRR, abs/1409.6110, 2014a. URL http://arxiv.org/abs/1409.6110.
  12. Best-arm identification in linear bandits. In Z. Ghahramani, M. Welling, C. Cortes, N. Lawrence, and K.Q. Weinberger (eds.), Advances in Neural Information Processing Systems, volume 27. Curran Associates, Inc., 2014b. URL https://proceedings.neurips.cc/paper_files/paper/2014/file/f387624df552cea2f369918c5e1e12bc-Paper.pdf.
  13. Thresholding bandit with optimal aggregate regret. In H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alché-Buc, E. Fox, and R. Garnett (eds.), Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc., 2019. URL https://proceedings.neurips.cc/paper_files/paper/2019/file/9a0684d9dad4967ddd09594511de2c52-Paper.pdf.
  14. lil’hdoc: An algorithm for good arm identification under small threshold gap. arXiv preprint arXiv:2401.15879, 2024.
  15. Fast online inference for nonlinear contextual bandit based on generative adversarial network. arXiv preprint arXiv:2202.08867, 2022.
  16. Differential good arm identification. arXiv preprint arXiv:2303.07154, 2023.

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.

Tweets

Sign up for free to view the 2 tweets with 0 likes about this paper.