Efficient Reinforcement Learning for Global Decision Making in the Presence of Local Agents at Scale
Abstract: We study reinforcement learning for global decision-making in the presence of local agents, where the global decision-maker makes decisions affecting all local agents, and the objective is to learn a policy that maximizes the joint rewards of all the agents. Such problems find many applications, e.g. demand response, EV charging, queueing, etc. In this setting, scalability has been a long-standing challenge due to the size of the state space which can be exponential in the number of agents. This work proposes the \texttt{SUBSAMPLE-Q} algorithm where the global agent subsamples local agents to compute a policy in time that is polynomial in . We show that this learned policy converges to the optimal policy in the order of as the number of sub-sampled agents increases, where is the Bellman noise. Finally, we validate the theory through numerical simulations in a demand-response setting and a queueing setting.
- Pseudorandomness of the Sticky Random Walk, 2023. URL https://arxiv.org/abs/2307.11104.
- The Bit Complexity of Dynamic Algebraic Formulas and their Determinants, 2024. URL https://arxiv.org/abs/2401.11127.
- Banach, S. Sur les opérations dans les ensembles abstraits et leur application aux équations intégrales. Fundamenta Mathematicae, 3(1):133–181, 1922. URL http://eudml.org/doc/213289.
- Neuro-Dynamic Programming. Athena Scientific, 1st edition, 1996. ISBN 1886529108. URL https://www.mit.edu/~dimitrib/NDP_Encycl.pdf.
- A Survey of Computational Complexity Results in Systems and Control. Automatica, 36(9):1249–1274, 2000. ISSN 0005-1098. doi: https://doi.org/10.1016/S0005-1098(00)00050-9. URL https://www.sciencedirect.com/science/article/abs/pii/S0005109800000509.
- Finite State Mean Field Games with Major and Minor Players, 2016. URL https://arxiv.org/abs/1610.05408.
- A Probabilistic Approach to Mean Field Games with Major and Minor Players, 2014. URL https://arxiv.org/abs/1409.7141.
- Model-free Mean-Field Reinforcement Learning: Mean-field MDP and mean-field Q-learning. The Annals of Applied Probability, 33(6B):5334 – 5381, 2023. doi: 10.1214/23-AAP1949. URL https://doi.org/10.1214/23-AAP1949.
- Sample complexity of policy-based methods under off-policy sampling and linear function approximation. In Camps-Valls, G., Ruiz, F. J. R., and Valera, I. (eds.), Proceedings of The 25th International Conference on Artificial Intelligence and Statistics, volume 151 of Proceedings of Machine Learning Research, pp. 11195–11214. PMLR, 28–30 Mar 2022. URL https://proceedings.mlr.press/v151/chen22i.html.
- Multi-agent Reinforcement Learning for Networked System Control. In International Conference on Learning Representations, 2020. URL https://openreview.net/forum?id=Syx7A3NFvH.
- Learning Graphon Mean Field Games and Approximate Nash Equilibria. In International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=0sgntlpKDOz.
- Multi-Agent Reinforcement Learning via Mean Field Control: Common Noise, Major Agents and Approximation Properties, 2023. URL https://arxiv.org/abs/2303.10665.
- Multi-Agent Learning with Heterogeneous Linear Contextual Bandits. In Thirty-seventh Conference on Neural Information Processing Systems, 2023. URL https://openreview.net/forum?id=7f6vH3mmhr.
- Asymptotic Minimax Character of the Sample Distribution Function and of the Classical Multinomial Estimator. The Annals of Mathematical Statistics, 27(3):642 – 669, 1956. doi: 10.1214/aoms/1177728174. URL https://doi.org/10.1214/aoms/1177728174.
- Towards Optimal Effective Resistance Estimation. In Thirty-seventh Conference on Neural Information Processing Systems, 2023. URL https://openreview.net/forum?id=KffE8iXAw7.
- On the Complexity of Adversarial Decision Making. In Oh, A. H., Agarwal, A., Belgrave, D., and Cho, K. (eds.), Advances in Neural Information Processing Systems, 2022. URL https://openreview.net/forum?id=pgBpQYss2ba.
- Model-Free Reinforcement Learning with the Decision-Estimation Coefficient. In Thirty-seventh Conference on Neural Information Processing Systems, 2023. URL https://openreview.net/forum?id=Ay3WvSrtpO.
- Correlation Decay in Random Decision Networks, 2009. URL https://arxiv.org/abs/0912.0338.
- Online Nonstochastic Model-Free Reinforcement Learning. In Thirty-seventh Conference on Neural Information Processing Systems, 2023. URL https://openreview.net/forum?id=B5LpWAaBVA.
- Thinking Fast and Slow: Optimization Decomposition Across Timescales. SIGMETRICS Perform. Eval. Rev., 45(2):27–29, oct 2017. ISSN 0163-5999. doi: 10.1145/3152042.3152052. URL https://doi.org/10.1145/3152042.3152052.
- Mean-Field Controls with Q-Learning for Cooperative MARL: Convergence and Complexity Analysis. SIAM Journal on Mathematics of Data Science, 3(4):1168–1196, 2021. doi: 10.1137/20M1360700. URL https://doi.org/10.1137/20M1360700.
- Dynamic Programming Principles for Mean-Field Controls with Learning, 2022a. URL https://pubsonline.informs.org/doi/abs/10.1287/opre.2022.2395?journalCode=opre.
- Mean-Field Multi-Agent Reinforcement Learning: A Decentralized Network Approach, 2022b. URL https://arxiv.org/pdf/2108.02731.pdf.
- Efficient Solution Algorithms for Factored MDPs. J. Artif. Int. Res., 19(1):399–468, oct 2003. ISSN 1076-9757. URL https://dl.acm.org/doi/10.5555/1622434.1622447.
- Hoeffding, W. Probability Inequalities for Sums of Bounded Random Variables. Journal of the American Statistical Association, 58(301):13–30, 1963. ISSN 01621459. URL http://www.jstor.org/stable/2282952.
- Graphon Mean-Field Control for Cooperative Multi-Agent Reinforcement Learning. Journal of the Franklin Institute, 360(18):14783–14805, 2023. ISSN 0016-0032. URL https://doi.org/10.1016/j.jfranklin.2023.09.002.
- Distributed Cooperative Multi-Agent Reinforcement Learning with Directed Coordination Graph. In 2022 American Control Conference (ACC), pp. 3273–3278, 2022. doi: 10.23919/ACC53348.2022.9867152.
- Approximately Optimal Approximate Reinforcement Learning. In Sammut, C. and Hoffman, A. (eds.), Proceedings of the Nineteenth International Conference on Machine Learning (ICML 2002), pp. 267–274, San Francisco, CA, USA, 2002. Morgan Kauffman. ISBN 1-55860-873-7. URL http://ttic.uchicago.edu/~sham/papers/rl/aoarl.pdf.
- Finite-Sample Convergence Rates for Q-Learning and Indirect Algorithms. In Kearns, M., Solla, S., and Cohn, D. (eds.), Advances in Neural Information Processing Systems, volume 11. MIT Press, 1998. URL https://proceedings.neurips.cc/paper_files/paper/1998/file/99adff456950dd9629a5260c4de21858-Paper.pdf.
- An Online Convex Optimization Approach to Real-time Energy Pricing for Demand Response. IEEE Transactions on Smart Grid, 8(6):2784–2793, 2017. doi: 10.1109/TSG.2016.2539948. URL https://ieeexplore.ieee.org/document/7438918.
- Deep Reinforcement Learning for Autonomous Driving: A Survey. IEEE Transactions on Intelligent Transportation Systems, 23(6):4909–4926, 2022. doi: 10.1109/TITS.2021.3054625. URL https://ieeexplore.ieee.org/document/9351818.
- Reinforcement Learning in Robotics: A Survey. The International Journal of Robotics Research, 32(11):1238–1274, 2013. doi: 10.1177/0278364913495721. URL https://doi.org/10.1177/0278364913495721.
- Mean Field Games. Japanese Journal of Mathematics, 2(1):229–260, March 2007. ISSN 1861-3624. doi: 10.1007/s11537-007-0657-8. URL https://link.springer.com/article/10.1007/s11537-007-0657-8.
- Sample Complexity of Asynchronous Q-Learning: Sharper Analysis and Variance Reduction. IEEE Transactions on Information Theory, 68(1):448–473, 2022. doi: 10.1109/TIT.2021.3120096.
- Beyond Black-Box Advice: Learning-Augmented Algorithms for MDPs with Q-Value Predictions. In Thirty-seventh Conference on Neural Information Processing Systems, 2023a. URL https://openreview.net/forum?id=RACcp8Zbr9.
- A Statistical Analysis of Polyak-Ruppert Averaged Q-Learning. In Ruiz, F., Dy, J., and van de Meent, J.-W. (eds.), Proceedings of The 26th International Conference on Artificial Intelligence and Statistics, volume 206 of Proceedings of Machine Learning Research, pp. 2207–2261. PMLR, 25–27 Apr 2023b. URL https://proceedings.mlr.press/v206/li23b.html.
- Distributed Reinforcement Learning in Multi-Agent Networked Systems. CoRR, abs/2006.06555, 2020. URL https://arxiv.org/abs/2006.06555.
- Multi-Agent Reinforcement Learning in Stochastic Networked Systems. In Thirty-fifth Conference on Neural Information Processing Systems, 2021. URL https://proceedings.neurips.cc/paper/2021/hash/412604be30f701b1b1e312 4c252065e6-Abstract.html.
- Online Adaptive Policy Selection in Time-Varying Systems: No-Regret via Contractive Perturbations. In Thirty-seventh Conference on Neural Information Processing Systems, 2023a. URL https://openreview.net/forum?id=hDajsofjRM.
- Learning-augmented Control via Online Adaptive Policy Selection: No Regret via contractive Perturbations. In ACM SIGMETRICS, Workshop on Learning-augmented Algorithms: Theory and Applications 2023, 2023b. URL https://learning-augmented-algorithms.github.io/papers/sigmetrics23-lata-posters-paper5.pdf.
- Littman, M. L. Markov Games as a Framework for Multi-Agent Reinforcement Learning. In Machine learning proceedings, Elsevier, pp. 157–163, 1994. URL https://courses.cs.duke.edu/spring07/cps296.3/littman94markov.pdf.
- Massart, P. The Tight Constant in the Dvoretzky-Kiefer-Wolfowitz Inequality. The Annals of Probability, 18(3):1269 – 1283, 1990. doi: 10.1214/aop/1176990746. URL https://doi.org/10.1214/aop/1176990746.
- Q-learning and Pontryagin’s Minimum Principle. In Proceedings of the 48h IEEE Conference on Decision and Control (CDC) held jointly with 2009 28th Chinese Control Conference, pp. 3598–3605, 2009. doi: 10.1109/CDC.2009.5399753. URL https://ieeexplore.ieee.org/document/5399753.
- Cooperative Multi-Agent Reinforcement Learning: Asynchronous Communication and Linear Function Approximation. In Krause, A., Brunskill, E., Cho, K., Engelhardt, B., Sabato, S., and Scarlett, J. (eds.), Proceedings of the 40th International Conference on Machine Learning, volume 202 of Proceedings of Machine Learning Research, pp. 24785–24811. PMLR, 23–29 Jul 2023. URL https://proceedings.mlr.press/v202/min23a.html.
- The Power of Two Choices in Randomized Load Balancing. PhD thesis, University of California, Berkeley, 1996. AAI9723118.
- A Survey of Distributed Optimization and Control Algorithms for Electric Power Systems. IEEE Transactions on Smart Grid, 8(6):2941–2962, 2017. doi: 10.1109/TSG.2017.2720471. URL https://ieeexplore.ieee.org/document/7990560.
- On the Approximation of Cooperative Heterogeneous Multi-Agent Reinforcement Learning (MARL) Using Mean Field Control (MFC). Journal of Machine Learning Research, 23(1), jan 2022. ISSN 1532-4435. URL https://jmlr.org/papers/volume23/21-1312/21-1312.pdf.
- Naaman, M. On the Tight Constant in the Multivariate Dvoretzky–Kiefer–Wolfowitz Inequality. Statistics & Probability Letters, 173:109088, 2021. ISSN 0167-7152. doi: https://doi.org/10.1016/j.spl.2021.109088. URL https://www.sciencedirect.com/science/article/pii/S016771522100050X.
- The Complexity of Optimal Queuing Network Control. Mathematics of Operations Research, 24(2):293–305, 1999. ISSN 0364765X, 15265471. URL http://www.jstor.org/stable/3690486.
- Powell, W. B. Approximate Dynamic Programming: Solving the Curses of Dimensionality (Wiley Series in Probability and Statistics). Wiley-Interscience, USA, 2007. ISBN 0470171553.
- Learning non-Markovian Decision-Making from State-only Sequences. In Thirty-seventh Conference on Neural Information Processing Systems, 2023. URL https://openreview.net/forum?id=cdlmsnQkZ9.
- Scalable Multi-Agent Reinforcement Learning for Networked Systems with Average Reward. In Proceedings of the 34th International Conference on Neural Information Processing Systems, NIPS’20, Red Hook, NY, USA, 2020a. Curran Associates Inc. ISBN 9781713829546. URL https://proceedings.neurips.cc/paper/2020/file/168efc366c449fab9c2843e9b54e2a18-Paper.pdf.
- Scalable Reinforcement Learning of Localized Policies for Multi-Agent Networked Systems. In Bayen, A. M., Jadbabaie, A., Pappas, G., Parrilo, P. A., Recht, B., Tomlin, C., and Zeilinger, M. (eds.), Proceedings of the 2nd Conference on Learning for Dynamics and Control, volume 120 of Proceedings of Machine Learning Research, pp. 256–266. PMLR, 10–11 Jun 2020b. URL https://proceedings.mlr.press/v120/qu20a.html.
- Reingold, O. Undirected Connectivity in Log-Space. J. ACM, 55(4), sep 2008. ISSN 0004-5411. doi: 10.1145/1391289.1391291. URL https://doi.org/10.1145/1391289.1391291.
- Decentralized Q-learning in Zero-sum Markov Games. In Beygelzimer, A., Dauphin, Y., Liang, P., and Vaughan, J. W. (eds.), Advances in Neural Information Processing Systems, 2021. URL https://openreview.net/forum?id=nhkbYh30Tl.
- Serfling, R. J. Probability Inequalities for the Sum in Sampling without Replacement. The Annals of Statistics, 2(1):39–48, 1974. ISSN 00905364. URL http://www.jstor.org/stable/2958379.
- Q-learning with Nearest Neighbors. In Bengio, S., Wallach, H., Larochelle, H., Grauman, K., Cesa-Bianchi, N., and Garnett, R. (eds.), Advances in Neural Information Processing Systems, volume 31. Curran Associates, Inc., 2018. URL https://proceedings.neurips.cc/paper_files/paper/2018/file/309fee4e541e51de2e41f21bebb342aa-Paper.pdf.
- Shapley, L. S. Stochastic Games*. Proceedings of the National Academy of Sciences, 39(10):1095–1100, 1953. doi: 10.1073/pnas.39.10.1095. URL https://www.pnas.org/doi/abs/10.1073/pnas.39.10.1095.
- Near-optimal time and sample complexities for solving markov decision processes with a generative model. In Bengio, S., Wallach, H., Larochelle, H., Grauman, K., Cesa-Bianchi, N., and Garnett, R. (eds.), Advances in Neural Information Processing Systems, volume 31. Curran Associates, Inc., 2018. URL https://proceedings.neurips.cc/paper_files/paper/2018/file/bb03e43ffe34eeb242a2ee4a4f125e56-Paper.pdf.
- Mastering the Game of Go with Deep Neural Networks and Tree Search. Nature, 529(7587):484–489, January 2016. ISSN 1476-4687. doi: 10.1038/nature16961. URL https://www.nature.com/articles/nature16961.
- Srinivasan, A. Probability and Computing. SIGACT News, 49(3):20–22, Oct 2018. ISSN 0163-5700. doi: 10.1145/3289137.3289142. URL https://doi.org/10.1145/3289137.3289142.
- Tsybakov, A. B. Introduction to Nonparametric Estimation. Springer Publishing Company, Incorporated, 1st edition, 2008. ISBN 0387790519. URL https://link.springer.com/book/10.1007/b13794.
- Q-learning. Machine Learning, 8(3):279–292, May 1992. ISSN 1573-0565. doi: 10.1007/BF00992698. URL https://link.springer.com/article/10.1007/BF00992698.
- Mean Field Multi-Agent Reinforcement Learning. In Dy, J. and Krause, A. (eds.), Proceedings of the 35th International Conference on Machine Learning, volume 80 of Proceedings of Machine Learning Research, pp. 5571–5580. PMLR, 10–15 Jul 2018. URL https://proceedings.mlr.press/v80/yang18d.html.
- Multi-Agent Reinforcement Learning: A Selective Overview of Theories and Algorithms, 2021. URL https://arxiv.org/abs/1911.10635.
- Control of Robotic Mobility-on-Demand Systems: A Queueing-Theoretical Perspective. The International Journal of Robotics Research, 35(1-3):186–203, 2016. doi: 10.1177/0278364915581863. URL https://doi.org/10.1177/0278364915581863.
Paper Prompts
Sign up for free to create and run prompts on this paper.