Policy Gradients for Optimal Parallel Tempering MCMC (2409.01574v2)
Abstract: Parallel tempering is a meta-algorithm for Markov Chain Monte Carlo that uses multiple chains to sample from tempered versions of the target distribution, enhancing mixing in multi-modal distributions that are challenging for traditional methods. The effectiveness of parallel tempering is heavily influenced by the selection of chain temperatures. Here, we present an adaptive temperature selection algorithm that dynamically adjusts temperatures during sampling using a policy gradient approach. Experiments demonstrate that our method can achieve lower integrated autocorrelation times compared to traditional geometrically spaced temperatures and uniform acceptance rate schemes on benchmark distributions.
- Agrawal, R. The continuum-armed bandit problem. SIAM Journal on Control and Optimization, 33:1926–1951, 1995. ISSN 0363-0129,1095-7138. doi: 10.1137/s0363012992237273. URL http://doi.org/10.1137/s0363012992237273.
- Towards optimal scaling of metropolis-coupled markov chain monte carlo. Statistics and Computing, 21:555–568, 10 2011. doi: 10.1007/s11222-010-9192-1.
- On the containment condition for adaptive markov chain monte carlo algorithms. Adv. Appl. Stat., 21, 01 2011.
- Bojesen, T. A. Policy-guided monte carlo: Reinforcement-learning markov chain dynamics. Physical Review E, 98(6), December 2018. ISSN 2470-0053. doi: 10.1103/physreve.98.063303. URL http://dx.doi.org/10.1103/PhysRevE.98.063303.
- Multi-agent reinforcement learning accelerated mcmc on multiscale inversion problem. ArXiv, abs/2011.08954, 2020. URL https://api.semanticscholar.org/CorpusID:227012696.
- Ensemble samplers with affine invariance. Communications in Applied Mathematics and Computational Science, 5, 01 2010. doi: 10.2140/camcos.2010.5.65.
- Feedback-optimized parallel tempering monte carlo. Journal of Statistical Mechanics: Theory and Experiment, 2006(03):P03018–P03018, March 2006. ISSN 1742-5468. doi: 10.1088/1742-5468/2006/03/p03018. URL http://dx.doi.org/10.1088/1742-5468/2006/03/P03018.
- Kofke, D. On the acceptance probability of replica-exchange monte carlo trials. jcph, 117:6911–, 10 2002. doi: 10.1063/1.1507776.
- Selection of temperature intervals for parallel-tempering simulations. The Journal of chemical physics, 122:206101, 06 2005. doi: 10.1063/1.1917749.
- Lipschitz bandits: Regret lower bounds and optimal algorithms, 2014.
- Adaptive parallel tempering algorithm, 2012.
- Markov-chain monte carlo: Some practical implications of theoretical results. Canadian Journal of Statistics, 26:5 – 20, 12 1997. doi: 10.2307/3315667.
- Coupling and ergodicity of adaptive markov chain monte carlo algorithms. Journal of Applied Probability, 44:458 – 475, 2007. URL https://api.semanticscholar.org/CorpusID:21227921.
- Ergodicity of combocontinuous adaptive mcmc algorithms. Method. Comput. Appl. Prob., 20(2):535–551, jun 2018. ISSN 1387-5841. doi: 10.1007/s11009-017-9574-3. URL https://doi.org/10.1007/s11009-017-9574-3.
- On the ergodicity of the adaptive metropolis algorithm on unbounded domains. The Annals of Applied Probability, 20(6):2178–2203, 2010. ISSN 10505164. URL http://www.jstor.org/stable/20799809.
- Replica-exchange molecular dynamics method for protein folding. Chemical Physics Letters, 314(1):141 – 151, 1999. ISSN 0009-2614. doi: http://dx.doi.org/10.1016/S0009-2614(99)01123-9. URL http://www.sciencedirect.com/science/article/pii/S0009261499011239.
- Dynamic temperature selection for parallel tempering in markov chain monte carlo simulations. Monthly Notices of the Royal Astronomical Society, 455(2):1919–1937, November 2015. ISSN 1365-2966. doi: 10.1093/mnras/stv2422. URL http://dx.doi.org/10.1093/mnras/stv2422.
- Reinforcement learning for adaptive mcmc, 2024.
Collections
Sign up for free to add this paper to one or more collections.
Paper Prompts
Sign up for free to create and run prompts on this paper using GPT-5.