Papers
Topics
Authors
Recent
Search
2000 character limit reached

Posterior Tempering Explains Variance Inflation in Linear and Generalized Linear Thompson Sampling

Published 2 Sep 2026 in stat.ML, cs.IT, cs.LG, and math.ST | (2609.01999v1)

Abstract: We study a variant of the Thompson Sampling (TS) algorithm, called αα-TS, for solving stochastic generalized linear bandit problems. Existing analyses of TS require inflating the posterior variance to derive near-optimal regret guarantees. We formalize the idea of variance inflation by introducing αα-TS that uses a fractional or αα-posterior instead of the standard posterior. Our main contribution is to identify general regularity conditions on the prior and reward distributions that enable a regret analysis of αα-TS without assuming any tractable approximation of the posterior distribution, unlike previous works. For a specific choice of αd<sup>1α\propto d<sup>{-1}, our general regret bound yields the best known regret bound of O(d<sup>3/2Tlog</sup>T)O(d<sup>{3/2}\sqrt{T}\log</sup> T) for both the exponential and sub-Gaussian families of reward distributions. We further provide an αα-dependent lower bound showing that the regret constant depends on the product αdαd, and that when αd<sup>1α\propto d<sup>{-1} the regret scales as Ω(d<sup>3/2T)Ω(d<sup>{3/2}\sqrt{T}), explaining the origin of the d<sup>3/2d<sup>{3/2} factor in the upper bound. Our proof technique adapts and combines recent advancements in the analysis of linear bandit problems with first- and second-order posterior concentration theory from the Bayesian statistics literature.

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.

Tweets

Sign up for free to view the 1 tweet with 3 likes about this paper.