Papers
Topics
Authors
Recent
Search
2000 character limit reached

Sharper Regret Bounds for Time-Varying Gaussian Process Bandits with Constant Exploration

Published 19 Aug 2026 in stat.ML and cs.LG | (2608.18863v1)

Abstract: We study Bayesian optimization in a time-varying environment where the unknown reward function evolves according to a Gaussian process drift model. Existing GP-UCB analyses in this setting typically require the exploration parameter to grow with the horizon to maintain uniform confidence bounds. Using per-round local confidence events, we show that GP-UCB can instead be run with a constant exploration parameter and obtain an expected-regret bound whose coefficient depends on the drift rate. We also derive a sharper time-varying maximum-information-gain bound. For the squared exponential kernel, it yields γ~T/T=O~(ε<sup>1/2)\tildeγ_T/T=\widetilde{\mathcal O}(ε<sup>{1/2}) and expected average regret O~(ε<sup>1/4)\widetilde{\mathcal O}(ε<sup>{1/4}) in the persistent-drift regime. The same constant-exploration analysis also yields realized-regret guarantees. Simulations support the predicted logarithmic dependence of the bound-suggested exploration parameter on $1/ε$.

Authors (2)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.