---
title: Arm-Specific Learning Rates
url: https://www.emergentmind.com/topics/arm-specific-learning-rates
type: topic
---

# Arm-Specific Learning Rates

Arm-specific learning rates refer to the practice of setting or adapting the learning rate parameter for each arm individually in online learning algorithms, particularly in stochastic and adversarial multi-armed bandit (MAB) problems. These rates may depend on the data and the current state of the algorithm, providing adaptivity and potentially yielding improved regret guarantees and empirical performance compared to global, time-homogeneous step sizes. The design, analysis, and implementation of arm-specific learning rates play a central role in contemporary bandit methodologies, with particular relevance to frameworks such as Follow-the-Regularized-Leader (FTRL) and Follow-the-Perturbed-Leader (FTPL).

## 1. Theoretical Foundations

Arm-specific learning rates emerged from the need to balance exploration and exploitation while responding to heterogeneity among arms, variation in rewards, or adversarial loss sequences. Rather than using a single learning rate $\eta_t$ for all arms and all rounds, algorithms may define rates $\eta_{t,i}$ potentially as functions of arm-specific statistics (e.g., estimated losses or probabilities of selection), or more generally as functions of the state of the algorithm.

Theoretical treatments often rely on concepts such as the probability simplex over arms, Markov chain evolution induced by the chosen rates, and trade-offs between stability and penalty components in regret decompositions. The paper "Regret Analysis of a Markov Policy Gradient Algorithm for Multi-arm Bandits" notes that learning rates can be made dependent on the current state of the algorithm, and such Markovian dependence enables refined regret and stability analyses (though the full formulas or theorems are not detailed) [2007.10229].

## 2. Surrogate-Probability Approaches and FTPL

A major challenge in FTPL-style bandits is the lack of closed-form for the true sampling probabilities $w_{t,i}$ of each arm under the perturbation mechanism. Recent work addresses this through computationally tractable surrogate probability functions, which are used both for estimating importance weights and designing adaptive arm-specific learning rates [2606.06043].

The surrogate probabilities $p_{t,i}$ and $q_{t,i}$ are defined as follows:
\[
p_{t,i} = \min\left\{(1+\eta_t \hat L_{t,i})^{-\alpha},\, \sigma_{t,i}^{-1}\right\},\quad
q_{t,i} = \min\left\{(1+\eta_{t-1} \hat L_{t,i})^{-\alpha},\, \sigma_{t,i}^{-1}\right\}
\]
where $\hat L_{t,i}$ is the cumulative (importance-weighted) loss of arm $i$, $\sigma_{t,i}$ is its rank by loss, and $\alpha>1$ is a perturbation parameter. These quantities serve as $O(1)$-computable proxies for $w_{t,i}$ and can guide the update of learning rates in a probability-adaptive manner.

## 3. Learning Rate Update Mechanisms

The update of learning rates in surrogate-probability-driven FTPL methods uses a stability–penalty matching principle. For each round, surrogate-based quantities $z_t$ and $h_t$ are computed (defined in terms of $q_{t,i}$), and the learning rate is updated as:
\[
\beta_{t+1} = 
\begin{cases}
\min\left\{2^{1/\alpha}\,\beta_t,\, \beta_t+\frac{z_t}{\beta_t\,h_t}\right\}, & \alpha \geq 2\\
\beta_t + \max\left\{\frac{z_t}{\beta_t\,h_t},\, \frac{4}{(2^{1/\alpha}-1)\,t}\right\}, & 1 < \alpha < 2
\end{cases}
,\quad \eta_t = 1/\beta_t
\]
No convex optimization steps are necessary, maintaining computational efficiency. These step sizes are global across arms for each round, but the surrogate measures retain strong per-arm adaptivity due to dependence on $\hat L_{t,i}$ [2606.06043].

## 4. Best-of-Both-Worlds Regret Guarantees

Adaptive learning rates tied to arm-specific surrogate probabilities achieve the so-called best-of-both-worlds (BOBW) regret bounds. For any Pareto shape parameter $\alpha > 1$, the FTPL algorithm with these rates enjoys:
- An adversarial regret bound:
  \[
  \Reg(T) \leq O\left(\frac{\alpha^2}{\alpha-1}\sqrt{K T}\right)
  \]
- For bandit settings with self-bounding constraints (e.g., corrupt stochastic rewards):
  \[
  \Reg(T) \leq O\bigl(\omega(\Delta)\ln T + \sqrt{C \omega(\Delta)\ln T} + c(\alpha, K)\bigr)
  \]
  where $\omega(\Delta) = O(\alpha^4 K / (\alpha-1)^2 \Delta_{\min})$, and $c(\alpha,K)$ is a constant [2606.06043].

A plausible implication is that arm-dependent learning rates via surrogates can yield both efficient procedures and optimal order regret across regimes, subsuming prior adversarial or stochastic-only approaches.

## 5. Algorithmic Implementation and Efficiency

Implementation with arm-specific surrogates proceeds with the following steps each round:
- Compute the ranks $\sigma_{t,i}$ of the cumulative losses $\hat L_{t,i}$
- For each arm, compute $q_{t,i}$
- Compute $z_{t-1}$ and $h_{t-1}$ using the surrogates
- Update $\beta_t$ and $\eta_t$ by the closed-form rule
- Draw perturbations and select $i_t$ using FTPL
- Perform importance-weighted loss updates using $O(\log t)$-time conditional resampling

All quantities, including the surrogates for probabilities and the learning rates, are closed-form and require no per-round optimization, which preserves scalability even in large action spaces [2606.06043].

## 6. Comparisons with Prior Methods

Key points of comparison include:
- Previous FTPL methods with fixed learning rates required tuning and achieved BOBW only for the special case $\alpha=2$. The surrogate-based construction generalizes to all $\alpha>1$ [2606.06043].
- FTRL with Tsallis entropy can match BOBW bounds via arm-probability-dependent rates, but the approach requires solving a convex program each round to compute arm selection probabilities, increasing computational burden.
- Surrogate-based rates maintain FTPL’s “optimization-free” character while achieving the same adaptive guarantees as FTRL with explicit arm probability weighting.

## 7. Extensions and Broader Applications

The surrogate-probability methodology for arm-specific or probability-dependent learning rates can be extended:
- To settings where probabilities are unavailable or hard to compute, as in heavy-tailed bandits, by using surrogates to threshold or adapt skipping arms
- To combinatorial, structured, or graph-structured bandit models, where FTRL incurs projection costs
- To contextual bandits and expert advice regimes, where unified policies benefit from the same surrogate‐SPM update rule

These extensions demonstrate the versatility and potential of surrogate-probability-driven, arm-specific learning rate schemes for both classic and modern bandit applications [2606.06043].

Source: https://www.emergentmind.com/topics/arm-specific-learning-rates