---
title: 'TS_Adam: Non-Stationary Forecasting Optimizer'
url: https://www.emergentmind.com/topics/ts_adam
type: topic
---

# TS_Adam: Non-Stationary Forecasting Optimizer

Searching arXiv for the specified paper and closely related optimizer/time-series forecasting context.
Attempting to retrieve arXiv metadata for 2603.10095.
TS_Adam is a lightweight variant of Adam designed for time-series forecasting under non-stationary objectives, especially settings with distributional drift in which the data distribution evolves over time. Its defining modification is the removal of Adam’s second-order bias correction from the learning-rate computation, while retaining the same two-moment recursion for gradients. In the reported formulation, this increases responsiveness to shifting loss landscapes, preserves the optimizer core structure, introduces no additional hyperparameters, and integrates directly into existing forecasting pipelines. The method is presented as a practical optimization strategy for real-world forecasting scenarios involving non-stationary data, with consistent gains across long- and short-term forecasting tasks [2603.10095].

## 1. Problem setting and optimization perspective

Time-series forecasting often operates in a non-stationary regime rather than under a stationary objective. The central difficulty emphasized for TS_Adam is distributional drift: the loss surface changes over time because the underlying data-generating process evolves. In this setting, an optimizer must not only attenuate gradient noise but also track a moving optimum [2603.10095].

The motivating claim is that adaptive optimizers such as Adam are typically designed for stationary objectives, and that this design assumption can become limiting when applied to forecasting workloads with persistent drift. In particular, the second-order bias correction in Adam is identified as a mechanism that can reduce responsiveness when the loss landscape changes. TS_Adam is proposed as a direct response to that diagnosis: rather than redesigning adaptive optimization from first principles, it modifies a single component of Adam’s update rule.

This framing places TS_Adam squarely within online and drifting-objective optimization. A plausible implication is that its intended advantage is not generic acceleration in all regimes, but improved tracking behavior when seasonality, regime changes, or abrupt concept drift alter the effective objective over training.

## 2. Formal definition relative to Adam

Let $f_t(\theta)$ denote the loss at step $t$, and let the stochastic gradient be
$$
g_t = \nabla_\theta f_t(\theta_{t-1}).
$$

Standard Adam maintains first- and second-moment estimates:
$$
m_t = \beta_1 m_{t-1} + (1-\beta_1) g_t,
\qquad
v_t = \beta_2 v_{t-1} + (1-\beta_2) g_t^2,
$$
with decay rates $\beta_1,\beta_2 \in [0,1)$ and defaults $\beta_1=0.9$, $\beta_2=0.999$. Adam then applies bias correction to both moments:
$$
\hat m_t = \frac{m_t}{1-\beta_1^t},
\qquad
\hat v_t = \frac{v_t}{1-\beta_2^t},
$$
and updates parameters by
$$
\theta_t = \theta_{t-1} - \alpha \frac{\hat m_t}{\sqrt{\hat v_t}+\epsilon},
$$
where $\alpha$ is the base learning rate and $\epsilon$ is a small constant such as $10^{-8}$ for numerical stability [2603.10095].

TS_Adam follows exactly the same two-moment scheme except that it omits the second-order bias correction in $\hat v_t$. The recursions are
$$
m_t = \beta_1 m_{t-1} + (1-\beta_1) g_t,
\qquad
v_t = \beta_2 v_{t-1} + (1-\beta_2) g_t^2,
$$
and only the first moment is bias-corrected:
$$
\hat m_t = \frac{m_t}{1-\beta_1^t},
\qquad
\hat v_t = v_t.
$$
The resulting parameter update is
$$
\theta_t = \theta_{t-1} - \alpha \frac{\hat m_t}{\sqrt{v_t}+\epsilon}.
$$

| Component | Adam | TS_Adam |
|---|---|---|
| First moment | $\hat m_t = m_t/(1-\beta_1^t)$ | $\hat m_t = m_t/(1-\beta_1^t)$ |
| Second moment | $\hat v_t = v_t/(1-\beta_2^t)$ | $\hat v_t = v_t$ |
| Update | $\theta_t=\theta_{t-1}-\alpha \hat m_t/(\sqrt{\hat v_t}+\epsilon)$ | $\theta_t=\theta_{t-1}-\alpha \hat m_t/(\sqrt{v_t}+\epsilon)$ |

The modification is therefore narrow but structurally consequential. TS_Adam is not presented as a new optimizer family with extra state, auxiliary schedules, or added tuning knobs; it preserves Adam’s core mechanism while deleting a single correction term.

## 3. Effective step size and dynamic-regret motivation

The paper rewrites Adam’s update in terms of an effective step-size modulation. In Adam,
$$
\eta_t^{\rm eff,Adam} = \frac{\sqrt{1-\beta_2^t}}{1-\beta_1^t},
$$
so the per-step update can be viewed as $\alpha\,\eta_t^{\rm eff,Adam}\,m_t/\sqrt{v_t}$. In TS_Adam, the factor $\sqrt{1-\beta_2^t}$ is removed, yielding
$$
\eta_t^{\rm eff,TS} = \frac{1}{1-\beta_1^t}
\;>\;
\eta_t^{\rm eff,Adam}
\quad\text{for } t>0.
$$
This expresses the method’s core heuristic: removing second-order bias correction raises the effective step size earlier in training [2603.10095].

The theoretical discussion is formulated through dynamic regret,
$$
R(T)=\sum_{t=1}^T \bigl[f_t(\theta_t)-f_t(\theta_t^*)\bigr],
$$
where $\theta_t^*$ is the minimizer of $f_t$. A standard bound, cited as Theorem 2, is given in the form
$$
R(T)\le
\underbrace{\tfrac{1}{2\eta_1}\|\theta_1-\theta_1^*\|^2}_{\text{init.\ error}}
+
\underbrace{\tfrac{G^2}{2}\sum_{t=1}^T \eta_t}_{\text{noise term}}
+
\underbrace{\delta \sum_{t=1}^T \frac{1}{\eta_t}+\cdots}_{\text{drift term}},
$$
where $\eta_t=\alpha\cdot\eta_t^{\rm eff}$, $G$ bounds the gradients, and $\delta$ bounds the per-step shift $\|\theta_{t+1}^*-\theta_t^*\|$.

The interpretation given is asymmetric across regimes. In stationary settings, choosing $\eta_t^{\rm eff}<1$ helps control the noise term. In non-stationary settings, however, the drift term $\sum 1/\eta_t$ encourages a larger $\eta_t^{\rm eff}$ so that the optimizer can track a moving optimum more responsively. Because Adam’s second-order bias correction keeps $\eta_t^{\rm eff,Adam}\ll 1$ for many steps, it is argued to damp updates excessively and accumulate drift-induced regret. TS_Adam, by removing the $\sqrt{1-\beta_2^t}$ factor, is intended to move $\eta_t^{\rm eff}$ closer to $1$ earlier and thus better balance noise suppression against drift adaptation.

## 4. Algorithmic form and integration into forecasting models

The algorithmic definition is deliberately minimal. In pseudo-Python/LaTeX form, the procedure is:

```python
Input: Base LR α, β1, β2, ε, θ0
m ← 0, v ← 0, t ← 0
while not converged:
    t ← t + 1
    g ← ∇ f_t(θ_{t-1})
    m ← β1 m + (1 - β1) g
    v ← β2 v + (1 - β2) g^2
    m_hat ← m / (1 - β1^t)
    v_hat ← v   # no bias-correction on 2nd moment
    θ_t ← θ_{t-1} - α m_hat / (sqrt(v_hat) + ε)
return θ_t
```

Integration into existing forecasting systems is described as trivial. In any PyTorch or TensorFlow model, including MICN, the standard Adam optimizer call can be replaced by a custom TS_Adam class that omits the second-moment bias correction. No new hyperparameters are introduced, and all model-specific hyperparameters remain identical [2603.10095].

The reported experimental defaults are the same as Adam’s: learning rate $\alpha \in \{10^{-4},10^{-3}\}$ tuned by grid, $\beta_1=0.9$, $\beta_2=0.999$, and $\epsilon=10^{-8}$. This point is operationally important because it distinguishes TS_Adam from methods whose empirical gains depend on optimizer-specific retuning. A common source of confusion is whether the method alters Adam’s entire adaptivity mechanism; it does not. The first-moment correction remains intact, and the only omitted term is the second-order bias correction.

## 5. Empirical evaluation in forecasting benchmarks

The reported evaluation covers multiple architectures and datasets. The models are MICN, PatchTST, and SegRNN. The datasets are ETTh1 and ETTh2 for hourly data, ETTm1 and ETTm2 for 15-minute data, plus ECL and Weather. Train/validation/test splits are 60\%/20\%/20\% for ETT and 70\%/20\%/10\% for ECL and Weather. Forecast horizons are $\{96, 192, 336, 720\}$ steps. Experiments use 3 independent runs with fixed seeds, and the metrics are MSE, MAE, and SMAPE [2603.10095].

For MICN, averaged across the four horizons, the key numerical results are as follows:

| Dataset | Adam | TS_Adam |
|---|---|---|
| ETTh1 | MSE $=0.569\pm0.023$, MAE $=0.528\pm0.010$ | MSE $=0.441\pm0.005$ ($\downarrow 22.4\%$), MAE $=0.460\pm0.003$ ($\downarrow 12.9\%$) |
| ETTh2 | MSE $=0.581\pm0.027$, MAE $=0.526\pm0.014$ | MSE $=0.533\pm0.010$ ($\downarrow 8.3\%$), MAE $=0.501\pm0.005$ ($\downarrow 4.8\%$) |
| ETTm1, ETTm2, ECL, Weather | Average MSE and MAE baseline | Average MSE $\downarrow 3$–$4\%$, MAE $\downarrow 2$–$3\%$ |

Across all ETT subsets with MICN, TS_Adam yields an average reduction of 12.8\% in MSE and 5.7\% in MAE over Adam, with all improvements statistically significant under Bonferroni-corrected $t$-tests. The abstract characterizes the gains as consistent across long- and short-term forecasting tasks, and the benchmark design supports that description because it spans several horizons, frequencies, and domains.

A plausible implication is that the optimizer effect is not restricted to a single architectural bias: MICN, PatchTST, and SegRNN represent distinct forecasting backbones, while ETT, ECL, and Weather expose different temporal structures. The paper’s strongest quantitative claims, however, are reported for MICN on the ETT datasets.

## 6. Practical properties, operating regime, and scope

The practical guidance emphasizes TS_Adam as a drop-in replacement. It requires no change to model code beyond swapping the optimizer, adds no hyperparameters, and preserves the default setting $\alpha=10^{-3}$, $\beta_1=0.9$, $\beta_2=0.999$, $\epsilon=10^{-8}$ for most tasks. The method is reported to be robust to batch size in the range 16–64, to different learning-rate schedules including cosine, step, and delayed-decay, and to weight-decay choices [2603.10095].

Its optimization profile is described as slightly higher initial oscillation but faster convergence and lower final loss in non-stationary tasks. It is noted as particularly effective where seasonality or abrupt concept drift is strong, with electricity and meteorology given as examples. The computational overhead is also explicitly addressed: TS_Adam requires no extra memory and is approximately 8\% cheaper per step because it performs one fewer vector division.

These claims delimit the intended scope of the method. TS_Adam is not presented as a universally superior replacement for Adam in all optimization settings. The theoretical discussion explicitly distinguishes stationary from non-stationary objectives, and the empirical evidence is centered on forecasting domains in which distribution shifts are material. This suggests that its main contribution lies in rebalancing adaptivity toward tracking rather than toward conservative early-step damping.

The implementation is publicly available at the repository listed with the paper, which further supports direct adoption in existing forecasting workflows.

Source: https://www.emergentmind.com/topics/ts_adam