Papers
Topics
Authors
Recent
Search
2000 character limit reached

AdEMAMix Optimizer: Dual Momentum for Deep Learning

Updated 19 March 2026
  • AdEMAMix Optimizer is defined as an optimizer for large-scale deep learning that uses two distinct EMA buffers (fast and slow) to overcome single-EMA limitations.
  • It employs a mixing strategy with a fast decay (β₁ ≈ 0.9) and a slow decay (β₃ ≈ 0.9999) that achieves up to a 95% token efficiency gain in language modeling.
  • Empirical studies demonstrate improved convergence rates, reduced model forgetting, and superior downstream performance in language and vision tasks with minimal overhead.

AdEMAMix is an optimizer designed for large-scale deep learning, augmenting the AdamW algorithm with a two-scale momentum mechanism. By maintaining distinct short-term (“fast”) and long-term (“slow”) exponential moving averages (EMAs) of past gradients and mixing them, AdEMAMix directly addresses a fundamental limitation of single-EMA optimizers. Empirical evaluation demonstrates that this approach yields substantially improved convergence rates, greater data efficiency, reduced model forgetting, and better downstream performance in high-iteration, large-data settings such as language modeling and image classification (Pagliardini et al., 2024, Morwani et al., 4 Feb 2025).

1. Motivation and Formulation

Momentum-based algorithms such as Adam and AdamW accumulate past gradients via a single EMA with decay rate β1\beta_1. This mechanism is reactive to either recent gradients (with small β1\beta_1) at the expense of forgetting history, or excessively sluggish (with high β1\beta_1) if trying to incorporate long-term information. No single β1\beta_1 parameter can achieve both rapid responsiveness and long memory simultaneously.

AdEMAMix addresses this by maintaining two EMA buffers:

  • m1m_1 (“fast”) with standard decay rate β1≈0.9\beta_1 \approx 0.9, capturing current curvature.
  • m2m_2 (“slow”) with a high decay (β3≫β1\beta_3 \gg \beta_1, e.g., β3=0.9999\beta_3 = 0.9999), retaining information over thousands of iterations.

The AdEMAMix update, with momentum buffers m1,m2m_1, m_2 and second-moment β1\beta_10: β1\beta_11 The parameter update is: β1\beta_12 Here, β1\beta_13 controls the mix between fast and slow EMA; β1\beta_14 and β1\beta_15 are ramped up via schedules for stability in early training (Pagliardini et al., 2024, Morwani et al., 4 Feb 2025).

2. Connections to Accelerated Gradient Methods

AdEMAMix operationalizes the accelerated stochastic gradient descent (SGD) template in the noise-dominated regime by decoupling the momentum coefficient from the weight on each new gradient. In the classical accelerated-SGD paradigm, the update is: β1\beta_16 For β1\beta_17 and β1\beta_18, provably accelerated convergence, β1\beta_19, is attained in the high-noise regime (Morwani et al., 4 Feb 2025).

By setting β1\beta_10 and β1\beta_11, AdEMAMix mirrors this two-scale mechanism. Unlike Adam, Lion, or MARS, AdEMAMix's explicit two-EMA mixture with independent scheduling more faithfully implements the theoretical acceleration template (Morwani et al., 4 Feb 2025).

3. Implementation, Hyperparameters, and Schedules

Key aspects of AdEMAMix implementation include:

  • Decay Rates: β1\beta_12 (fast), β1\beta_13 (second moment), β1\beta_14 in β1\beta_15 (slow).
  • Mixing Weight β1\beta_16: Governs slow EMA influence; effective values are β1\beta_17. Both β1\beta_18 and β1\beta_19 are ramped up over the total training window β1\beta_10 with linear or “half-life” schedules, preventing instability in early iterations.
  • Step Size β1\beta_11: Standard AdamW schedules apply (warmup + cosine/linear decay).
  • Gradient Clipping: Norm clipping remains beneficial.
  • Memory Overhead: Maintaining β1\beta_12 doubles the first-moment state, unless β1\beta_13 (in which case β1\beta_14 is dropped).
  • Bias Correction: Only β1\beta_15 receives bias correction; the slow buffer β1\beta_16 is slowly ramped and not bias-corrected, avoiding cold-start issues.

Pseudocode mirrors AdamW with the addition of the β1\beta_17 update and mixing step. In practice, failure to schedule large β1\beta_18 or β1\beta_19 can cause unstable jumps (“explosions”) early in training. These issues are remediated via appropriate scheduler design (Pagliardini et al., 2024).

4. Empirical Performance and Comparative Studies

Large-scale experiments substantiate the advantages of AdEMAMix:

  • Language Modeling: On the RedPajama v2 corpus, Transformer architectures with m1m_10M, m1m_11M, and m1m_12B parameters attain the same validation loss as AdamW-trained models using roughly half the number of tokens (m1m_1395% token efficiency gain). For example, a m1m_14B parameter model trained with AdEMAMix on m1m_15B tokens matches AdamW’s performance at m1m_16B tokens.
  • Generalization and Slower Forgetting: In “forgetting” experiments (injecting a held-out batch at time m1m_17), AdEMAMix retains information about the batch over thousands of steps, whereas AdamW forgets rapidly.
  • Architectural Breadth: AdEMAMix halves the steps to convergence for Mamba on FineWeb and yields lower train/test loss for ViTs trained on ImageNet-21k. No systematic benefit is observed where data is extremely scarce or overfitting dominates.
  • Training Overhead: The addition of m1m_18 increases training time per step by less than m1m_19, offset by substantial token efficiency.
  • Model Switching: Switching from AdamW to AdEMAMix mid-training (with β1≈0.9\beta_1 \approx 0.90) immediately improves convergence, especially with earlier switching.
  • In-Context Performance: On zero/few-shot tasks (HellaSwag, ARC, MMLU, PubMedQA, RewardBench), AdEMAMix outperforms AdamW baselines, sometimes by several percentage points.

Table: Empirical Comparison in Language Modeling and Vision Tasks (Pagliardini et al., 2024, Morwani et al., 4 Feb 2025)

Task AdEMAMix AdamW Notes
LM, 1.3B, 101B tokens matches AdamW @ 197B 197B tokens needed β1≈0.9\beta_1 \approx 0.9195% token reduction
Mamba, FineWeb Halves steps - β1≈0.9\beta_1 \approx 0.92, β1≈0.9\beta_1 \approx 0.93
ViT, ImageNet-21k Lower loss Higher loss Large-scale, data-rich; AdEMAMix advantage
ViT, ImageNet-1k (scarce) No improvement - Overfitting regime

AdEMAMix’s closest comparators include AdamW (single EMA), Lion (coordinate-wise sign momentum), MARS (aggregated gradients), and schedule-free Adam variants. Unlike these, AdEMAMix maintains two independently scheduled momenta, matching the two-timescale dynamics required for accelerated SGD in stochastic regimes.

  • AdamW: Ties the current gradient’s weight to the momentum term; cannot separate fast/slow behavior.
  • Lion and MARS: Implement variants of acceleration or sign normalization, but lack explicit two-buffer mixtures or direct theoretical matching to accelerated SGD.
  • Schedule-Free AdamW: Employs a single momentum buffer with time-scheduled β1≈0.9\beta_1 \approx 0.94, but lacks stable acceleration in large-batch regimes.

Ablation studies show that dropping β1≈0.9\beta_1 \approx 0.95 (setting β1≈0.9\beta_1 \approx 0.96) suffices in small-batch/noisy regimes, but in large-batch scenarios, both buffers are required to retain the acceleration and stability properties (Morwani et al., 4 Feb 2025).

6. Extensions: Simplified-AdEMAMix

Building on the dual-EMA design, "Simplified-AdEMAMix" collapses both buffers into a single theory-style EMA. The update: β1≈0.9\beta_1 \approx 0.97

β1≈0.9\beta_1 \approx 0.98

Here, β1≈0.9\beta_1 \approx 0.99 is warmed to m2m_20 and m2m_21 can be fixed or zero. Empirically, for both small and large batch regimes, setting m2m_22 recovers the full performance of the original AdEMAMix (even at scale), eliminating the need for two buffers and reducing implementation complexity. This establishes that the efficacy of AdEMAMix lies in flexibly scheduled momentum, not strictly the duplication of EMA state (Morwani et al., 4 Feb 2025).

7. Limitations, Open Questions, and Research Directions

AdEMAMix excels in high-iteration, high-data regimes. Its strengths are less pronounced for low-iteration or distribution-shifted settings; in such cases, tuning m2m_23 downward or reverting to AdamW may be recommended. Maintaining both buffers increases memory overhead and can introduce early training instability without proper scheduling. Theoretical questions remain regarding the trade-off between noise accumulation, generalization, and the influence of alternative memory kernels (e.g., power-law decay). These observations motivate continued investigation of multi-timescale momentum mechanisms beyond EMAs and deeper analysis of their generalization properties in diverse learning regimes (Pagliardini et al., 2024, Morwani et al., 4 Feb 2025).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (2)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to AdEMAMix Optimizer.