Papers
Topics
Authors
Recent
Search
2000 character limit reached

AI-SARAH: Adaptive Stochastic Gradient Method

Updated 28 January 2026
  • AI-SARAH is an adaptive, implicit, tune-free stochastic recursive gradient optimizer designed for large-scale convex finite-sum minimization.
  • It automatically adjusts step-sizes online by exploiting local smoothness estimates and directional derivatives, enhancing convergence efficiency.
  • The method outperforms traditional SARAH variants by reducing the need for manual hyperparameter tuning while ensuring robust variance reduction.

AI-SARAH is an adaptive, implicit, and tune-free stochastic recursive gradient optimization method designed for large-scale convex finite-sum minimization problems in machine learning. Developed as a practical advancement over the original SARAH and its variants such as SARAH⁺ and iSARAH, AI-SARAH introduces a mechanism to adapt step-sizes online to local smoothness, efficiently leveraging information from stochastic directional derivatives to accelerate convergence without the requirement of manual hyperparameter tuning or prior knowledge of global smoothness or strong convexity parameters (Shi et al., 2021).

1. Problem Setting and Algorithmic Background

AI-SARAH addresses unconstrained finite-sum minimization problems of the form:

P(w)=1n∑i=1nfi(w),P(w) = \frac{1}{n} \sum_{i=1}^n f_i(w),

where w∈Rdw \in \mathbb{R}^d, fi(w)f_i(w) is the loss for sample ii, and P(w)P(w) is the empirical risk. The classical assumption is that each fif_i is convex and twice continuously differentiable. The objective may be either μ\mu-strongly convex (μ>0\mu > 0) or merely convex (μ=0\mu = 0). Traditional smoothness is defined globally: ∥∇fi(x)−∇fi(y)∥≤Li∥x−y∥\|\nabla f_i(x) - \nabla f_i(y)\| \leq L_i \|x-y\|, but AI-SARAH explicitly exploits local smoothness along line segments by considering estimates of the smallest constant w∈Rdw \in \mathbb{R}^d0 satisfying

w∈Rdw \in \mathbb{R}^d1

The SARAH family, introducing the stochastic recursive gradient estimator

w∈Rdw \in \mathbb{R}^d2

motivates AI-SARAH’s efficient variance reduction and memoryless advantages (Nguyen et al., 2017).

2. Adaptive and Implicit Step-Size Mechanism

Unlike SARAH, which employs a fixed step-size w∈Rdw \in \mathbb{R}^d3, AI-SARAH adaptively selects step-sizes at each inner iteration. At every inner step, the algorithm defines the subproblem

w∈Rdw \in \mathbb{R}^d4

where w∈Rdw \in \mathbb{R}^d5 is a sampled mini-batch. The step-size w∈Rdw \in \mathbb{R}^d6 is chosen (approximately) to minimize w∈Rdw \in \mathbb{R}^d7. A one-step Newton update at w∈Rdw \in \mathbb{R}^d8 yields

w∈Rdw \in \mathbb{R}^d9

where \begin{align*} \xi_t'(0) &= -2 v_{t-1}T \nabla2 f_{i_t}(w_{t-1}) v_{t-1}, \ \xi_t''(0) &= 2 v_{t-1}T [\nabla2 f_{i_t}(w_{t-1})]2 v_{t-1} + 2 v_{t-1}T \nabla3 f_{i_t}(w_{t-1})[v_{t-1}] v_{t-1}. \end{align*} To stabilize the step-size, an exponential moving average of the harmonic mean of past reciprocals fi(w)f_i(w)0 maintains an upper bound fi(w)f_i(w)1. The step-size for each iteration is then set as fi(w)f_i(w)2 (Shi et al., 2021).

AI-SARAH thereby adjusts to local curvature and smoothness, estimating effective step-sizes dynamically without access to global fi(w)f_i(w)3 or fi(w)f_i(w)4. This adaptation is not present in base SARAH or other fixed step-size variance-reduced methods.

3. Algorithmic Structure and Stopping Criteria

AI-SARAH operates in epochs (outer loops), each with up to fi(w)f_i(w)5 inner steps, but can terminate earlier if fi(w)f_i(w)6 for a default threshold fi(w)f_i(w)7. Pseudocode details include:

  • Mini-batch size fi(w)f_i(w)8 (default fi(w)f_i(w)9 or ii0),
  • Exponential smoothing parameter ii1 (default ii2),
  • No explicit requirement for full knowledge of global ii3 or ii4,
  • Updates using autodiff for efficient computation of ii5 and ii6.

The algorithm’s stopping rule for the inner loop is identical in spirit to the SARAH⁺ variant, leveraging the observed decay in the squared norm of ii7 to avoid unnecessary iterations and maintain variance reduction efficiency (Nguyen et al., 2017).

4. Theoretical Convergence Properties

Under strong convexity, with ii8 ii9-strongly convex and each P(w)P(w)0 convex (and using P(w)P(w)1-local smoothness), AI-SARAH achieves linearly decaying expected gradient norm per outer epoch:

P(w)P(w)2

where

P(w)P(w)3

and total inner-loop step-sum P(w)P(w)4, with P(w)P(w)5. When local smoothness P(w)P(w)6 is much less than the global P(w)P(w)7, P(w)P(w)8 is larger, leading to potentially faster convergence compared to standard SARAH. In particular, the rate recovers the classical SARAH rate when P(w)P(w)9 but improves upon it where local geometry allows (Shi et al., 2021).

5. Computational Complexity

Each AI-SARAH inner iteration consists of:

  • One mini-batch gradient at fif_i0 and fif_i1 (fif_i2),
  • Two extra directional derivatives via automatic differentiation for fif_i3 and fif_i4 (fif_i5 each).

Thus, per-inner-iteration cost is approximately fif_i6 equivalent stochastic gradients, compared to fif_i7 for base SARAH. The outer full gradient fif_i8 is computed once per epoch (fif_i9). For μ\mu0-accuracy in the strongly convex case, both AI-SARAH and classical SARAH require μ\mu1 gradient-equivalents due to μ\mu2 in practical parameterization. The factor-of-3 cost is offset by the improved adaptation and lack of hyperparameter tuning (Shi et al., 2021).

6. Empirical Performance and Robustness

Extensive experiments on logistic regression tasks, both regularized and unregularized, were conducted on 10 LIBSVM datasets (e.g., ijcnn1, rcv1, news20, gisette, mushrooms). Competitors included fine-tuned SARAH, SARAH+, SVRG, Adam, and SGD-momentum, with up to 5,000 hyperparameter configurations evaluated per dataset for non-AI-SARAH methods.

Key comparative results:

  • AI-SARAH, with default settings (μ\mu3, μ\mu4, μ\mu5), uniformly outperformed or matched the best tuned SARAH, SARAH+, and SVRG algorithms in convergence per effective pass and wall-clock time.
  • AI-SARAH was competitive or faster than Adam and SGD-momentum in reaching lower gradient norms on convex problems.
  • AI-SARAH exhibited robust performance across all problems with no per-dataset tuning.

An example comparison table for final μ\mu6 after 20 passes (regularized setting):

Dataset AI-SARAH Best SARAH Best SVRG Best Adam Best SGD-m
ijcnn1 1.2e–6 8.7e–6 1.0e–5 3.5e–6 5.1e–6
rcv1 2.3e–7 1.5e–6 2.1e–6 4.4e–7 8.9e–7

Values represent mean outcomes over 10 seeds, indicating that AI-SARAH converges more rapidly without requiring hyperparameter tuning (Shi et al., 2021).

7. Implementation Guidelines and Practical Considerations

Recommended implementation details for AI-SARAH include:

  • Default parameters (μ\mu7, μ\mu8, μ\mu9 or μ>0\mu > 00) are robust across tasks.
  • The subproblem for step-size can be solved via a one-step Newton update at μ>0\mu > 01, evaluable by two backward-mode autodiff passes; for quadratic loss, a closed form is available.
  • Exponential smoothing on μ>0\mu > 02 is used to update μ>0\mu > 03; i.e., μ>0\mu > 04.
  • Stopping the inner loop upon μ>0\mu > 05 or after μ>0\mu > 06 steps.
  • No requirement to estimate global μ>0\mu > 07 or μ>0\mu > 08.
  • Efficient integration into deep learning frameworks (e.g., PyTorch, TensorFlow) by treating the step-size μ>0\mu > 09 as a leaf variable in the computational graph and computing required derivatives through backpropagation (Shi et al., 2021).

8. Relation to SARAH, SARAH⁺, and iSARAH

AI-SARAH shares the stochastic recursive gradient paradigm with SARAH, which achieves linear convergence under strong convexity and offers the unique property that the inner-loop estimator μ=0\mu = 00 exhibits linear convergence in expectation within a single loop (Nguyen et al., 2017). SARAH⁺ introduces adaptive inner-loop termination based on μ=0\mu = 01 norm decay, a design inherited by AI-SARAH. The inexact SARAH (iSARAH) generalizes the methodology to expectation minimization problems beyond finite sums by using stochastic, rather than exact, full gradients, and adjusts batch sizes per theoretical analysis (Nguyen et al., 2018).

The defining advancement in AI-SARAH lies in automating step-size adjustment guided by local geometry—rendering it entirely tune-free in practice—while preserving the memory efficiency and convergence guarantees characteristic of the SARAH lineage.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to AI-SARAH.