---
title: 'GRADSTOP: Gradient-based Stopping in ML'
url: https://www.emergentmind.com/topics/gradstop
type: topic
---

# GRADSTOP: Gradient-based Stopping in ML

GRADSTOP is an overloaded label in contemporary machine learning and optimization. In the literature considered here, it refers not to a single canonical algorithm but to several gradient-based stopping, stabilization, and steering mechanisms. The common pattern is that gradient information is used not only to update model parameters, but also to determine when optimization should halt, when acceleration should restart, when a trajectory has become stable, or when a predictor should be regularized against overfitting or harmful behavior [2601.01853][2508.19028][2412.08501].

## 1. Terminological scope and recurring structure

Across recent work, GRADSTOP appears in at least three direct senses: as asymptotic “stopping” behavior of adaptive stochastic optimizers; as a named validation-free early stopping rule based on posterior sampling; and as a named label-free early stopping rule for deep unsupervised outlier detection. Closely related work uses gradient norms or directional derivatives as stopping or restart signals in accelerated optimization, saddle-point methods, stochastic approximation, and boosting [2401.07672][2307.09921][2606.02740][2004.00475].

| Usage | Core mechanism | Representative paper |
|---|---|---|
| Adaptive-optimizer “stopping” | Vanishing gradients and shrinking updates in AdaGrad-Norm and RMSProp | [2601.01853] |
| Validation-free early stopping | Approximate posterior sampling from gradient information | [2508.19028] |
| Label-free UOD early stopping | Gradient cohesion and divergence during training | [2412.08501] |
| Restart/stopping in optimization | Gradient-triggered restart or norm-based termination | [2401.07672], [2307.09921], [2606.02740], [2004.00475] |

This multiplicity matters. In one line of work, “stop” means that an optimizer asymptotically ceases to make meaningful updates because $\|\nabla g(\theta_n)\| \to 0$. In another, it means a finite-time model-selection rule that selects an iterate before overfitting. In a third, it means a gradient-triggered restart or termination criterion inside an optimization algorithm. A plausible implication is that GRADSTOP is best treated as a family of gradient-governed stopping ideas rather than as a single method name.

## 2. Asymptotic stopping in adaptive gradient methods

One technically precise meaning of GRADSTOP arises in the asymptotic analysis of adaptive stochastic optimization. The underlying problem is
\[
\min_{\theta \in \mathbb{R}^d} g(\theta),
\]
with $g$ continuously differentiable, non-negative, and $L$-smooth,
\[
\|\nabla g(\theta) - \nabla g(\theta')\| \le L \|\theta - \theta'\|\quad\forall \theta,\theta',
\]
together with coercivity,
\[
\lim_{\|\theta\|\to\infty} g(\theta)=+\infty,
\]
a non-flatness-at-infinity condition,
\[
\liminf_{\|\theta\|\to\infty}\|\nabla g(\theta)\|>\tilde{\delta},
\]
and stochastic gradients satisfying unbiasedness and an affine variance bound,
\[
\mathbb{E}[\nabla g(\theta_n,\xi_n)\mid\mathscr{F}_{n-1}] = \nabla g(\theta_n),
\]
\[
\mathbb{E}\big[\|\nabla g(\theta_n,\xi_n)\|^2\mid\mathscr{F}_{n-1}\big]
\le \sigma_0 \|\nabla g(\theta_n)\|^2 + \sigma_1.
\]
Within this setting, AdaGrad-Norm is written as
\[
\begin{aligned}
S_n &= S_{n-1} + \|\nabla g(\theta_n,\xi_n)\|^2,\\
\theta_{n+1} &= \theta_n - \frac{\alpha_0}{\sqrt{S_n}}\nabla g(\theta_n,\xi_n),
\end{aligned}
\]
with scalar stepsize $\alpha_n=\alpha_0/\sqrt{S_n}$, while RMSProp is studied in the coordinate-wise form
\[
\begin{aligned}
v_n^{(i)} &= \beta_n v_{n-1}^{(i)} + (1-\beta_n)\big(\nabla_i g(\theta_n,\xi_n)\big)^2,\\
\theta_{n+1}^{(i)} &= \theta_n^{(i)} - \frac{\alpha_n^{(0)}}{\sqrt{v_n^{(i)}+\epsilon}}\nabla_i g(\theta_n,\xi_n),
\end{aligned}
\]
with the analyzed schedule
\[
\alpha_n^{(0)} = \frac{1}{\sqrt{n}},\qquad \beta_n = 1-\frac{1}{n}\quad (n\ge 2).
\]
The central technical device is a stopping-time partition built from a Lyapunov function. For AdaGrad-Norm, the Lyapunov term is
\[
\zeta(n) := \frac{\|\nabla g(\theta_n)\|^2}{\sqrt{S_{n-1}}},\qquad
\hat g(\theta_n) := g(\theta_n) + \frac{\sigma_0\alpha_0}{2}\zeta(n),
\]
and the trajectory is partitioned by thresholds $\Delta_\tau$ and $2\Delta_\tau$ into intervals that localize excursions of $\hat g$. This is used to prove
\[
\mathbb{E}\Big[\sup_{n\ge 1} g(\theta_n)\Big] < \infty
\]
and, under coercivity,
\[
\sup_{n\ge 1}\|\theta_n\| < \infty \quad\text{a.s.}
\]
Building on these stability results, the analysis establishes
\[
\lim_{n\to\infty}\|\nabla g(\theta_n)\| = 0 \quad \text{a.s.},
\qquad
\lim_{n\to\infty}\mathbb{E}\|\nabla g(\theta_n)\|^2 = 0,
\]
for AdaGrad-Norm, and the same type of almost sure and mean-square convergence for RMSProp under the stated hyperparameter choices. The practical interpretation is explicit: “in practice, this means that step magnitudes $\|\theta_{n+1}-\theta_n\|$ shrink to zero, so the algorithm effectively stops making large updates” [2601.01853].

A precursor AdaGrad analysis introduced the same stopping-time viewpoint and also derived the near-optimal non-asymptotic bound
\[
\frac{1}{T}\sum_{n=1}^T \mathbb{E}\|\nabla g(\theta_n)\|^2
\le \mathcal{O}\!\left(\frac{\ln T}{\sqrt{T}}\right),
\]
together with almost sure and mean-square convergence under related assumptions. This places GRADSTOP, in the adaptive-optimizer sense, at the intersection of Lyapunov analysis, stochastic approximation, and stopping-time control of loss spikes [2409.05023].

## 3. Validation-free early stopping via posterior sampling

A second, explicitly named, GRADSTOP is a stochastic early stopping rule for gradient descent that uses only gradient information produced “for free” during training. The setup is a supervised learning problem with empirical loss
\[
L(\theta)=\sum_{i=1}^n l_i(\theta),
\]
linked to a Bayesian model by
\[
l_i(\theta) = - \log p(z_i \mid \theta) - \frac{1}{n}\log p(\theta),
\qquad
p(\theta\mid D)\propto e^{-L(\theta)}.
\]
The method interprets early stopping as drawing a single approximate sample from the posterior, restricted to the optimization path $\{\theta_1,\dots,\theta_T\}$.

The key statistic is a gradient-based approximation of posterior credibility. Let
\[
g_i(\theta)=\nabla_\theta l_i(\theta),\qquad
\overline g(\theta)=\frac{1}{n}\sum_{i=1}^n g_i(\theta),
\]
and define the empirical gradient covariance
\[
\Sigma_G(\theta)=\frac{1}{n}\sum_{i=1}^n
\bigl(g_i(\theta)-\overline g(\theta)\bigr)
\bigl(g_i(\theta)-\overline g(\theta)\bigr)^\top.
\]
Under a local quadratic approximation,
\[
L(\theta)\approx \frac{1}{2}(\theta-\theta^*)^\top H(\theta-\theta^*),
\qquad
p(\theta\mid D)\approx \mathcal N(\theta^*,H^{-1}),
\]
and the Mahalanobis radius is approximated by
\[
z(\theta)=n\,\overline g(\theta)^\top \Sigma_G(\theta^*)^{-1}\overline g(\theta).
\]
The resulting credibility approximation is
\[
\hat s(\theta\mid D)
=
1-F_{\chi_d^2}\bigl(z(\theta)\bigr),
\]
where $F_{\chi_d^2}$ is the $\chi_d^2$ cdf. GRADSTOP draws $u\sim \mathrm{Uniform}(0,1)$ and returns the iterate whose $\hat s_t$ is closest to $u$; in the deterministic variant used in experiments, a fixed threshold is used and training is stopped when $\hat s_t$ exceeds it.

Algorithmically, each iteration requires the per-example gradient matrix $\mathbf G_t\in\mathbb R^{n\times d}$, Oracle Approximating Shrinkage for covariance regularization,
\[
\hat\Sigma_G = (1-\epsilon)\Sigma_G + \epsilon \frac{\mathrm{tr}(\Sigma_G)}{d} I_d,
\]
and then
\[
z = n\,\overline g^\top \hat\Sigma_G^{-1}\overline g,
\qquad
\hat s(\mathbf G)=1-F_{\chi_d^2}(z).
\]
The method is validation-free, uses all data for training, and was evaluated on small-data tabular tasks and transfer-learning settings. Reported test-loss examples include: Heart disease with logistic regression, where no stopping gives $1.468$, validation stopping gives $1.138$, and GRADSTOP gives $1.055$; Diabetes with an MLP, where no stopping gives $1.988$, validation gives $0.907$, and GRADSTOP gives $0.785$; and Hepatitis with an MLP, where no stopping gives $2.758$, validation gives $0.657$, and GRADSTOP gives $0.625$ [2508.19028].

The same framework also yields uncertainty estimates. For an affine scalar functional $f(\theta)$, the posterior standard deviation is approximated by
\[
\sigma[f(\theta)]
\approx
\sqrt{f'(\theta^*)^\top \Sigma_G(\theta^*)^{-1} f'(\theta^*)/n}.
\]
This places GRADSTOP in a Bayesian-asymptotic lineage rather than in the Lyapunov or stationary-point lineage of optimization-theoretic uses of the term.

## 4. Label-free early stopping in unsupervised outlier detection

A third named GradStop is a label-free early stopping algorithm for deep unsupervised outlier detection trained on contaminated data. The data are unlabeled,
\[
\mathcal D = \mathcal D_{in}\cup \mathcal D_{out},
\qquad
|\mathcal D_{out}|\ll |\mathcal D_{in}|,
\]
and a model $M$ outputs anomaly scores
\[
v_j = f_M(x_j).
\]
The target outlier-detection objective is the AUC-like ordering probability
\[
P(v^-<v^+)=P\big(f_M(x_{in})<f_M(x_{out})\big),
\]
whereas training minimizes an unsupervised loss
\[
\mathcal L(M;B)=\frac{1}{|B|}\sum_{x\in B}\mathcal J_M(x).
\]
The central difficulty is the “misalignment between the model's direct optimization goal and the final performance goal of Outlier Detection (OD) task,” together with the “inlier priority phenomenon,” under which models fit inliers faster than outliers.

GradStop operationalizes this training-dynamics view with a sampling step and two gradient statistics. Given per-sample gradients
\[
G=\{g_i : g_i=\nabla_{\Theta_t}\mathcal J_M(x_i;\Theta_t)\},
\]
it sorts them by norm and forms
\[
G_t^{\text{last}}=\{g_i:\text{$k$ smallest }\|g_i\|\},
\qquad
G_t^{\text{top}}=\{g_i:\text{$k$ largest }\|g_i\|\}.
\]
The “last” set is intended to be enriched in inliers, and the “top” set in outliers. The cohesion metric is
\[
\mathbf C(G)=\frac{\left\|\sum_{i=1}^k g_i\right\|}{\sum_{i=1}^k \|g_i\|},
\]
and the divergence metric is the angle
\[
\mathbf D(G^1,G^2)=
\angle\left(\sum_{i=1}^k g_i^1,\sum_{i=1}^k g_i^2\right).
\]
The monitored quantity is the cohesion difference
\[
C^{diff}[t]
=
\mathbf C\big(G_t^{\text{last}}\big)-\mathbf C\big(G_t^{\text{top}}\big).
\]
The algorithm stops when this difference reaches a local maximum and then enters a downtrend, interpreting that point as the moment at which inlier-priority learning is strongest and the model is about to start overfitting outliers.

The theoretical analysis introduces class-level gradients $\nabla_i=\|\nabla f^i(\omega_t)\|$, $\nabla_o=\|\nabla f^o(\omega_t)\|$, the ratio
\[
r_t=\frac{\nabla_i}{\nabla_o},
\]
the angle
\[
\theta_t=\angle\big(\nabla f^i(\omega_t),\nabla f^o(\omega_t)\big),
\]
and the class-size ratio
\[
R=\frac{|C_i|}{|C_o|}.
\]
A sufficient condition for strengthening inlier priority is
\[
r_t >
\cos\theta_t\,R
+
\sqrt{\cos^2\theta_t\,R^2 + 2R + 1}.
\]
Empirically, the method was evaluated on 4 deep UOD algorithms and 47 real-world datasets. For the AE family, VanillaAE has mean AUC $0.758 \pm 0.004$ and GradAE has mean AUC $0.775 \pm 0.003$, while DeepSVDD improves from AUC $0.502$ to $0.648$ under GradStop. The paper states that “AE enhanced by GradStop achieves better performance than itself, other SOTA UOD methods, and even ensemble AEs” [2412.08501].

This use of GRADSTOP is therefore neither asymptotic convergence control nor Bayesian posterior sampling. It is a label-free model-selection rule derived from gradient geometry during representation learning.

## 5. Gradient-triggered stopping and restart in adjacent optimization frameworks

Several adjacent lines of work use the same logic—small or misaligned gradients as a principled stop or restart signal—without always adopting the exact GRADSTOP name.

In accelerated composite optimization, gradient restart is triggered by
\[
(x_{k+1}-x_k,\; y_k-x_{k+1}) > 0 \quad\Rightarrow\quad \text{restart at step }k.
\]
For the smooth case, this becomes
\[
(\nabla f(y_k),\,x_{k+1}-x_k)>0.
\]
The method resets momentum when the inertial direction becomes detrimental. The discrete-time analysis proves global $R$-linear convergence for strongly convex composite problems, and the continuous-time analysis shows that the gradient-restarted ODE has global linear convergence for quadratic convex objectives, whereas the non-restarted trajectory does not enjoy this property [2401.07672].

For saddle-point problems
\[
\min_x \max_y f(x,y),
\]
under a two-sided Polyak–Łojasiewicz condition, a nested gradient scheme uses an inner stopping rule
\[
\|\nabla_y f(x_k,y_m)\| \le \mu_2\gamma
\]
and an outer stopping rule
\[
\|\nabla_x f(x_k,y_m)\| \le L_{12}\gamma\sqrt{6}.
\]
The inner rule guarantees
\[
\|y_m-y^*(x_k)\|\le \gamma,
\]
while the outer rule implies
\[
g(x_k)-g(x^*) \le \frac{7L_{12}^2\gamma^2}{\mu_1},
\qquad
\|\hat x-x^*\| \le \frac{L_{12}\gamma\sqrt{14}}{\mu_1},
\]
for the outer objective $g(x)=\max_y f(x,y)$ [2307.09921].

In gradient boosted decision trees, ScoreStop casts early stopping as a score test for the null hypothesis that the current predictor is already the population risk minimizer. For a direction $h$ and score contributions $s(W;h,f_*)$, the statistic is
\[
T_n(h,f_*)
=
\frac{n\,[s(W;h,f_*)]^2}{[s(W;h,f_*)^2]},
\]
or, in the influence-function setting,
\[
T_n(h,f_*)
=
\frac{n\,[\hat\varphi(W;h,f_*,\hat P_n)]^2}
{[\hat\varphi(W;h,f_*,\hat P_n)^2]}.
\]
Under the null, it has an asymptotic $\chi_1^2$ law, and the rule is to stop when
\[
T_n(\hat h_m,\hat f_m)\le c_\alpha.
\]
Because the statistic is scale-invariant in the update direction, the same construction applies to explicit losses, LambdaRank, and Cox regression [2606.02740].

For SGD on Bottou–Curtis–Nocedal functions, two further stopping rules are developed after proving
\[
\|\dot F(\beta_k)\|_2 \to 0 \quad \text{a.s.}
\]
The first stops when the norm of the mean of fresh stochastic gradients falls below a threshold,
\[
\left\|
\frac{1}{N_j}\sum_{i=1}^{N_j}\dot f(\beta_{T_j},Z_{ij})
\right\|_2 \le \epsilon,
\]
and the second stops when the empirical fraction of “small” stochastic gradients exceeds a vote threshold,
\[
\frac{1}{N_j}\sum_{i=1}^{N_j}
\mathbf 1\{\|\dot f(\beta_{T_j},Z_{ij})\|_2 \le \epsilon\}
\ge \delta_j.
\]
Both are shown to trigger in finite time with probability $1$ under the stated assumptions [2004.00475].

Taken together, these works show that gradient-based stopping can mean at least four distinct things: restart of acceleration, certification of approximate stationarity, early termination of iterative fitting, and test-based model selection.

## 6. Informal and adjacent usages

The terminological spread of GRADSTOP extends beyond stopping in the narrow sense. In the provided material, “Gradient Reversal Against Discrimination” is described as “sometimes referred to informally as ‘GRADSTOP’.” That method is not an early stopping rule; it is an adversarial-training technique for fairness in which a trunk representation feeds both a target branch and an attribute branch, and gradients from the attribute branch are reversed before reaching the trunk:
\[
R_\lambda(h)=h,\qquad
\frac{\partial R_\lambda(h)}{\partial h} = -\lambda I.
\]
Equivalently, the effective trunk gradient is
\[
\nabla_{\theta_{\text{trunk}}}\ell_{\text{eff}}
=
\nabla_{\theta_{\text{trunk}}}\ell_t(y)
-
\lambda \nabla_{\theta_{\text{trunk}}}\ell_p(a_p).
\]
The goal is not to stop optimization, but to suppress protected-attribute information in the learned representation while retaining predictive information for the main task [1807.00392].

A related but different extension appears in safety guardrails for large language models. “Gradient-Controlled Decoding” is described as belonging to the same family as “GRADSTOP-style ‘gradient-guided stopping/steering’ methods.” There, gradients with respect to two anchor tokens—an acceptance anchor such as “Sure” and a refusal anchor such as “Sorry”—are used to detect unsafe prompts, after which refusal tokens are preset-injected before decoding resumes. The mitigation step guarantees first-token safety because the first emitted tokens are not sampled but fixed [2604.05179].

These adjacent usages reinforce a broader interpretation. The stable core is not the exact algorithmic form, but the idea that gradients can be used as a control signal for halting, redirecting, or structurally constraining learning or decoding. In that sense, GRADSTOP denotes a methodological motif whose exact realization depends on the surrounding problem: non-convex stochastic optimization, Bayesian early stopping, unsupervised anomaly detection, fairness, or safety steering.

Source: https://www.emergentmind.com/topics/gradstop