---
title: Run-Wise Early-Stopping Certificates
url: https://www.emergentmind.com/topics/run-wise-early-stopping-certificates
type: topic
---

# Run-Wise Early-Stopping Certificates

Run-wise early-stopping certificates are per-execution stopping justifications attached to a concrete training, inference, or search trajectory. Rather than terminating because a fixed iteration cap or patience window has been reached, they halt when an explicit criterion certifies that further computation is either statistically unsupported, operationally unnecessary, or provably unable to improve the incumbent outcome. In gradient boosted decision trees, "ScoreStop: Gradient-based early stopping using functional score tests" formalizes this idea by testing whether the current predictor is the population risk minimizer and issuing a certificate when the estimated directional derivative is statistically indistinguishable from zero [2606.02740]. Related work instantiates the same broad idea with e-processes, conformal thresholds, frontier invariants, run-length statistics, and anytime-valid confidence sequences, but the certified object differs substantially across domains.

## 1. Definition and scope

In the most direct statistical formulation, a run-wise early-stopping certificate is a decision that the current state of an iterative procedure is consistent with a formally defined target. For ScoreStop, that target is the population risk minimizer \(f_0\), and the certificate states that the current predictor \(f_t\) is consistent with being the population minimizer when the null hypothesis \(H_0: f_t=f_0\) is not rejected on validation data [2606.02740]. This replaces the standard patience rule, which halts after \(P\) consecutive non-improvements in validation loss, but whose parameter has no calibrated statistical meaning and behaves idiosyncratically across datasets, sample sizes, losses, learning rates, and tree complexities.

The literature shows that the phrase is broader than a single testing framework. Sequential-EDFL for language-model generation realizes a run-wise certificate through an anytime-valid e-process whose threshold crossing \(E_t \ge 1/\delta\) triggers stopping with type I error controlled at level \(\delta\) regardless of stopping time [2510.06478]. Statistical early stopping for reasoning models uses either a renewal-process approximation or a nonparametric maxwise conformal rule, with the latter giving a finite-sample per-run bound on the probability of halting too early on well-posed queries under exchangeability [2602.13935]. Auditable routing certificates in agentic search are deterministic and ledger-verified: halting is sound when every unexpanded leaf is upper-bounded by a frontier key under the same realized randomness [2509.10550].

A common misconception is that all such certificates are anytime-valid or certify correctness. The available work does not support that generalization. ScoreStop’s \(\chi^2_d\) calibration is per iteration rather than uniform over sequential stopping times [2606.02740]. Sequential-EDFL certifies information sufficiency relative to a skeleton baseline, not factual correctness [2510.06478]. Reasoning-model certificates bound false-positive early halts on well-posed queries rather than truth of the final answer [2602.13935]. This suggests that “run-wise early-stopping certificate” is best understood as a family of stopping semantics, not a single guarantee.

## 2. ScoreStop and the functional score-test formulation

ScoreStop is defined for iterative learners that produce predictors \(f_1,f_2,\ldots,f_m\), with a focus on gradient boosted decision trees. The construction applies to smooth losses, nonsmooth or implicit losses such as LambdaRank, and data-dependent losses such as Cox regression, where “gradients” may be subgradients or influence-function-based [2606.02740]. The underlying population problem is specified by letting \(W=(Y,X)\), \(f:X\to\mathbb{R}\) in a Hilbert space \(H\), and unit loss \(L\), with
\[
R_0(f)=E[L\{f(X),W\}],
\qquad
f_0=\arg\min_{f\in H}E[L\{f(X),W\}].
\]

At iteration \(t\), the learner holds the current predictor \(f_t \equiv \hat f_m\), and the next update direction \(h_t \equiv \hat h_m\) is fitted to approximate the gradient function
\[
h^*(x;\hat f_m)=E\!\left[\nabla L\{\hat f_m(X),W\}\mid X=x\right].
\]
ScoreStop then tests the null hypothesis \(H_0:f_0=f_*\), where \(f_*\) is typically \(f_t\). The relevant local quantity is the Gateaux derivative
\[
R'_0(h;f)=\frac{\partial R_0(f+th)}{\partial t}\bigg|_{t=0}
=
E[h(X)\,\nabla L\{f(X),W\}],
\]
with score function
\[
s(W;h,f)=h(X)\,\nabla L\{f(X),W\}.
\]

Under \(H_0\) and fixed \(h\neq 0\), the validation average satisfies
\[
\sqrt{n}R'_n(h;f_*) \overset{d}{\to}
N\!\left(0;E[s(W;h,f_*)^2]\right).
\]
The operational interpretation is that if the directional derivative along the learned update direction is statistically indistinguishable from zero, then the current model is consistent with being the population minimizer and training can stop. Because the construction uses gradients rather than loss values, it avoids the dependence of patience-based rules on noisy or implicitly defined validation losses.

A central design property is scale-invariance. Replacing \(h\to ch\) or \(\nabla L\to c\nabla L\) for any \(c\neq 0\) leaves the test statistic unchanged. The paper states that learning-rate choices and sign conventions therefore do not affect the test, and equivalently
\[
T_n(\hat h_m,\hat f_m)=T_n(\hat f_{m+1}-\hat f_m,\hat f_m),
\]
independent of \(\eta_m\) [2606.02740].

## 3. Test statistics, stopping rule, and theoretical guarantees

The single-direction ScoreStop statistic uses a plug-in variance estimate:
\[
T_n(h,f_*)=
\frac{n\,[s(W;h,f_*)]^2}{[s(W;h,f_*)^2]}.
\]
Under \(H_0:f_*=f_0\), it satisfies
\[
T_n(h,f_*) \overset{d}{\to} \chi^2_1.
\]
For vector-valued predictors \(f:X\to\mathbb{R}^K\), the score generalizes to
\[
s(W;h,f)=\sum_{k=1}^{K} h_k(X)\,\nabla_k L\{f(X),W\}.
\]
For multiple directions \(h_1,\ldots,h_d\), ScoreStop forms a quadratic statistic
\[
T_n(\mathbf{h},f_*)=
n\,[\hat{\boldsymbol\varphi}(\mathbf{h},f_*)]^\top
\hat{\Sigma}_n(\mathbf{h},f_*)^{-1}
[\hat{\boldsymbol\varphi}(\mathbf{h},f_*)],
\]
with asymptotic null law \(\chi^2_d\) [2606.02740].

The per-iteration stopping rule is operationally simple. At iteration \(t\), the base learner \(\hat h_t\) is fit on the training set. On the held-out validation set, one computes \(T_n(\hat h_t,\hat f_t)\). If \(T_n \le c_\alpha\), training stops and returns \(\hat f_t\); otherwise the update \(\hat f_{t+1}=\hat f_t-\eta_t\hat h_t\) is accepted and training continues. The threshold is calibrated as
\[
c_\alpha = F^{-1}_{\chi^2_d}(1-\alpha),
\]
with \(d=1\) for the forward and backward variants and \(d=2\) for the stabilized two-direction variant. For \(d=1\), the paper notes the convenient parameterization \(c_\alpha=z^2\) with two-sided \(z=\Phi^{-1}(1-\alpha/2)\) [2606.02740].

Three algorithmic variants are distinguished. Forward ScoreStop (FWD-SS) tests \(T_n(\hat h_m,\hat f_m)\). Backward ScoreStop (BWD-SS) tests \(T_n(\hat h_{m-1},\hat f_m)\) for \(m\ge 2\). Stabilized ScoreStop (STAB-SS) uses the two-direction statistic \(T_n((\hat h_m,\hat h_{m-1}),\hat f_m)\). The computational overhead is one validation pass per iteration to compute \([s]\) and \([s^2]\), or \([\varphi]\) and \([\varphi^2]\) in the data-dependent case; no training refits are needed for the test itself.

The asymptotic theory is organized around assumptions on influence-function centrality, direction linearity and a Riesz representer, finite variance, regular asymptotic linearity, and moment consistency. Under these conditions, the null distribution is \(\chi^2_1\) or \(\chi^2_d\) as appropriate. The power result states that if \(\langle h,h^*(f_*)\rangle \neq 0\) and the variance is finite, then \(P\{T_n(h,f_*)>c_\alpha\}\to 1\) for any fixed \(c_\alpha\), so the test continues when the current predictor is not the minimizer and the direction is aligned with improvement [2606.02740]. By Cauchy–Schwarz, \(|\langle h,h^*(f_*)\rangle|\) is maximized by \(h\propto h^*(f_*)\), which the paper identifies with the base learner direction in boosting; this is the sense in which FWD-SS tests in a near-optimal direction.

The main theoretical caveat is sequential. The paper explicitly recommends choosing \(c_\alpha\) as a performance-tuned regularization parameter, interpretable via \(\chi^2_d\), rather than treating it as a strict test level over the entire stopping time. There are no anytime-valid guarantees for the online use of one test per iteration [2606.02740].

## 4. Implicit losses, data-dependent losses, and influence-function generalization

A notable feature of ScoreStop is that it does not depend on explicit evaluation of a scalar loss value. This is crucial for LambdaRank, where the training signal is an implicit document-wise gradient, and for Cox proportional hazards, where the loss depends on nuisance quantities estimated from the data [2606.02740].

For data-dependent losses, the score \(s\) is replaced by a mean-zero influence function \(\varphi(W;h,f_*,P_0)\) satisfying
\[
\sqrt{n}R'_n(h;f_*) \overset{d}{\to}
N\!\left(0;E[\varphi(W;h,f_*,P_0)^2]\right),
\]
and the corresponding one-step debiased statistic is
\[
T_n(h,f_*)=
\frac{n\,[\varphi(W;h,f_*,\hat P_n)]^2}
{[\varphi(W;h,f_*,\hat P_n)^2]}.
\]
This reduces to the basic statistic when the score is known and there is no data dependence.

In LambdaRank, a query \(W=(N,Y,X)\) with scores \(\underline f_N(X)=(f(X_1),\ldots,f(X_N))\) uses
\[
s(W;h,f)=\sum_{k=1}^{N} h(X_k)\,\lambda_k\{\underline f_N(X),W\},
\]
where the pairwise gradient is
\[
\lambda_{ij}\{\underline f_N(X),W\}
=
\frac{-\sigma}{1+\exp\{\sigma(f(X_i)-f(X_j))\}}
\,\big|\Delta \mathrm{NDCG}_{ij}\{\underline f_N(X),W\}\big|,
\]
and
\[
\lambda_k=
\sum_{j:(k,j)\in\mathcal{I}(W)} \lambda_{kj}
-
\sum_{j:(j,k)\in\mathcal{I}(W)} \lambda_{jk}.
\]
The paper states that Poincaré symmetry ensures an implicit loss exists locally, but ScoreStop uses \(\lambda_k\) directly and thereby avoids explicit evaluation of nonsmooth NDCG.

For Cox proportional hazards, the unit loss and gradient are
\[
L\{f(X),W;\Lambda_0\}=\exp\{f(X)\}\Lambda_0(T)-Df(X),
\qquad
\nabla L\{f(X),W;\Lambda_0\}=\exp\{f(X)\}\Lambda_0(T)-D.
\]
On the validation set, the Breslow baseline-hazard estimator is
\[
\hat{\Lambda}_{n,f}(t)=\left[\frac{N(t)}{S_n^{(0)}(T;f)}\right],
\qquad
S_n^{(0)}(t;f)=[Y(t)\exp\{f(X)\}],
\]
and the efficient influence function under \(H_0\) is
\[
\varphi(W;h,f_*,P_0)=
s(W;h,f_*,\Lambda_0)
+
D\,\bar h(T)
-
\exp\{f_*(X)\}\int_0^T \bar h(t)\,\mathrm d\Lambda_0(t),
\]
with
\[
\bar h(t)=
\frac{E[h(X)\exp\{f_*(X)\}Y(t)]}
{E[\exp\{f_*(X)\}Y(t)]}.
\]
Plugging this \(\varphi\) into the one-step statistic yields a nuisance-adjusted ScoreStop test. The broader implication is that run-wise certificates need not be restricted to explicit, smooth, scalar losses; they can be built from derivative representations that remain valid under ranking and survival objectives.

## 5. Empirical behavior, limitations, and operational guidance

In synthetic experiments across regression, classification, ranking, and survival tasks with known \(f_0\), ScoreStop in its forward, backward, and stabilized forms is competitive with patience-based stopping at small \(P\), and is notably stronger in learning-to-rank, where loss-based stopping tracks a nonsmooth metric not directly tied to training gradients [2606.02740]. STAB-SS often outperforms FWD-SS and BWD-SS on regression and survival, which the paper attributes to reduced reliance on a single base learner. Near the test-loss minimum iteration, \(T_n\) aligns well with \(\chi^2_d\) in QQ plots, indicating good finite-sample calibration of the null distribution. At moderate learning rates such as \(\eta=0.05\), ScoreStop and patience are comparable; very large \(\eta\) degrades all methods. Across 13 real datasets spanning regression, classification, multiclass, Poisson, quantile, ranking, and survival, ScoreStop with \(z=0.05\) is broadly competitive with patience baselines, especially when validation losses are noisy or implicit.

The same experiments motivate practical threshold choices. The paper recommends \(z\in\{0.05,0.1\}\), corresponding to \(c_\alpha=z^2\) for the single-direction case. It also recommends communicating significance through \(z\), interpreted in standard deviation units, while treating \(c_\alpha\) as a regularization parameter calibrated to \(\chi^2_d\) rather than as a strict familywise sequential error level.

Several failure modes are explicit. Directional power depends on alignment: if \(\langle h,h^*(f_*)\rangle = 0\), the test has no power against that alternative. STAB-SS mitigates this by testing multiple directions, but successive update directions can be nearly collinear, leading to rank-deficient covariance, especially at very small learning rates. The calibration presumes a sufficiently large, independent validation set, so small validation samples or dependence between training and validation can harm calibration. For Cox and other data-dependent losses, reliable nuisance estimation is required on the validation set, and the one-step debiased construction depends on the asymptotic linearity conditions. For LambdaRank, the paper advises handling ties and truncation carefully [2606.02740].

The recommended practice is correspondingly conservative: use an independent validation set; prefer FWD-SS or STAB-SS with \(z\approx 0.05\)–\(0.1\); compute nuisance estimates on validation data for data-dependent losses; monitor near-collinearity if using STAB-SS; and interpret the threshold as a scale-calibrated regularizer rather than a strict test-level \(\alpha\). Within those constraints, the run-wise certificate supplied by ScoreStop is a principled substitute for patience counts because it is anchored to an interpretable null hypothesis and a calibrated test statistic.

## 6. Broader landscape and methodological contrasts

Subsequent work shows that run-wise early-stopping certificates are now used across generation, optimization, certification, routing, and quantum learning, but the stopping evidence and guarantee class vary [2510.06478] [2602.13935] [2605.29860] [2606.27694] [2512.13123] [2509.10550] [2602.10648].

| Setting | Certificate mechanism | Scope of guarantee |
|---|---|---|
| Gradient boosting | Functional score test \(T_n \le c_\alpha\) | Per-iteration \(\chi^2_d\) calibration |
| LLM generation | E-process threshold \(M_t \ge 1/\delta\) | Anytime-valid \(\delta\)-level control |
| Reasoning models | Maxwise conformal or renewal thresholds | False-positive halt control on well-posed queries |
| Projected SGD | Confidence sequence \(U_t^{\mathrm{obs}}(\alpha)\le \varepsilon\) | \(\varepsilon\)-optimality with probability at least \(1-\alpha\) |
| Agentic routing | Frontier key bounded by incumbent | Deterministic replayable halting soundness |

Sequential-EDFL is the clearest anytime-valid analogue. It constructs a self-normalized empirical-Bernstein e-process over clipped token-wise information lift relative to a fixed skeleton baseline, uses mixtures over \(\lambda\), supports adaptive resets, and stops at the first time the e-process reaches \(1/\delta\). On six benchmarks it reduces generation by \(22\)–\(28\%\) versus sequential baselines with \(12\%\) computational overhead, but the paper emphasizes that the certificate controls information sufficiency, not factual correctness; even with a correctness gate, \(10.9\%\) of stopped sequences remain incorrect [2510.06478]. This is a markedly different semantic object from ScoreStop’s consistency-with-\(f_0\) certificate.

For reasoning models, uncertainty-aware stopping monitors arrivals of uncertainty keywords. The nonparametric maxwise conformal rule gives the finite-sample guarantee
\[
\mathbb{P}\!\left(\exists j:\,u_{n+1}(L_j)>\tau^\star\right)\le \alpha,
\]
which the paper interprets as a run-wise early-stopping certificate with false-positive risk at most \(\alpha\) on well-posed queries under exchangeability [2602.13935]. Empirically, the keyword-based rules achieve up to \(\sim 70\%\) token savings on ill-posed math tasks while maintaining low false-positive rates.

In reinforcement learning, ESPO derives a trajectory-wise certificate from logits already computed during sampling. It forms a normalized and smoothed surrogate regret \(z_t\) and halts when
\[
z_t > \beta \cdot \max\!\big(V_\phi(s_t),\varepsilon_v\big).
\]
Truncated trajectories are mapped to absorbing failure states with a terminal penalty, concentrating negative temporal-difference errors near the detected failure step. On DeepSeek-R1-Distill-Qwen-7B for mathematical reasoning, ESPO exceeds PPO on AIME 2024, AMC 2023, and MATH-500 while saving more than \(20\%\) rollout tokens cumulatively [2605.29860].

Run-wise certificates also appear in fixed-input certification. In randomized smoothing, anytime-valid per-input stopping is obtained by inverting mixture e-processes for Bernoulli success probabilities and mapping the resulting lower confidence sequence to a certified radius \(r_t=\sigma\Phi^{-1}(\max\{\underline p_t,0.5\})\). With meta-learned priors for the mixture, the method reports a \(20\)-fold reduction in sample complexity relative to traditional fixed-sample randomized smoothing while maintaining rigorous statistical guarantees [2606.27694]. For projected SGD with general convex objectives, anytime-valid confidence sequences yield an online stopping rule \(\tau(\varepsilon,\alpha)=\inf\{t\ge 1:U_t^{\mathrm{obs}}(\alpha)\le \varepsilon\}\) such that the returned weighted-average iterate is \(\varepsilon\)-optimal with probability at least \(1-\alpha\), and the stopping time is almost surely finite under standard stochastic approximation stepsizes [2512.13123].

Other variants are non-statistical but still run-wise. Ledger-verified routing certificates for perturb-and-MAP best-first search halt when the maximum frontier key is no larger than the incumbent realized value, with validator replay checking the frontier invariant and halting soundness under the same realized exponential race [2509.10550]. In quantum learning under one-bit feedback, the certificate is a run-length event: halt after \(M_H\) consecutive successes under frozen control. Its usefulness depends sharply on noise, with the paper identifying the feasibility threshold \(qM_H \gtrsim 1\), beyond which the expected halting time grows exponentially [2602.10648].

This broader landscape makes two distinctions especially important. First, some certificates are time-uniform and optional-stopping safe, while others are only calibrated at a fixed or per-iteration horizon. Second, the object being certified may be statistical optimality, information sufficiency, non-premature halting, frontier dominance, or robustness radius. Run-wise early-stopping certificates are therefore best treated as a methodological class defined by auditable stopping semantics, not by a single proof technique or guarantee format.

Source: https://www.emergentmind.com/topics/run-wise-early-stopping-certificates