---
title: Uniform Argument Stability
url: https://www.emergentmind.com/topics/uniform-argument-stability
type: topic
---

# Uniform Argument Stability

to=arxiv_search 娱乐彩票注册  彩神争霸怎么json
{"query":"\"uniform argument stability\" OR \"argument stability\" learning algorithm generalization", "max_results": 10, "sort_by": "relevance"}
to=arxiv_search  大发快三豹子json
{"query":"ti:\"Algorithmic stability and hypothesis complexity\" OR abs:\"argument stability\"", "max_results": 10, "sort_by": "relevance"}
to=arxiv_search  彩神争霸是json
{"query":"all:\"uniform stability\" Banach martingale generalization stability of learning algorithms", "max_results": 10, "sort_by": "relevance"}
to=arxiv_search  天天中彩票足彩 娱乐开号json
{"query":"1702.08712", "max_results": 5, "sort_by": "relevance"}
Uniform argument stability is a notion of algorithmic stability that controls the perturbation of the learned hypothesis in its ambient normed function space rather than only the perturbation of the loss. In the formulation introduced in "Algorithmic stability and hypothesis complexity" [1702.08712], a learning algorithm \(\mathcal A:S\mapsto h_S\) taking an i.i.d. sample \(S\in \mathcal Z^n\) into a hypothesis \(h_S\) in a separable Banach space \((\mathcal B,\|\cdot\|)\) is \(\alpha(n)\)-uniformly argument stable when replacing any single training example changes the output by at most \(\alpha(n)\) almost surely. This notion is strictly stronger than classical uniform stability under standard Lipschitz assumptions, and it yields high-probability generalization bounds through Banach-space martingale concentration, Rademacher complexity of a data-dependent hypothesis class, and localization arguments [1702.08712].

## 1. Formal definition and basic comparison

Let \(S=(Z_1,\dots,Z_n)\in\mathcal Z^n\) be an i.i.d. training sample from \(P\) on \(\mathcal Z=\mathcal X\times\mathcal Y\), let \(S^i=(Z_1,\dots,Z_{i-1},Z'_i,Z_{i+1},\dots,Z_n)\) be the sample in which the \(i\)-th point is replaced by an independent copy \(Z'_i\sim P\), and let \(\mathcal A:S\mapsto h_S\in H\subseteq(\mathcal B,\|\cdot\|)\). The algorithm is called \(\alpha(n)\)-uniformly argument stable if, for every \(i\in\{1,\dots,n\}\),
\[
\|h_S-h_{S^i}\|\le \alpha(n)
\quad\text{almost surely.}
\]
This definition measures the sensitivity of the output hypothesis itself to a one-sample perturbation [1702.08712].

The standard comparator is classical uniform stability in the sense of Bousquet–Elisseeff, which bounds the loss difference
\[
\sup_{Z\in\mathcal Z}\bigl|\ell(h_S,Z)-\ell(h_{S^i},Z)\bigr|\le \beta(n).
\]
Under the assumptions that \(\ell\) is \(L\)-Lipschitz in the prediction \(\langle h,x\rangle\) and \(\|X\|_*\le B\) almost surely, argument stability implies loss stability through
\[
\|h_S-h_{S^i}\|\le \alpha(n)
\;\Longrightarrow\;
\bigl|\ell(h_S,Z)-\ell(h_{S^i},Z)\bigr|\le L\,B\,\alpha(n).
\]
The reverse implication does not hold in general [1702.08712]. A common misconception is therefore to treat the two notions as interchangeable. The stated implication shows that uniform argument stability is stronger: it controls the entire hypothesis perturbation in norm, not merely its effect on a particular bounded loss.

## 2. Banach-space framework and concentration around the mean

The generalization theory built around uniform argument stability in [1702.08712] is formulated in a geometric-probabilistic framework. The Banach space \((\mathcal B,\|\cdot\|)\) is assumed to be \((2,D)\)-smooth, meaning that for all \(h,h'\in\mathcal B\),
\[
\|h+h'\|^2+\|h-h'\|^2
\le
2\|h\|^2+2D^2\|h'\|^2,
\]
with Hilbert spaces as the special case \(D=1\). The dual space is assumed to be of type \(p\ge 1\) with constant \(C_p\), so that for any \(x_1,\dots,x_n\in\mathcal X\),
\[
\mathbb E_\sigma\Big\|\sum_{i=1}^n \sigma_i x_i\Big\|_*
\le
C_p\Bigl(\sum_{i=1}^n \|x_i\|_*^p\Bigr)^{1/p}.
\]
In addition, one assumes \(\|X\|_*\le B\) almost surely, and that the loss \(\ell(h,z)\) is bounded by \(M\) and \(L\)-Lipschitz in \(\langle h,x\rangle\) [1702.08712].

A central step is a martingale concentration argument due to Pinelis. If \(D_1,\dots,D_n\) is a martingale-difference sequence in a \((2,D)\)-smooth Banach space and \(\sum_{t=1}^n \|D_t\|_\infty^2\le c^2\), then for every \(\epsilon>0\),
\[
\Pr\Bigl\{\Bigl\|\sum_{t=1}^n D_t\Bigr\|\ge c\,\epsilon\Bigr\}
\le
2\exp\!\bigl(-\tfrac{\epsilon^2}{2D^2}\bigr).
\]
Applied to the Doob martingale of \(h_S-\mathbb E h_S\), this yields the concentration lemma
\[
\Pr\bigl\{\|h_S-\mathbb E h_S\|\le D\,\alpha(n)\sqrt{2n\log(2/\delta)}\bigr\}\ge 1-\delta.
\]
Thus uniform argument stability implies that the random hypothesis is localized, with high probability, in a ball centered at its mean [1702.08712].

This localization is the basis for what the paper calls the algorithmic hypothesis class,
\[
r=r(n,\delta)=D\,\alpha(n)\sqrt{2n\log(2/\delta)},
\qquad
B_r=\{\,h\in H:\|h-\mathbb E h_S\|\le r\,\}.
\]
The generalization problem is then reduced to controlling the complexity of \(B_r\) rather than of the full ambient class.

## 3. Rademacher bounds and high-probability generalization

The localization step leads to an explicit Rademacher-complexity estimate for \(B_r\). Under the preceding assumptions,
\[
\mathcal R(B_r)
\le
D\,C_p\,B\,\sqrt{2\log(2/\delta)}\,\alpha(n)\,n^{-1/2+1/p}.
\]
In the Hilbert case, where \(D=1\), \(p=2\), and \(C_2=1\), this simplifies to
\[
\mathcal R(B_r)\le B\,\sqrt{2\log(2/\delta)}\,\alpha(n).
\]
The proof shifts the class by \(\mathbb E h_S\), uses
\[
\sup_{\|h-\mathbb E h_S\|\le r}\frac1n\sum \sigma_i \langle h-\mathbb E h_S,X_i\rangle
=
\frac r n \Big\|\sum \sigma_i X_i\Big\|_*,
\]
and then invokes the type-\(p\) inequality [1702.08712].

For Hilbert spaces, this yields a direct high-probability excess-risk estimate. If the loss is bounded by \(M\) and \(L\)-Lipschitz, then with probability at least \(1-2\delta\),
\[
R(h_S)-R_S(h_S)
\le
2LB\,\sqrt{2\log(2/\delta)}\,\alpha(n)
+
M\sqrt{\tfrac{\log(1/\delta)}{2n}}.
\]
The first term is the stability-driven term and the second is the standard sampling term. The significance of the result is structural: the dependence on the learning algorithm enters only through \(\alpha(n)\), while the remainder of the bound is mediated by the geometry of the hypothesis space and by boundedness and Lipschitz assumptions [1702.08712].

The same framework also supports a localized fast-rate statement. For any \(a>1\), with probability \(1-2\delta\),
\[
R(h_S)-\frac a{a-1}R_S(h_S)
\le
8LB\,\sqrt{2\ln(2/\delta)}\,\alpha(n)
+
\frac{(6a+8)M\ln(1/\delta)}{3n}.
\]
The proof uses Bartlett–Mendelson locality on a deformed class. In this sense, uniform argument stability is not merely a replacement for loss-stability estimates; it is a route to localized high-probability control through an algorithm-dependent neighborhood around \(\mathbb E h_S\) [1702.08712].

## 4. Instantiations: regularized ERM and stochastic gradient descent

The abstract framework becomes concrete once \(\alpha(n)\) is computed for specific algorithms. For regularized empirical risk minimization,
\[
R_{S,\lambda}(h)=\frac1n\sum_{i=1}^n \ell(h,Z_i)+\lambda N(h),
\]
assume the penalty \(N\) satisfies, for some \(C>0\) and \(\xi>1\),
\[
N(h_S)+N(h_{S^i})-2N\!\Bigl(\frac{h_S+h_{S^i}}2\Bigr)
\ge
C\,\|h_S-h_{S^i}\|^\xi.
\]
Then one obtains
\[
\alpha(n)=\Bigl(\frac{LB}{C\lambda n}\Bigr)^{1/(\xi-1)}.
\]
For \(N(h)=\|h\|^2\), one takes \(\xi=2\), giving \(\alpha(n)=O((\lambda n)^{-1})\) [1702.08712]. Combined with the localized generalization theorem, this yields explicit high-probability rates driven by regularization strength and sample size.

For stochastic gradient descent in \(\mathbb R^d\), with \(\ell\) both \(L\)-Lipschitz and \(s\)-smooth, the paper imports a stability estimate of the form
\[
\|h_T-h_T^i\|
\le
\frac{1+1/(sc)}{n-1}\,(2cBL)^{1/(sc+1)}\,T^{sc/(sc+1)}
\]
after \(T\) updates with suitably decaying \(\alpha_t\). Substituting this into the general theorem gives a corresponding high-probability deformed bound. Variants for convex or strongly convex losses yield the simpler rates \(\alpha(n)=O(\tfrac1n\sum_t \alpha_t)\) or \(\alpha(n)=O(\tfrac1n)\), and hence \(O(1/n)\) generalization [1702.08712].

These examples clarify the operational role of uniform argument stability. It is not a property attached only to a hypothesis class; rather, it is an algorithm-level sensitivity parameter that can be derived from optimization or regularization structure and then inserted into a generalization theorem.

## 5. Relation to classical uniform stability and worst-case limits

Classical uniform stability is formulated directly at the level of losses. In "Stability of Multi-Task Kernel Regression Algorithms" [1306.3905], if \(Z=\{(x_1,y_1),\dots,(x_m,y_m)\}\), \(Z^{\setminus i}=Z\setminus\{(x_i,y_i)\}\), and \(A:Z\mapsto f_Z\in H\), then \(\beta\)-uniform stability means
\[
\bigl|c(y,f_Z,x)-c(y,f_{Z^{\setminus i}},x)\bigr|\le \beta
\]
for all \(m\), all \(i\), all training sets, and all test points \((x,y)\) independent of \(Z\). In an operator-valued RKHS \(H\subset \mathcal Y^{\mathcal X}\) with regularized empirical risk
\[
R_{\mathrm{emp}}(f,Z)=\frac1m\sum_{k=1}^m c(y_k,f,x_k),
\qquad
R_{\mathrm{reg}}(f,Z)=R_{\mathrm{emp}}(f,Z)+\lambda \|f\|_H^2,
\]
the minimizer \(f_Z=\arg\min_{f\in H}R_{\mathrm{reg}}(f,Z)\) is uniformly stable under the bounded-kernel, well-posedness, convexity, and Lipschitz assumptions (H1)–(H3), with
\[
\beta=\frac{C^2\kappa^2}{2m\lambda}.
\]
This leads, under bounded loss \(M\), to the standard Bousquet–Elisseeff high-probability estimate
\[
\mathbb E_{(X,Y)}[c(Y,f_Z,X)]
\le
R_{\mathrm{emp}}(f_Z,Z)+2\beta+(4m\beta+M)\sqrt{\frac{\ln(1/\delta)}{2m}}
\]
[1306.3905]. The contrast with uniform argument stability is conceptual: the latter first controls \(\|h_S-h_{S^i}\|\), then derives loss stability and generalization from geometry and Lipschitz structure.

The broader theory of uniform stability has sharp worst-case limitations. "A Tight Lower Bound for Uniformly Stable Algorithms" proves that for any uniformly \(\gamma\)-stable algorithm and \(L\)-bounded loss, recent upper bounds give
\[
R_{\rm pop}(A_S)-R_{\rm emp}(A_S)
=
\widetilde O\!\Bigl(\gamma+\frac{L}{\sqrt n}\Bigr),
\]
more precisely
\[
R_{\rm pop}(A_S)-R_{\rm emp}(A_S)
=
O\!\Bigl(\gamma\log n\log\tfrac1\delta+\tfrac{L}{\sqrt n\,\sqrt{\log(1/\delta)}}\Bigr),
\]
and the paper matches this up to logarithmic factors with a lower bound: for any \(0<\gamma\le L\) and integer \(n\), there exist a domain, a distribution, an \(L\)-bounded loss, and a \(\gamma\)-stable algorithm such that with constant probability,
\[
R_{\rm pop}(A_S)-R_{\rm emp}(A_S)\ge \frac{\gamma}{4}+\frac{L}{32\sqrt n}
=
\Omega\!\bigl(\gamma + n^{-1/2}L\bigr)
\]
[2012.13326]. This shows that the dependence on \(\gamma\) and \(L/\sqrt n\) cannot be improved, beyond logarithmic factors, for uniformly stable algorithms.

Because argument stability implies uniform stability under \(L\)-Lipschitzness and \(\|X\|_*\le B\), a plausible implication is that any reduction from argument stability to classical loss stability inherits these worst-case barriers at the loss-stability layer. At the same time, the Banach-space localization method of [1702.08712] extracts additional structure from the norm perturbation itself rather than stopping at the induced loss difference.

## 6. Adjacent developments and methodological extensions

Subsequent work on optimization-driven stability has emphasized direct analysis of iterate dynamics. In the \(\beta\)-smooth, \(\gamma\)-strongly convex regime, "A Unified Lyapunov-IQC Framework for Uniform Stability of Smooth Quadratic First-Order Accelerated Optimizers" defines \(\epsilon_n\)-uniform stability for a first-order optimizer \(A\) by
\[
\sup_{S,S' \text{ differing in one example}}
\mathbb E\bigl[|\ell(w_n,z)-\ell(w_n',z)|\bigr]
\le \epsilon_n,
\]
where \(w_n\) and \(w_n'\) are the terminal iterates on neighboring samples. Under \(G\)-Lipschitzness of \(\ell\) in its first argument,
\[
| \mathbb E[f(w_n)-f(w_n')] |
\le
G\,\mathbb E[\|w_n-w_n'\|].
\]
For SGD with fixed step size \(\eta\le 1/\beta\), the classical coupling argument produces \(\epsilon_n=O(1/n)\). For Nesterov accelerated gradient, the presence of momentum breaks the simple one-step contraction argument, and the paper develops a quadratic Lyapunov and then a Lyapunov–IQC certification via an LMI solvable by SDP. In the smooth-quadratic regime this yields \(O(1/\sqrt n)\) uniform stability for NAG [2605.08488].

These developments do not replace uniform argument stability; they delineate a neighboring methodological direction. Uniform argument stability analyzes the perturbation of the output hypothesis in a Banach or Hilbert norm and then leverages hypothesis-space geometry. Lyapunov–IQC methods analyze optimizer state trajectories directly and certify contraction through control-theoretic machinery. This suggests a broader taxonomy of stability analyses: hypothesis-level stability, loss-level stability, and state-space dynamical stability.

Within that taxonomy, uniform argument stability occupies a distinctive position. It is stronger than classical loss stability, it naturally interfaces with Banach-space concentration, and it supports localized generalization analyses that are sensitive to the geometry of the hypothesis space and to algorithm-specific perturbation bounds. Its principal limitation is equally clear: useful generalization theorems still require boundedness, Lipschitzness, and structural assumptions on the ambient space, and any passage through loss stability is subject to the sharp worst-case limits known for uniformly stable algorithms [1702.08712; 2012.13326].

Source: https://www.emergentmind.com/topics/uniform-argument-stability