---
title: Small-Covariance Noise-to-State Stability (scNSS)
url: https://www.emergentmind.com/topics/small-covariance-noise-to-state-stability-nss
type: topic
---

# Small-Covariance Noise-to-State Stability (scNSS)

Small-covariance noise-to-state stability (scNSS) is a stochastic stability notion for systems whose state remains, with high probability, in a covariance-dependent neighborhood of an equilibrium, but only when the driving noise covariance is sufficiently small. In the formulation introduced for stochastic gradient dynamics, scNSS preserves the standard NSS template
\[
\mathbb P\!\left\{ \mathcal V(\chi(t))\le \beta(\mathcal V(\chi(0)),t)+\gamma(\Sigma\Sigma^\top) \right\}\ge 1-\epsilon,
\]
while restricting the admissible covariance magnitude to
\[
\|\Sigma\Sigma^\top\|_\infty<d,
\]
with \(\gamma\in\mathcal K_{[0,d)}\) rather than \(\gamma\in\mathcal K\). The concept was introduced to capture robustness regimes in which stability is available only for sufficiently small covariance, especially in stochastic optimization and Langevin-type dynamics [2509.24277].

## 1. Formal definition and state-space setting

The scNSS framework is posed for stochastic systems of the form
\[
d\chi(t)=f(\chi(t))\,dt+g(\chi(t))\,\Sigma(t)\,dB(t),
\]
where \(\chi(t)\in\mathcal S\subset\mathbb R^n\), \(\mathcal S\) is open and diffeomorphic to \(\mathbb R^n\), \(B(t)\) is an \(m\)-dimensional standard Brownian motion, \(f:\mathcal S\to\mathbb R^n\) and \(g:\mathcal S\to\mathbb R^{n\times m}\) are locally bounded and locally Lipschitz, \(f(\chi^*)=0\) at an equilibrium \(\chi^*\), and \(\Sigma:\mathbb R_+\to\mathbb R^{m\times m}\) is Borel measurable and locally essentially bounded. The instantaneous noise covariance is \(\Sigma(t)\Sigma(t)^\top\) [2509.24277].

A size function \(\mathcal V:\mathcal S\to\mathbb R_+\) is twice continuously differentiable, positive definite with respect to \(\chi^*\), and coercive in the sense that
\[
\chi_k\to \partial\mathcal S \text{ or } \|\chi_k\|\to\infty
\quad\Rightarrow\quad
\mathcal V(\chi_k)\to\infty.
\]
This function supplies the state-energy variable in the stability estimate.

The distinction between NSS and scNSS is entirely in the disturbance channel. Standard NSS in probability requires the bound
\[
\mathbb P\!\left\{ \mathcal V(\chi(t))\le \beta(\mathcal V(\chi(0)),t)+\gamma(\Sigma\Sigma^\top) \right\}\ge 1-\epsilon
\]
for some \(\beta\in\mathcal{KL}\), \(\gamma\in\mathcal K\), all \(t\ge 0\), all initial conditions, and all \(\epsilon\in(0,1)\). By contrast, scNSS in probability uses the same probabilistic form but only for sufficiently small covariance, with \(\gamma\in\mathcal K_{[0,d)}\) and the side condition \(\|\Sigma\Sigma^\top\|_\infty<d\). The definition therefore encodes a local-in-covariance robustness property rather than a global-in-covariance one. The same source also recalls integral NSS (iNSS),
\[
\mathbb P\!\left\{ \mathcal V(\chi(t))\le \beta(\mathcal V(\chi(0)),t)+\int_0^t \gamma(\Sigma(\tau)\Sigma(\tau)^\top)\,d\tau \right\}\ge 1-\epsilon,
\]
which weakens the disturbance channel further by allowing an accumulated energy term instead of an instantaneous covariance gain.

## 2. Lyapunov characterization and hierarchy of stochastic robustness

The central Lyapunov inequalities separate NSS, scNSS, and iNSS by the admissible comparison function class in the dissipation term. An NSS-Lyapunov function satisfies
\[
\mathcal L[\mathcal V](\xi,\Theta) \le -\alpha(\mathcal V(\xi))+\gamma(\Theta\Theta^\top),
\]
for \(\alpha\in\mathcal K_\infty\), \(\gamma\in\mathcal K\), all \(\xi\in\mathcal S\), and all \(\Theta\). An scNSS-Lyapunov function satisfies the same inequality with \(\alpha\in\mathcal K\), \(\gamma\in\mathcal K_{[0,d)}\), and only for \(\Theta\Theta^\top<d\). An iNSS-Lyapunov function weakens dissipation further to \(\alpha\in\mathcal{PD}\), \(\gamma\in\mathcal K\) [2509.24277].

The corresponding implication results are direct: existence of an scNSS-Lyapunov function implies scNSS, existence of an NSS-Lyapunov function implies NSS, and existence of an iNSS-Lyapunov function implies iNSS. In the scNSS proof, stopping times and supermartingale arguments are used to show that trajectories enter and remain in a neighborhood
\[
D=\left\{\xi\in\mathcal S:\ \mathcal V(\xi)\le \alpha^{-1}\!\circ c\,\gamma(\Sigma\Sigma^\top)\right\},
\]
whose size is controlled by the covariance. A representative estimate is
\[
\mathbb P\!\left\{ \mathcal V(\chi(t)) \le \beta_1(\mathcal V(\chi(0)),t) +\frac{2}{\epsilon}\,\alpha^{-1}\!\circ c\,\gamma(\Sigma\Sigma^\top) \right\}\ge 1-\epsilon,
\]
valid when \(\Sigma\Sigma^\top\) is below a small-covariance threshold \(d_1\).

Within this framework, NSS is the strongest notion, scNSS is weaker because it only covers \(\Sigma\Sigma^\top<d\), and iNSS is weaker still because it only requires positive-definite dissipation and admits an integral disturbance term. The same hierarchy appears at the level of optimization geometry:
\[
K_\infty\text{-PL} \Rightarrow \text{NSS},\qquad
K\text{-PL} \Rightarrow \text{scNSS},\qquad
\mathcal{PD}\text{-PL} \Rightarrow \text{iNSS}.
\]
A common misconception is to treat scNSS as merely NSS with smaller constants. The formal difference is sharper: the gain itself is only defined on a bounded covariance interval, so the guarantee is structurally local in covariance magnitude.

## 3. Gradient dynamics and generalized Polyak–Łojasiewicz structure

The principal application domain for scNSS is stochastic gradient dynamics of the overdamped Langevin type,
\[
dz(t)=-\nabla J(z(t))\,dt+G(z(t))\,\Sigma(t)\,dB(t).
\]
The standing assumptions are that \(J\in C^2\), \(J\) is bounded below and coercive, and \(G\) is locally Lipschitz and locally bounded. When \(\nabla J\) is globally \(L\)-Lipschitz and \(G\) is globally bounded by \(K_G\), the Lyapunov candidate
\[
V(z)=J(z)-J^*
\]
satisfies
\[
\mathcal L[V](z,\Sigma) \le -\|\nabla J(z)\|^2 +\frac12\,L K_G^2\,\Sigma\Sigma^\top.
\]
This makes the generalized Polyak–Łojasiewicz condition the decisive geometric hypothesis [2509.24277].

The paper distinguishes three variants:
\[
\|\nabla J(z)\|\ge \mu(J(z)-J^*),
\]
with \(\mu\in\mathcal K_\infty\) for \(K_\infty\)-PL, \(\mu\in\mathcal K\) for \(K\)-PL, and \(\mu\in\mathcal{PD}\) for positive-definite-PL. Under global Lipschitzness and bounded noise gain, the theorem for overdamped diffusion states that \(K_\infty\)-PL implies NSS, \(K\)-PL implies scNSS and iNSS, and \(\mathcal{PD}\)-PL implies iNSS. The neighborhood size around the optimum is determined by the covariance term in the Lyapunov inequality; smaller covariance produces a smaller ultimate region.

The same trichotomy extends to underdamped Langevin diffusion, or stochastic heavy-ball dynamics,
\[
dz(t)=v(t)\,dt,\qquad
dv(t)= -\eta\nabla J(z(t))\,dt - c\,v(t)\,dt + G(z(t),v(t))\Sigma(t)\,dB(t).
\]
Here the analysis uses the mixed Lyapunov function
\[
V_2(z,v)=J(z)-J^*+\lambda_1\langle v,\nabla J(z)\rangle+\frac{\lambda_2}{2}\|v\|^2,
\]
and obtains an estimate of the form
\[
\mathcal L[V_2] \le -\lambda_1\eta\,\mu_1(\cdot)+\frac{\lambda_2}{2}K_G^2\,\Sigma\Sigma^\top.
\]
Under global Lipschitz gradient and bounded \(G\), the same three-way conclusion holds: \(K_\infty\)-PL yields NSS, \(K\)-PL yields scNSS and iNSS, and \(\mathcal{PD}\)-PL yields iNSS.

## 4. Non-globally-Lipschitz objectives and step-size tuning

When \(\nabla J\) is not globally Lipschitz, the fixed-drift overdamped model is replaced by a state-dependent step-size dynamics,
\[
dz(t)=-\eta(J(z(t)))\nabla J(z(t))\,dt+G(z(t))\Sigma(t)\,dB(t).
\]
The analysis is organized through sublevel sets
\[
Z_h=\{z:\ J(z)-J^*\le h\},
\]
and curvature envelopes
\[
\bar L(h)=\frac12\max_{z\in Z_h}\nabla^2J(z)\,G(z)^2,\qquad
\tilde L(h)=\bar L(h)-\bar L(0).
\]
Under the \(K\)-PL condition, the ratio
\[
\frac{\tilde L(h)}{\eta(h)\mu(h)^2}
\]
determines whether the system is merely small-covariance stable or globally NSS [2509.24277].

The resulting dichotomy is explicit. If
\[
\lim_{h\to\infty}\frac{\tilde L(h)}{\eta(h)\mu(h)^2}=d_2>0,
\]
then the dynamics are scNSS. If instead
\[
\lim_{h\to\infty}\frac{\tilde L(h)}{\eta(h)\mu(h)^2}=0,
\]
the dynamics are NSS. The interpretation given is that the learning rate \(\eta(h)\) must dominate the local curvature growth \(\tilde L(h)\) relative to the PL gain. Two concrete examples are provided: \(\eta(h)=\tilde L(h)\) gives scNSS, whereas \(\eta(h)=h\,\tilde L(h)\) gives NSS.

For the underdamped non-globally-Lipschitz case, a different Lyapunov construction is used,
\[
V_3(z,v)=\varphi_2(J(z)-J^*)+\langle \nabla J(z),v\rangle+\|v\|^2,
\]
where \(\varphi_2\) is obtained by smoothing a \(K_\infty\) bound on the gradient via local curvature. Under suitable tuning of \(\eta\) and \(c\), the same \(K_\infty/K/\mathcal{PD}\) trichotomy again yields NSS, scNSS, and iNSS. This suggests that scNSS is not tied to global smoothness, but rather to the balance among curvature growth, dissipation, and covariance magnitude.

## 5. Canonical applications: LQR and logistic regression

For continuous-time linear quadratic regulator policy optimization, the objective is
\[
\min_{K\in G} J_2(K)=\operatorname{Tr}(P_K),
\]
where \(P_K\) solves
\[
(A-FK)^\top P_K + P_K(A-FK) + Q + K^\top R K = 0,
\]
and the gradient is
\[
\nabla J_2(K)=2(RK-F^\top P_K)Y_K,
\]
with \(Y_K\) solving
\[
(A-FK)Y_K + Y_K(A-FK)^\top + I=0.
\]
The objective \(J_2\) is coercive and analytic, \(\nabla J_2\) is only locally Lipschitz, and \(J_2\) satisfies the \(K\)-PL condition
\[
\|\nabla J_2(K)\|\ge \mu_5(J_2(K)-J_2^*),\qquad
\mu_5(h)=\frac{h}{b_1 h+b_2}.
\]
For the overdamped stochastic policy dynamics,
\[
dK(s)= -2\eta(J(K(s))-J^*)(RK(s)-F^\top P(s))Y(s)\,ds+\Sigma_1(s)\,dW(s),
\]
the criterion is
\[
\lim_{h\to\infty}\frac{h^3}{\eta(h)}=0
\quad\Rightarrow\quad \text{NSS},
\]
while
\[
0<\lim_{h\to\infty}\frac{h^3}{\eta(h)}<\infty
\quad\Rightarrow\quad \text{scNSS}.
\]
For underdamped or heavy-ball LQR policy dynamics, the same general underdamped theorem and the \(K\)-PL property yield scNSS.

For logistic regression, the loss is
\[
J_3(\theta)=\frac1N\sum_{i=1}^N \ell_i(\theta), \qquad
p_i(\theta)=\frac{1}{1+e^{-\theta^\top x_i}}.
\]
Under the assumptions that the data matrix \(X\) has full row rank and the data are nonseparable, \(J_3\) is coercive, strictly convex, and has globally Lipschitz gradient with constant
\[
\frac{1}{4N}\|XX^\top\|.
\]
It also satisfies a \(K\)-PL condition
\[
\|\nabla J_3(\theta)\|\ge \mu_6(J_3(\theta)-J_3^*)
\]
for some \(K\)-function \(\mu_6\), but it does not satisfy the \(K_\infty\)-PL condition because \(\|\nabla J_3(\theta)\|\) is uniformly bounded while \(J_3(\theta)\to\infty\) as \(\|\theta\|\to\infty\). Consequently, both the overdamped diffusion
\[
d\theta(s)=-\nabla J_3(\theta(s))\,ds+\Sigma(s)\,dB(s)
\]
and the underdamped dynamics
\[
d\theta(s)=v(s)\,ds,\qquad
dv(s)=-\eta\nabla J_3(\theta(s))\,ds-cv(s)\,ds+\Sigma(s)\,dB(s)
\]
are concluded to be scNSS rather than NSS. The operational interpretation is that the stochastic training trajectory remains, with high probability, in a covariance-dependent neighborhood of the logistic optimum, and that neighborhood shrinks as the covariance decreases [2509.24277].

## 6. Relation to earlier NSS theory and adjacent extensions

scNSS sits within a broader NSS lineage for stochastic systems with persistent or state-dependent noise. Earlier NSS theory for SDEs with persistent noise already expressed state bounds in terms of the maximum covariance magnitude through \(\Sigma(t)\), using NSS in probability and \(p\)th-moment NSS with Lyapunov conditions of the form
\[
LV(x,t)\le -W(x)+\sigma(\Sigma(t)).
\]
That framework did not isolate a separate small-covariance notion, but it already encoded the principle that smaller covariance implies a smaller disturbance-dependent state bound [1501.05008].

A different precursor is the time-domain analysis of noise-to-state exponentially stable systems. For SDEs
\[
dx(t)=f(x)\,dt+h(x)\Sigma(t)\,d\mathcal B(t),
\]
the resident-time ratio
\[
D(r)\triangleq \liminf_{T\to+\infty}\frac{1}{T}\int_0^T \mathbf{1}_{\{\|x(t)\|_2\le r\}}\,dt
\]
was introduced to quantify the fraction of time a trajectory spends in a bounded region. The resulting almost-sure lower bound improves as the noise intensity decreases and tends to \(1\) as the noise vanishes. That work does not study a small-covariance regime in the perturbative sense, but it provides a time-average small-noise enhancement of NSS-type behavior and complements fixed-time probabilistic boundedness by a long-run residence guarantee [1607.02800].

More recent neighboring developments broaden the same theme rather than replacing it. Stochastic contraction results give incremental noise- and input-to-state stability in weighted \(\ell_2\)-norms, with asymptotic mean-square errors scaling linearly with diffusion magnitude, such as terms of order \(\sigma_x^2/c\) and \((\ell^2/c^2)(\sigma_\xi^2/c)\); this is closely aligned with the intuition behind small-covariance robustness [2602.18382]. In Wasserstein space, distributional ISS has been proposed for measure-valued dynamics, with the claim that on compact domains it recovers particle-level ISS and NSS, including stochastic particle systems whose disturbance magnitude is \(\|\Sigma\Sigma^\top\|_\infty\) [2603.28910]. For discrete-time linear systems, risk-aware stability introduces NSS-type offsets measured through coherent risk functionals or a mean-conditional-variance functional, so that covariance and higher-order centered noise statistics affect the residual state-risk bound [2211.12416]. In randomly switching linear systems with additive Gaussian noise, lifted covariance dynamics and invariant ellipsoids yield computable bounds on state covariance trajectories; this has been interpreted as a small-covariance NSS-type property because smaller mode-dependent covariances \(Q_i\) lead to tighter covariance confinement [1905.09427].

Taken together, these results delimit the scope of scNSS. It is neither an exact stationary-distribution theory nor a generic ergodic theorem, and it is distinct from fixed-time NSS, time-average residence bounds, distributional Wasserstein robustness, and risk-aware formulations. Its specific role is to formalize the regime in which covariance-dependent ultimate boundedness is valid only below a noise threshold, a setting that arises naturally in stochastic gradient systems and related control and optimization dynamics.

Source: https://www.emergentmind.com/topics/small-covariance-noise-to-state-stability-nss