---
title: Normalized Directional Sharpness (NDS)
url: https://www.emergentmind.com/topics/normalized-directional-sharpness-nds
type: topic
---

# Normalized Directional Sharpness (NDS)

Searching arXiv for the cited papers to ground the article in current preprints.
Normalized Directional Sharpness (NDS) is a direction-sensitive sharpness construct that appears in several distinct technical settings rather than as a single universal formula. In recent arXiv usage, it denotes at least four related normalizations: a learning-rate-normalized mini-batch curvature statistic for the Edge of Stochastic Stability (EoSS) in momentum training, $\mathrm{NDS}=\eta\,\mathrm{BS}$; a Hessian Rayleigh quotient along an optimizer’s realized update direction, $\Delta\theta^\top H\Delta\theta/\|\Delta\theta\|^2$; a normalized fluctuation statistic over short SAM or ASAM probe trajectories, $\mathrm{Std}(\{\log(r_t+\varepsilon)\})$ with $r_t=s_t^2/\mathcal L_{\xi_t}$; and, in satellite image quality assessment, a normalized decay rate of gradients along pronounced edges [2604.14108] [2606.04662] [2606.25004] [2410.10488]. The shared theme is normalization of a directional sharpness signal so that operative comparisons are less confounded by learning rate, step magnitude, loss scale, parameterization, contrast, or exposure.

## 1. Taxonomy of definitions

The term covers multiple non-equivalent quantities whose common structure is directional measurement plus explicit normalization. In stochastic optimization near EoSS, the underlying directional statistic is Batch Sharpness,
\[
\mathrm{BS}(\theta)
:=
\mathbb{E}_{B\sim \mathcal P_b}\!\left[
\frac{g_B(\theta)^\top H_B(\theta)\,g_B(\theta)}{\|g_B(\theta)\|_2^2}
\right],
\]
and the natural normalization is
\[
\mathrm{NDS}(\theta):=\eta\cdot \mathrm{BS}(\theta).
\]
In optimizer comparison, NDS is the Hessian Rayleigh quotient along the actual update,
\[
\mathrm{NDS}(\Delta\theta):=\frac{\Delta\theta^\top H\Delta\theta}{\|\Delta\theta\|^2}.
\]
In certification, the central object is a directional-sharpness time series generated by short SAM-style dynamics and normalized by per-step loss,
\[
r_t:=\frac{s_t^2}{\mathcal L_{\xi_t}},
\qquad
\mathrm{NDS}_{\mathcal C}(\vec w_0,D,r)
:=
\mathrm{Std}\Big(\{\log(r_t+\varepsilon)\}_{t=0}^{T-1}\Big).
\]
In satellite imagery, NDS is used as a compact name for the normalized decay rate of gradients along pronounced edges, with directional scores
\[
\mathrm{NDS}_x=\frac{1}{n}\sum_{i=1}^n \Delta G_{x,i},
\qquad
\mathrm{NDS}_y=\frac{1}{n}\sum_{i=1}^n \Delta G_{y,i}.
\]
These definitions are not interchangeable, but each treats sharpness as a quantity that must be evaluated along an operational direction rather than by a purely global worst-case surrogate [2604.14108] [2606.04662] [2606.25004] [2410.10488].

| Setting | Direction | Normalization |
|---|---|---|
| Momentum EoSS | Mini-batch gradient direction $g_B$ | $\mathrm{NDS}=\eta\,\mathrm{BS}$ |
| Optimizer comparison | Realized update $\Delta\theta$ | Divide by $\|\Delta\theta\|^2$ |
| Certification | SAM/ASAM probe trajectory | $\mathrm{Std}(\log(s_t^2/\mathcal L_{\xi_t}))$ |
| Satellite imagery | Pronounced edge directions | Divide gradient decay by original edge gradient |

A common misconception is to treat NDS as a synonym for maximum curvature or for a single Hessian-based scalar. The relevant papers instead define it through optimizer directions, stochastic mini-batch directions, SAM probe directions, or edge-normal directions, and the normalization target changes with the application.

## 2. Learning-rate-normalized directional sharpness at the Edge of Stochastic Stability

In "Momentum Further Constrains Sharpness at the Edge of Stochastic Stability" [2604.14108], the operative curvature statistic is Batch Sharpness, the expected directional mini-batch curvature along the stochastic descent direction. The paper emphasizes that this quantity differs from $\lambda_{\max}$, the largest eigenvalue of the full-batch Hessian. In full-batch GD, $\lambda_{\max}\approx 2/\eta$ marks the deterministic Edge of Stability, but in mini-batch training $\lambda_{\max}$ may plateau far below $2/\eta$ and does not change instantaneously with batch-size interventions, whereas catapult events align with Batch Sharpness crossing its operating plateau. This is why the directional statistic, rather than the maximum curvature, is treated as the empirically relevant quantity for EoSS [2604.14108].

The normalization
\[
\mathrm{NDS}(\theta):=\eta\cdot \mathrm{BS}(\theta)
\]
removes the trivial $1/\eta$ scaling that repeatedly appears in the observed curvature thresholds. Under SGDM and SGDN, the paper identifies two batch-size-dependent plateau regimes. In the small-batch, noise-dominated regime,
\[
\mathrm{BS}_{\mathrm{plateau}}\approx \frac{2(1-\beta)}{\eta}
\quad\Longrightarrow\quad
\mathrm{NDS}_{\mathrm{plateau}}\approx 2(1-\beta).
\]
In the large-batch regime, the plateau depends on the momentum variant. For SGDM,
\[
\mathrm{BS}_{\mathrm{plateau}}\approx \frac{2(1+\beta)}{\eta}
\quad\Longrightarrow\quad
\mathrm{NDS}_{\mathrm{plateau}}\approx 2(1+\beta),
\]
whereas for SGDN,
\[
\mathrm{BS}_{\mathrm{plateau}}\approx \frac{2(1+\beta)}{\eta(1+2\beta)}
\quad\Longrightarrow\quad
\mathrm{NDS}_{\mathrm{plateau}}\approx \frac{2(1+\beta)}{1+2\beta}.
\]
For vanilla SGD, $\beta=0$ recovers
\[
\mathrm{BS}_{\mathrm{plateau}}\approx \frac{2}{\eta}
\quad\Longrightarrow\quad
\mathrm{NDS}_{\mathrm{plateau}}\approx 2.
\]

These plateau values are central because they show that momentum does not merely shift a single stability threshold. In small batches, momentum lowers the operative normalized plateau to $2(1-\beta)$, whereas in large batches SGDM raises it to $2(1+\beta)$. The paper explicitly highlights this qualitative flip: momentum favors flatter regions in the noise-dominated regime but sharper regions in the near-deterministic regime [2604.14108].

The measurement protocol is also directional. At parameter iterate $\theta$, one computes or approximates $g_B(\theta)$ and $H_B(\theta)$ on random mini-batches $B\sim\mathcal P_b$, forms the Rayleigh quotient
\[
\frac{g_B(\theta)^\top H_B(\theta)\,g_B(\theta)}{\|g_B(\theta)\|_2^2},
\]
and averages over batches to obtain $\mathrm{BS}(\theta)$. Tracking $\mathrm{BS}$ over training reveals progressive sharpening followed by stabilization at the batch-regime-dependent plateau; multiplying by $\eta$ yields NDS, which exposes the optimizer dependence through $\beta$.

## 3. Linear stability interpretation and hyperparameter coupling

The same EoSS work derives the momentum thresholds from linearized dynamics near a minimizer $\theta^\star$. With $v_t$ the velocity and $x_t:=\theta_t-\theta^\star$, SGDM is written as
\[
v_{t+1}=\beta v_t+\widehat H_t\,x_t,
\qquad
x_{t+1}=x_t-\eta\,v_{t+1},
\]
where $\widehat H_t=\tfrac1b\sum_{j\in B_t}H_j$ is the random mini-batch Hessian under quadratic approximation and interpolation [2604.14108].

In one dimension, with random curvature $h_t$ of mean $a=\mathbb E[h_t]$ and variance $\sigma_b^2=\mathrm{Var}(h_t)$, the dominant eigenvalue of the mean-square operator expands as
\[
\lambda_\star(\eta)
=
1-\frac{\eta}{1-\beta}\,2a
+\frac{\eta^2}{(1-\beta)^2}\,\sigma_b^2
+
O\!\left(\frac{\eta^2 a^2}{(1-\beta)^3}\right).
\]
The resulting exact boundary for mean-square stability interpolates between deterministic and stochastic limits:
\[
\frac{1}{\eta_{\max}}
=
\frac{a}{2(1+\beta)}
+
\frac{\sigma_b^2}{2a(1-\beta)}.
\]
As $\sigma_b^2\to 0$, this recovers the deterministic heavy-ball threshold $\eta_{\max}=2(1+\beta)/a$. As $\sigma_b^2\gg a^2$, the effective step size
\[
\eta_{\mathrm{eff}}=\frac{\eta}{1-\beta}
\]
governs stability, yielding the small-batch threshold $\mathrm{BS}_{\mathrm{plateau}}\approx 2(1-\beta)/\eta$ and thus $\mathrm{NDS}_{\mathrm{plateau}}\approx 2(1-\beta)$.

The multidimensional slow-mode constraint in the noise-dominated regime is
\[
\rho\!\left(I_{d^2}-\eta_{\mathrm{eff}}\,K+\eta_{\mathrm{eff}}^2\,G\right)<1,
\qquad
\eta_{\mathrm{eff}}=\frac{\eta}{1-\beta},
\]
with $\bar H=\mathbb E[\widehat H_t]$, $G=\mathbb E[\widehat H_t\otimes \widehat H_t]$, and $K=\bar H\otimes I_d+I_d\otimes \bar H$. The paper notes that this coincides with the vanilla-SGD mean-square stability condition evaluated at $\eta_{\mathrm{eff}}$, explaining why the small-batch momentum threshold matches SGD’s EoSS evaluated at the amplified effective step size [2604.14108].

This stability picture has direct tuning implications. In the small-batch regime, keeping $\eta/(1-\beta)$ approximately fixed maintains the NDS plateau and avoids crossing the stochastic edge. Mid-run changes that lower the effective threshold, such as increasing $\eta$, increasing $\beta$, or decreasing $b$, produce catapults exactly when Batch Sharpness rises above the new plateau. Stabilizing changes reopen progressive sharpening until $\mathrm{BS}$ approaches the new, higher plateau. The paper’s empirical demonstrations on MLPs and CNNs up to ResNet-18 on CIFAR-10 and SVHN therefore present NDS not only as a descriptive statistic but also as a control variable for instability-adjacent training [2604.14108].

The main limitation is theoretical. Unlike vanilla SGD, where crossing $\mathrm{BS}>2/\eta$ is a proved sufficient instability criterion on quadratics, the momentum analysis does not provide a direct momentum-specific theorem making BS a formal certificate. Instead, the small-batch SGDM result is reduced to vanilla SGD with $\eta_{\mathrm{eff}}=\eta/(1-\beta)$.

## 4. Hessian Rayleigh-quotient NDS in optimizer comparison

In "Why Muon Outperforms Adam: A Curvature Perspective" [2606.04662], NDS is defined as the Rayleigh quotient of the Hessian along the optimizer’s realized update direction:
\[
\mathrm{NDS}(\Delta\theta):=
\frac{\Delta\theta^\top H\Delta\theta}{\|\Delta\theta\|^2}.
\]
The local second-order model is
\[
L(\theta+\Delta\theta)
\approx
L(\theta)+\nabla L(\theta)^\top \Delta\theta+\frac12\Delta\theta^\top H(\theta)\Delta\theta,
\]
so the curvature penalty can be decomposed as
\[
\frac12\Delta\theta^\top H\Delta\theta
=
\frac12\|\Delta\theta\|^2\cdot \mathrm{NDS}.
\]
Within this framework, the first-order gain is $\nabla L(\theta)^\top\Delta\theta$, and NDS isolates the second-order curvature actually paid by the realized step rather than by a global spectral extremum [2606.04662].

This formulation is used to explain optimizer differences at matched validation loss on a 124M-parameter NanoGPT trained on FineWeb-10B. The paper reports that Adam and Muon have comparable first-order decreases, while Muon consistently incurs a smaller curvature penalty. Decomposing the penalty shows that update norms are comparable, so the curvature-penalty gap is driven by lower NDS for Muon rather than by smaller steps. At matched validation loss, the Adam/Muon ratio of NDS averages $1.76$, whereas the ratio of squared update norm stays near $1$; over aligned training steps, the mean NDS ratio is $2.94$ [2606.04662].

The computational methodology is explicitly directional and avoids full Hessian materialization. At a training step $t$, with parameters $W_t$ and optimizer update $Z_t$, the experiments compute the Hessian–vector quadratic form $\langle Z_t,\mathcal H[Z_t]\rangle$ and normalize by $\|Z_t\|_F^2$. Directional sharpness is computed every $500$ steps through Hessian–vector products. For Muon, $Z_t^{\mathrm{Muon}}=\eta_t O_t$, where $O_t$ is the spectrally normalized momentum matrix, the polar factor $U_tV_t^\top$ of $B_t$; Muon uses Newton–Schulz orthogonalization with $5$ iterations, momentum warmed from $0.85$ to $0.95$, and leaves embeddings and $\mathrm{lm\_head}$ to Adam in the reported run [2606.04662].

The paper also decomposes total NDS into within-layer and cross-layer components,
\[
S_F(W_t;Z_t)=S_F^{\mathrm{within}}(W_t;Z_t)+S_F^{\mathrm{cross}}(W_t;Z_t),
\]
with
\[
S_F^{\mathrm{within}}(W_t;Z_t)
=
\sum_{\ell=1}^L
\frac{\langle Z_{t,\ell},\mathcal H_{\ell\ell}[Z_{t,\ell}]\rangle}{\|Z_t\|_F^2},
\]
\[
S_F^{\mathrm{cross}}(W_t;Z_t)
=
\sum_{\ell\neq \ell'}
\frac{\langle Z_{t,\ell},\mathcal H_{\ell\ell'}[Z_{t,\ell'}]\rangle}{\|Z_t\|_F^2}.
\]
For Muon, the cross-layer component drops faster, so the within-layer fraction rises from about $14\%$ early to about $44\%$ later, while Adam’s fraction stays comparatively stable at approximately $27\%\to 34\%$. A layerwise localization analysis attributes roughly $70\%$ of the within-layer Adam–Muon NDS gap to boundary layers $L1$ and $L12$, approximately $28\%$ to deep layers $L8$–$L11$, and approximately $2\%$ to middle layers $L2$–$L7$ [2606.04662].

The same paper links NDS to data imbalance. On Zipf-Probabilistic Context-Free Grammar data with imbalance exponent $s\in\{0,0.5,1\}$, the trajectory-averaged NDS rises with imbalance for both optimizers but much more for Adam. After normalization by Muon’s NDS at $s=0$, Adam’s normalized NDS increases from $1.63$ to $2.38$, whereas Muon’s increases from $1.00$ to $1.25$; the normalized gap grows from $0.63$ to $1.13$. The theoretical explanation is given through a quadratic model with heterogeneous positive curvatures $w_i$ and gradient alignment toward top-curvature modes. Muon’s spectral normalization equalizes amplitudes across active singular modes, so its per-step NDS becomes a group-size average,
\[
S_F(Z_t^{\mathrm{Muon}})
=
\alpha w_{\rm H}+(1-\alpha)w_{\rm L},
\]
whereas GD’s NDS is a residual-energy-weighted average,
\[
S_F(Z_t^{\mathrm{GD}})
=
P_t^{\mathrm{GD}}w_{\rm H}+(1-P_t^{\mathrm{GD}})w_{\rm L}.
\]
Under the stated heterogeneity and alignment conditions, Muon has smaller average NDS than GD for any finite horizon and, when curvature heterogeneity is sufficiently strong, also attains lower local quadratic loss after the same number of steps [2606.04662].

## 5. Dynamic NDS for certification, auditing, and SAM stability

"Certification of Machine Learning Models via Directional Sharpness" [2606.25004] introduces directional sharpness as a dynamic, SAM-based generalization metric and adopts an empirical normalization by default. The procedure starts from initial parameters $\vec w_0$, dataset $D$, public seed $r$, batch size $B$, number of steps $T$, a per-step sharpness function $S(\cdot)$, and a SAM update operator $\Phi_\Theta(\cdot)$. For each step $t$, it samples a mini-batch $\xi_t$, computes the mini-batch loss $\mathcal L_{\xi_t}(w_t)$, computes per-step sharpness $s_t$, and updates
\[
w_{t+1}:=\Phi_\Theta(w_t,\xi_t).
\]
For SAM,
\[
s_t=\rho\|\nabla \mathcal L_{\xi_t}(w_t)\|_q,
\]
and for ASAM,
\[
s_t=\rho\|\mathbf T_{w_t}\nabla \mathcal L_{\xi_t}(w_t)\|_q,
\qquad
\mathbf T_w:=\mathrm{diag}(|w|).
\]
The probe history is then aggregated by a fluctuation statistic. The default normalization is
\[
r_t:=\frac{s_t^2}{\mathcal L_{\xi_t}},
\qquad
\mathrm{NDS}_{\mathcal C}(\vec w_0,D,r)
=
\mathrm{Std}\Big(\{\log(r_t+\varepsilon)\}_{t=0}^{T-1}\Big),
\]
which is intended to reduce sensitivity to overall loss scale [2606.25004].

The paper’s theoretical analysis places this dynamic NDS in a local linearization regime near a minimum $w^\star$, under smoothness, a PL inequality, and bounded gradient noise:
\[
\|H(\vec w)\|_2\le L_s,
\qquad
\|g(\vec w)\|_2^2\ge 2\mu\,\mathcal L(\vec w),
\]
\[
\mathbb E[\|g_\xi(\vec w)-g(\vec w)\|_2^2]
\le
\frac{\sigma^2}{B}\,\mathcal L(\vec w).
\]
For mini-batch SAM sharpness with $p=q=2$ and radius $\rho$, the per-step squared sharpness satisfies the sandwich bound
\[
\underline\kappa_\rho \cdot \mathcal L(\vec w_t)
\le
\mathbb E[s_t^2\mid \vec w_t]
\le
\overline\kappa_\rho \cdot \mathcal L(\vec w_t),
\]
with
\[
\underline\kappa_\rho=2\rho^2\mu,
\qquad
\overline\kappa_\rho=\rho^2\Big(2L_s+\frac{\sigma^2}{B}\Big).
\]
If a minimum is linearly SAM-stable, then $\mathbb E[s_t^2]$ remains bounded by a constant times $\mathcal L(\vec w_0)$; conversely, exponential growth of $\mathbb E[s_t^2]$ implies SAM-instability [2606.25004].

The detection mechanism is explicitly mini-batch-sensitive. With per-example gradients $g_i$, average gradient $g$, and coherence
\[
c_g(\vec w):=
\frac{\|g(\vec w)\|_2^2}{\frac1N\sum_{i=1}^N\|g_i(\vec w)\|_2^2},
\]
the gap between RMS mini-batch SAM sharpness and full-batch sharpness is
\[
(\mathsf S_{\mathrm{rms}})^2-\mathsf S^2
=
\rho^2\cdot \frac{N-B}{B(N-1)}\cdot (1-c_g(\vec w))\cdot \frac1N\sum_{i=1}^N \|g_i\|_2^2.
\]
When $B\ll N$ and $c_g(\vec w)\ll 1$, this is approximately
\[
(\mathsf S_{\mathrm{rms}})^2-\mathsf S^2
\approx
\rho^2\cdot \frac1B\cdot \frac1N\sum_{i=1}^N\|g_i\|_2^2.
\]
This explains why dynamic directional sharpness can expose incoherence that static full-batch sharpness averages away [2606.25004].

Empirically, on CIFAR-10/100 with VGG-13/16/19-BN and WRN28-10 across SGD, Adam, SAM, and ASAM, the dynamic measure outperforms static baselines in correlation with generalization. Reported values include Spearman $\rho=0.902$, Kendall $\tau=0.702$, Kendall $\Psi=0.694$ for directional sharpness with $B=16$, $T=5$, compared with lower values for ASAM and magnitude-aware worst-case sharpness. In matched-accuracy benign/faulty pairs, the benign/faulty ratio is substantially smaller for the directional metric than for static sharpness in noisy labels, spurious-feature overfitting, backdoors, post-quantization quality, and small-test-set settings. The same work also makes the metric proof-compatible: a prover can commit to $D$ and $\vec w_0$, execute the public NDS computation in-circuit, and prove the predicate
\[
\mathsf{Pred}_\tau(\vec w_0,D,r)\equiv [\mathrm{NDS}_{\mathcal C}(\vec w_0,D,r)\le \tau]
\]
without revealing the data. For $T=5$ and $B=8$, proving NDS with the cited backend is reported as up to $80{,}000\times$ faster than proving an entire training run [2606.25004].

## 6. NDS in no-reference satellite image sharpness assessment

In "A Novel No-Reference Image Quality Metric For Assessing Sharpness In Satellite Imagery" [2410.10488], NDS is a compact name for the paper’s "normalized decay rate of gradients along pronounced edges." The grayscale image is
\[
I(x,y)\in\mathbb R,\qquad (x,y)\in \Omega\subset \mathbb Z^2,
\]
with spatial gradient
\[
\nabla I(x,y)=
\begin{bmatrix}
I_x(x,y)\\
I_y(x,y)
\end{bmatrix},
\qquad
\|\nabla I(x,y)\|=\sqrt{I_x(x,y)^2+I_y(x,y)^2}.
\]
Edge detection and gradient estimation are performed with a Sobel operator $(5\times 5)$. At an edge pixel, the unit normal and tangent are
\[
\mathbf n(x,y)=\frac{\nabla I(x,y)}{\|\nabla I(x,y)\|},
\qquad
\mathbf t(x,y)=
\begin{bmatrix}
-n_y(x,y)\\
n_x(x,y)
\end{bmatrix}.
\]
The directional gradient along the normal is
\[
g_n(x,y)=\nabla I(x,y)\cdot \mathbf n(x,y)=\|\nabla I(x,y)\|.
\]

The operational pipeline begins with preprocessing. High-frequency anomalies are filtered by a neighbors-based rule,
\[
p'_{i,j}=
\begin{cases}
\dfrac{1}{N}\sum_{(k,l)\in\mathcal N(i,j)} p_{k,l}
& \text{if }
\left|
\dfrac{p_{i,j}-\mu_{\mathcal N(i,j)}}{\mu_{\mathcal N(i,j)}}
\right|>\theta,
\\[6pt]
p_{i,j}
& \text{otherwise},
\end{cases}
\]
where
\[
\mu_{\mathcal N(i,j)}=\dfrac1N\sum_{(k,l)\in\mathcal N(i,j)}p_{k,l}.
\]
A low-high intensity mask excludes saturated and underexposed pixels,
\[
M_{LH}(x,y)=
\begin{cases}
1 & \text{if } \tau_L\le I(x,y)\le \tau_H,\\
0 & \text{otherwise},
\end{cases}
\qquad
I_p(x,y)=I(x,y)\,M_{LH}(x,y).
\]
Gradients on $I_p$ are filtered by percentile masks $M_{P_x},M_{P_y}$, typically in the $98.5$th to $99.5$th percentiles, to isolate pronounced edges.

The core directional score is based on the change in edge gradients after controlled Gaussian blur. With Gaussian kernel $G_\sigma$ of size $5\times 5$ and $\sigma=1$,
\[
B=I_p*G_\sigma.
\]
Using the same masks, the normalized decays are
\[
\Delta G_x=\frac{G_{x,\mathrm f}-G_{B_x}}{G_{x,\mathrm f}+\epsilon},
\qquad
\Delta G_y=\frac{G_{y,\mathrm f}-G_{B_y}}{G_{y,\mathrm f}+\epsilon},
\]
and the directional sharpness scores in the paper’s percent scale are
\[
S_x=100\times \frac1n\sum_{i=1}^n \Delta G_{x,i},
\qquad
S_y=100\times \frac1n\sum_{i=1}^n \Delta G_{y,i}.
\]
The normalized scores are
\[
\mathrm{NDS}_x=\frac1n\sum_{i=1}^n \Delta G_{x,i},
\qquad
\mathrm{NDS}_y=\frac1n\sum_{i=1}^n \Delta G_{y,i},
\]
so $S_x=100\times \mathrm{NDS}_x$ and $S_y=100\times \mathrm{NDS}_y$ [2410.10488].

The normalization divides decay by the original edge gradient and therefore acts as a local contrast normalization. The paper also defines a representativeness indicator using a heavier blur with $\sigma_R\approx 5$,
\[
B^{(R)}=I*G_{\sigma_R},
\]
\[
R_x=\frac1{N_x}\sum \left(\nabla_x(B^{(R)}\cdot M_{LH})\cdot M_{P_x}\right),
\qquad
R_y=\frac1{N_y}\sum \left(\nabla_y(B^{(R)}\cdot M_{LH})\cdot M_{P_y}\right),
\]
to filter images lacking sufficient usable edges. The intended application is constellation-wide quality monitoring, and the directional separation $(\mathrm{NDS}_x,\mathrm{NDS}_y)$ supports diagnosis of motion blur versus defocus [2410.10488].

The paper further provides an analytical step-edge model. For a unit step of contrast $\Delta I$ blurred by a Gaussian PSF of width $\sigma_0$,
\[
E(d)=\frac{\Delta I}{2}\left[1+\mathrm{erf}\!\left(\frac{d}{\sqrt{2}\sigma_0}\right)\right],
\]
\[
g_n(d)=\frac{\Delta I}{\sqrt{2\pi}\sigma_0}\exp\!\left(-\frac{d^2}{2\sigma_0^2}\right),
\qquad
g_n(0)=\frac{\Delta I}{\sqrt{2\pi}\sigma_0}.
\]
Applying additional Gaussian blur of width $\sigma_b$ changes the effective width to $\sigma'=\sqrt{\sigma_0^2+\sigma_b^2}$, so the peak normalized decay becomes
\[
\Delta
=
1-\frac{g_n'(0)}{g_n(0)}
=
1-\frac{\sigma_0}{\sqrt{\sigma_0^2+\sigma_b^2}}.
\]
Sharper edges, corresponding to smaller $\sigma_0$, therefore exhibit larger normalized decay upon additional blur [2410.10488].

## 7. Conceptual relations, distinctions, and limitations

Across these papers, NDS is consistently a directional quantity, but the direction and the normalization target are domain-specific. In momentum EoSS analysis, the direction is the stochastic mini-batch gradient and the normalization removes $1/\eta$ scaling. In Muon-versus-Adam analysis, the direction is the optimizer’s actual update and the normalization divides out step magnitude. In certification, the direction is induced by a short SAM or ASAM trajectory, and normalization divides out per-step loss scale before taking a temporal fluctuation statistic. In satellite imagery, the directions are pronounced edge normals or their axis-aligned approximations, and normalization divides gradient decay by original edge strength. This suggests that NDS is better understood as a design pattern for directional sharpness than as a single canonical scalar.

Several distinctions are technically important. First, directional sharpness is not the same as $\lambda_{\max}(H)$. The EoSS paper explicitly shows that $\lambda_{\max}$ may fail to diagnose mini-batch instability transitions, whereas Batch Sharpness tracks catapult events [2604.14108]. Second, the optimizer-comparison NDS of [2606.04662] is not a dynamic generalization certificate; it is a local Rayleigh quotient tied to a realized update. Third, the certification NDS of [2606.25004] is not a Hessian quadratic form at all, but a normalized fluctuation statistic over short SAM-style dynamics. Fourth, the image-quality NDS of [2410.10488] is unrelated to learning dynamics despite sharing the same directional-normalized motif.

The limitations are likewise definition-specific. The EoSS momentum analysis relies on linearization near a minimizer, quadratic approximation of per-sample losses, interpolation, and mean-square stability with random curvature; extending the analysis to very large models remains future work [2604.14108]. The Muon analysis studies local second-order behavior and stylized quadratic problems with heterogeneous curvature; its empirical results are reported for NanoGPT-scale models and controlled Zipf-PCFG settings [2606.04662]. The certification framework gives a one-way stability-to-bounded-sharpness implication and requires threshold calibration for binary certificates, while dataset forgery is out of scope [2606.25004]. The satellite metric can become unreliable under sparse edges, cloud cover, near-Nyquist texture, extreme noise, heavy compression, or spatially varying blur [2410.10488].

Taken together, these works establish NDS as a versatile but non-unified concept. In each case, it is the normalization of a directionally meaningful sharpness signal that makes the resulting quantity operational: learning-rate invariant for EoSS boundaries, step-size invariant for optimizer-local curvature penalties, loss-scale stabilized for dynamic generalization auditing, and contrast-normalized for no-reference image sharpness assessment.

Source: https://www.emergentmind.com/topics/normalized-directional-sharpness-nds