---
title: Scale-Invariant Self-Normalized Bounds
url: https://www.emergentmind.com/topics/scale-invariant-self-normalized-bounds
type: topic
---

# Scale-Invariant Self-Normalized Bounds

Scale-invariant self-normalized bounds are estimates whose controlling quantity is unchanged under the natural rescaling of the problem. Across probability, statistics, PDE, online learning, bandits, and neural-network generalization, the common mechanism is to normalize by an observed quadratic form, empirical norm, information divergence, or critical function-space norm, so that the resulting statistic or bound no longer depends on an external scale parameter. Canonical examples include \(S_k/V_n(\beta)\) with \(V_n(\beta)=\big(\sum_{i=1}^n |\xi_i|^\beta\big)^{1/\beta}\), \(S_n/\sqrt{[S]_n}\), \(N(n)I(X(n);\mu)\), \(\|(\lambda I+V_t)^{-1/2}M_t\|\), and elliptic estimates whose constants depend only on scale-critical Lorentz norms such as \(\|b-c\|_{L^{n,1}}\) [1611.08436][1712.03667][1309.3376][2511.03606][1904.04770].

## 1. Canonical forms of self-normalization

In the probabilistic literature, self-normalization usually means dividing a partial sum by a random norm built from the same sample. For independent, symmetric, nondegenerate variables \((\xi_i)_{i=1}^n\), one central construction is
\[
V_n(\beta)=\Big(\sum_{i=1}^n |\xi_i|^\beta\Big)^{1/\beta},\qquad
\max_{1\le k\le n}\frac{S_k}{V_n(\beta)},\qquad S_k=\sum_{i=1}^k \xi_i,
\]
with \(\beta>1\). If \(\xi_i' = c\xi_i\), then \(S_k' = cS_k\) and \(V_n'(\beta)=|c|V_n(\beta)\), so the upper-tail law of the self-normalized statistic is unchanged under global rescaling [1611.08436]. This is the basic finite-dimensional model of scale invariance.

A closely related Student-type normalization is
\[
T_n=\frac{S_n}{V_n},\qquad V_n=\sqrt{\sum_{j=1}^n X_j^2},
\]
for i.i.d. \(X_j\) with \(E X_1=0\) and \(E X_1^2=1\). The same ratio appears in asymptotic expansions, local limit theorems, and entropic CLTs for self-normalized sums, and remains invariant under \(X_j\mapsto \sigma X_j\) because numerator and denominator scale by the same factor [2207.14402]. In martingale form, the analogous statistics are
\[
\frac{S_n}{\sqrt{[S]_n}},\qquad \frac{S_n}{\sqrt{(S)_n}},
\]
where \([S]_n=\sum_{i=1}^n X_i^2\) and \((S)_n=\sum_{i=1}^n E(X_i^2\mid\mathcal F_{i-1})\). These are invariant under \(X_i\mapsto cX_i\) for the same reason [1712.03667].

A distinct but related construction replaces quadratic normalization by information normalization. In a filtered setting with empirical mean \(X_t=S_t/t\), the self-normalized quantity is
\[
t\,I(X_t;\mu),
\]
where \(I(\cdot;\mu)\) is the rate function defined from the conditional log-mgf. With a random number of observations,
\[
S(n)=\sum_{t=1}^n \varepsilon_t X_t,\qquad N(n)=\sum_{t=1}^n \varepsilon_t,\qquad X(n)=\frac{S(n)}{N(n)},
\]
the relevant statistic becomes \(N(n)\,I(X(n);\mu)\) [1309.3376]. Here the normalization is internal to the deviation functional itself rather than an explicit denominator.

These constructions have a common formal pattern: the numerator is a cumulative fluctuation, while the denominator or rate functional is computed from the same path. This suggests that “self-normalized” is best understood as a structural principle rather than a single formula.

## 2. Non-asymptotic deviation bounds and Gaussian approximation

For independent symmetric variables, the basic maximal deviation inequality is explicit. If \(0<x\le n^{(\beta-1)/\beta}\), then
\[
\mathbf P\!\left(\max_{1\le k\le n}\frac{S_k}{V_n(\beta)}\ge x\right)\le B_n(\beta,x),
\]
where
\[
B_n(\beta,x)=\exp\Big\{-\lambda x + n\,g\!\Big(\frac{\lambda}{n^{1/\beta}}\Big)\Big\},\qquad
g(u)=\log(\cosh u),
\]
with
\[
\lambda=\frac{\beta}{\beta-1}\log\frac{n^{(\beta-1)/\beta}+x}{n^{(\beta-1)/\beta}-x}.
\]
The range \(x\le n^{(\beta-1)/\beta}\) is optimal, and for \(\beta\in(1,2]\) this yields the simpler sub-Gaussian-type bound
\[
\mathbf P\!\left(\max_{1\le k\le n}\frac{S_k}{V_n(\beta)}\ge x\right)\le
\exp\!\left\{-\frac{x^2}{2}\,\beta^{1/(1-\beta)}\right\},
\]
which for \(\beta=2\) becomes \(\exp(-x^2/2)\) [1611.08436].

For martingale differences, the Berry–Esseen theory has the same self-normalized flavor. If
\[
N_n=\sum_{i=1}^n E|X_i|^{2p}+E|(S)_n-1|^p,\qquad p>1,
\]
then
\[
\sup_{x\in\mathbb R}\left|\mathbf P\!\left(\frac{S_n}{\sqrt{[S]_n}}\le x\right)-\Phi(x)\right|
\le C_p\,N_n^{1/(2p+1)},
\]
and the same rate holds for \(S_n/\sqrt{(S)_n}\). The paper also proves optimality: there exist martingale difference sequences for which the order \(N_n^{1/(2p+1)}\) cannot be improved [1712.03667].

At the level of refined asymptotics, self-normalized sums admit non-uniform Edgeworth expansions. If \(X_1,\dots,X_n\) are symmetric, non-singular, and \(E|X_1|^s<\infty\) for some real \(s\ge 2\), with \(m=\lfloor s\rfloor\), then
\[
\sup_{x\in\mathbb R}(1+|x|)^m\,|F_n(x)-\Phi^Q_{m,n}(x)|
=
O\Big(n^{-(s-2)/2}(\log n)^{(s+2m)/2}\Big),
\]
where \(\Phi^Q_{m,n}(x)=\Phi(x)+\sum_{r=1}^{m-2}Q_r(x)n^{-r/2}\) is the Edgeworth approximation for the cdf of \(T_n=S_n/V_n\). Parallel weighted bounds hold for the density and feed into entropic CLTs and total variation convergence [2207.14402].

A high-dimensional analogue considers the coordinate-wise maximum of self-normalized statistics,
\[
\|T_n\|_\infty
=
\max_{1\le j\le d}
\frac{\left|\sum_{i=1}^n e_j^\top X_i\right|}
{\sqrt{\sum_{i=1}^n (e_j^\top X_i)^2}}.
\]
With finite third absolute moment, the Gaussian approximation error satisfies
\[
\Delta_n \le C\,\log^{5/4}(e d)\,n^{-1/8}
\left(\mathbb E\max_{1\le j\le d}|e_j^\top X_1/\sigma_j|^3\right)^{1/4},
\]
so \(\Delta_n\to 0\) as long as \(\log(d)=o(n^{1/10})\) [2501.08979]. The dependence on standardized coordinates rather than raw \(\sigma_j\) again encodes scale invariance.

## 3. Information-normalized and vector-valued concentration in sequential problems

Information-normalized bounds provide uniform-in-time confidence sequences without explicit variance normalization. In the martingale/exponential-moment setting,
\[
\mathbb P\big(\exists t\in\{1,\dots,n\}: t\,I(X_t;\mu)\ge \delta\big)
\le 2e\,\lceil \delta\log n\rceil e^{-\delta},
\]
and for a random sample size \(N(n)\),
\[
\mathbb P\big(I(X(n);\mu)\ge \delta/N(n)\big)
\le 2e\,\lceil \delta\log n\rceil e^{-\delta}.
\]
For bounded \([0,1]\)-valued variables the rate function becomes the Bernoulli KL divergence \(\mathrm{kl}(x,\mu)\), giving confidence sets of the form
\[
\{\mu\in[0,1]: \mathrm{kl}(X_t,\mu)\le \delta/t\},
\]
which directly underlie KL-UCB-type indices in stochastic bandits [1309.3376].

In Hilbert-space form, the self-normalized statistic is
\[
S_t=\big\|(\lambda I + V_t)^{-1/2}M_t\big\|,\qquad
V_t=\sum_{i<t} X_iX_i^\top,\qquad
M_t=\sum_{i<t} X_i\varepsilon_i.
\]
For light tails beyond sub-Gaussianity, a supermartingale construction yields time-uniform bounds of the form
\[
\Pr\Big(\forall t\ge 1:\ \|(\lambda I+V_t)^{-1/2}M_t\|\le J_t(\delta)\Big)\ge 1-\delta,
\]
with \(J_t(\delta)\) derived under Bernstein or Bennett assumptions and depending on \(\log\det(I+\lambda^{-1}V_t)\) or the corresponding information gain \(\gamma_t(\lambda)\) [2511.03606]. Here the normalization is matrix-valued and geometric: the empirical Gram matrix determines both the scale and the shape of the confidence region.

This perspective becomes essential in generalized kernelized bandits. For bounded martingale noise \(\epsilon_t\) with predictable heteroscedastic variance \(\sigma_t^2\), the paper introduces
\[
S_t=\sum_{s=1}^{t-1}\epsilon_s\phi(x_s),\qquad
\mathcal H_t(\lambda)=\sum_{s=1}^{t-1}\sigma_s^2\phi(x_s)\phi(x_s)^\top+\lambda I,
\]
and proves a self-normalized Bernstein-like dimension-free inequality for
\[
\|S_t\|_{\mathcal H_t(\lambda)^{-1}}.
\]
This yields regret \(\widetilde O(\gamma_T\sqrt{T/\kappa_*})\) for GKB-UCB, matching, up to multiplicative constants and logarithmic terms, the state-of-the-art bounds for both kernelized bandits and generalized linear bandits [2508.01681]. A plausible implication is that variance-weighted self-normalization is the right replacement for Hoeffding-type normalization whenever the reward model is heteroscedastic.

## 4. Elliptic PDEs and critical-space normalization

In elliptic theory, scale-invariant self-normalized bounds refer to estimates whose constants depend only on norms that remain unchanged under the natural PDE scaling. The model operator is
\[
\mathcal{L}u=-\operatorname{div}(A\nabla u + b u)+ c\cdot \nabla u + d u
\quad\text{in }\Omega\subset\mathbb R^n,\ n\ge 3,
\]
with \(A\) measurable, bounded, and uniformly elliptic, \(b,c\in L^{n,q}(\Omega)\) for some \(q<\infty\), \(d\in L^{n/2,\infty}(\Omega)\), and either \(d\ge \operatorname{div} b\) or \(d\ge \operatorname{div} c\) in distributions. The decisive critical hypothesis is
\[
b-c\in L^{n,1}(\Omega),
\]
because under the elliptic scaling \(x\mapsto x_0+r x\),
\[
\|b_r-c_r\|_{L^{n,1}(\Omega_r)}=\|b-c\|_{L^{n,1}(\Omega)}.
\]
In this precise sense, \(L^{n,1}\) is the scale-critical Lorentz space for first-order terms [1904.04770].

Under these assumptions there exist nonnegative Green’s functions \(G(x,y)\) for \(\mathcal L\) and \(g(x,y)\) for \(\mathcal L^t\) such that, for all \(x\neq y\),
\[
0\le G(x,y)\le C'|x-y|^{2-n},\qquad
0\le g(x,y)\le C'|x-y|^{2-n},
\]
and for almost every fixed \(y\),
\[
\|G(\cdot,y)\|_{L^{\frac n{n-2},\infty}(\Omega)}
+
\|\nabla_x G(\cdot,y)\|_{L^{\frac n{n-1},\infty}(\Omega)}
\le C,
\]
with analogous bounds for \(g\). The constants satisfy
\[
C=C\big(n,\lambda,\|b-c\|_{L^{n,1}(\Omega)}\big),\qquad
C'=C'\big(n,\lambda,\|A\|_\infty,\|b-c\|_{L^{n,1}(\Omega)}\big),
\]
and are independent of the domain size, including \(\mathbb R^n\) and half-spaces [1904.04770].

The scaling is explicit. If
\[
\widetilde G(\widetilde x,\widetilde y)=r^{n-2}G(x_0+r\widetilde x,x_0+r\widetilde y),
\]
then the bound \(G(x,y)\le C'|x-y|^{2-n}\) becomes
\[
\widetilde G(\widetilde x,\widetilde y)\le C'|\widetilde x-\widetilde y|^{2-n},
\]
with the same constant. The weak-\(L^{n/(n-2)}\) and weak-\(L^{n/(n-1)}\) norms are likewise invariant at the critical exponents. The paper states that, in this sense, the bounds are self-normalized: changing the spatial scale does not change the size of the constants [1904.04770].

The same Green’s function bounds yield scale-invariant maximum principles for subsolutions of
\[
\mathcal L u \le -\operatorname{div}f+g.
\]
For finite \(|\Omega|\), if \(d\ge \operatorname{div} c\), \(f\in L^{n,1}(\Omega)\), \(g\in L^{n/2,1}(\Omega)\), and \(u\in W^{1,2}(\Omega)\) is a subsolution, then
\[
\sup_\Omega u
\le
C\left(\sup_{\partial\Omega}u^+ + \|f\|_{L^{n,1}(\Omega)} + \|g\|_{L^{n/2,1}(\Omega)}\right).
\]
A local version on \(B_r\) gives
\[
\sup_{B_{r/2}} |u|
\le
C\left(\fint_{B_r}|u|+\|f\|_{L^{n,1}(B_r)}+\|g\|_{L^{n/2,1}(B_r)}\right),
\]
with the same scaling invariance [1904.04770].

## 5. Learning-theoretic and optimization viewpoints

In online learning, scale invariance is enforced algorithmically. The normalized online learning algorithms NG and NAG are designed to be independent of feature scales, with regret bounds dependent on the ratio of scales existent in the data rather than the absolute scale [1305.6646]. The setup introduces a scaling adversary that chooses a positive definite matrix \(S\) and restricts inputs by \(\|S^{1/2}x_t\|_p\le 1\). The comparator class is
\[
\mathcal W_X^C=\{w\in\mathbb R^d:\|S^{-1/2}w\|_q\le C\},\qquad \frac1p+\frac1q=1.
\]
For the minimax optimal diagonal conditioner in hindsight, the regret takes the scale-invariant form
\[
R_T \le C\sum_{i=1}^d \sqrt{S_{ii}\sum_{t=1}^T g_{t,i}^2},
\]
which is unchanged under coordinatewise rescaling because \(S_{ii}\) and \(g_{t,i}\) transform inversely [1305.6646]. In the streaming case the only additional dependence is through scale ratios \(\Delta_i=\max_t|x_{t,i}|/|x_{t_0^i,i}|\).

In PAC-Bayesian analyses of neural networks, the same idea appears as normalization by parameter geometry. “Normalized flat minima” introduces the objective
\[
\min_{\sigma,\sigma'} \sum_{i,j}
\left(
h_{i,j}(\sigma_i\sigma'_j)^2
+
\frac{W[i,j]^2}{2\lambda(\sigma_i\sigma'_j)^2}
\right),
\]
where \(h_{i,j}\) are Hessian-diagonal terms and the variances factor into row and column scales. The minimized value is invariant under row–column rescalings that preserve the ReLU network function, and serves as a scale-invariant sharpness term in the PAC-Bayes bound [1901.04653].

A related construction replaces raw parameters by “connectivity” coordinates \(c\) defined through
\[
\psi=\theta^*+\theta^*\odot c.
\]
With an isotropic Gaussian prior on \(c\),
\[
\mathbb P_{\theta^*}(c)=\mathcal N(c\mid 0,\alpha^2 I_P),
\]
and a Gaussian posterior from Bayesian linear regression in the linearized model, the resulting PAC-Bayes-CTK bound depends on the eigenvalues of the Connectivity Tangent Kernel
\[
\mathbf C_{\mathcal X}^{\theta^*}
=
\mathbf J_\theta(\mathcal X,\theta^*)\operatorname{diag}(\theta^*)^2
\mathbf J_\theta(\mathcal X,\theta^*)^\top.
\]
Because \(\mathbf J_\theta(x,\mathcal T(\psi))\operatorname{diag}(\mathcal T(\psi))
=
\mathbf J_\theta(x,\psi)\operatorname{diag}(\psi)\) for function-preserving diagonal scalings \(\mathcal T\), the prior, posterior, CTK, and KL term are invariant under those scalings [2209.15208].

Self-normalization also appears in conditional exponential families. In self-normalized log-linear models one constrains or penalizes the variance of the log-normalizer \(A(X,\eta)\), and the likelihood gap satisfies
\[
\Delta_\ell(\hat\eta,\hat\eta_\delta)
\le
\left(1-\frac{\delta}{R\|\hat\eta\|_2}\right)
\mathbb E_X \mathrm{KL}(p_{\hat\eta}(\cdot\mid X)\,\|\,\mathrm{Unif}),
\]
so the controlling quantity is the normalized tolerance \(\delta/(R\|\hat\eta\|_2)\), which is invariant under inverse rescaling of features and parameters that preserves \(p_\eta(y\mid x)\) [1506.04147].

At the optimization level, scale-invariant first-order methods based on linear minimization oracles explicitly normalize directions by the geometry of a norm ball. In matrix optimization with heavy-tailed noise and spectral norm geometry, the paper proves that any scale-invariant first-order method requires
\[
\Omega\!\left(\min\{m,n\}\,\epsilon^{-\frac{3p-2}{p-1}}\right)
\]
oracle calls, and that a batched Scion method with spectral norm achieves the matching upper bound
\[
O\!\left(\min\{m,n\}\,\epsilon^{-\frac{3p-2}{p-1}}\right).
\]
With Hessian Lipschitzness, a transported Scion method improves this to
\[
O\!\left(\min\{m,n\}\,\epsilon^{-\frac{5p-3}{2p-2}}\right)
\]
[2605.18528]. The controlling constant is a self-normalized martingale factor
\[
\tau(\|\cdot\|_\star,m,n,p)
=
\sup \frac{\mathbb E\big\|\sum_{t=1}^T Z_t\big\|_\star}
{\mathbb E\big(\sum_{t=1}^T \|Z_t\|_\star^p\big)^{1/p}},
\]
which is dimension-free for Frobenius geometry but scales as \(\Theta(\min\{m,n\}^{1-1/p})\) for the nuclear norm, making the price of scale invariance geometry dependent.

## 6. Sharpness, impossibility, and structural conditions

Several literatures show that scale-invariant self-normalization is a sharp property rather than an automatic consequence of dividing by a random quantity. In elliptic PDE, the condition
\[
b-c\in L^{n,1}(\Omega)
\]
is both necessary and optimal for scale-invariant Green’s function bounds. The counterexample
\[
c(x)=-\frac{x}{|x|^2\ln|x|}
\]
lies in \(L^{n,q}\) for every \(q>1\) but not in \(L^{n,1}\), and for operators such as \(-\Delta+\delta c\cdot\nabla\) the Green’s function fails to satisfy the pointwise and weak-type bounds uniformly in the pole even when \(\delta\) is arbitrarily small [1904.04770]. The paper’s conclusion is that the Lorentz refinement, not merely scale invariance of the Lebesgue exponent \(L^n\), is essential.

In probability, self-normalization does not eliminate substantive assumptions. The maximal inequality for \(S_k/V_n(\beta)\) requires independent, symmetric, nondegenerate variables and \(\beta>1\) [1611.08436]. The non-uniform Edgeworth theory for \(T_n=S_n/V_n\) requires symmetry, non-singularity, and moment assumptions such as \(E|X_1|^s<\infty\) [2207.14402]. The Berry–Esseen bound for self-normalized martingales requires \(E|X_i|^{2p}<\infty\) for some \(p>1\) and involves the deviation of the predictable quadratic variation from \(1\) through
\[
N_n=\sum_{i=1}^n E|X_i|^{2p}+E|(S)_n-1|^p
\]
[1712.03667]. In high dimension, the best current Berry–Esseen bound for the coordinate-wise maximum of self-normalized sums is slower than the classical \(n^{-1/2}\) rate and requires \(\log(d)=o(n^{1/10})\) [2501.08979].

The strongest impossibility result concerns vector-valued self-normalized martingales. For
\[
R_T=\|S_T\|_{V_T^\dagger}^2,\qquad
S_T=\sum_{i=1}^T Y_iX_i,\qquad
V_T=\sum_{i=1}^T X_iX_i^\top,
\]
the paper proves that nontrivial scale-invariant upper bounds exist only in dimension \(d=1\) without further assumptions. In \(d=1\), there is an \(O(\log T)\) scale-invariant bound for dyadic martingales and an explicit algorithm with \(O(\log T)\) doubly-uniform regret in online linear regression. In contrast, for \(d>1\), no nontrivial scale-invariant bound can hold in full generality, and sublinear doubly-uniform regret is impossible [2605.01628].

That same paper also identifies a recovery mechanism. Under a smoothness condition requiring bounded Radon–Nikodym derivatives of the conditional covariate laws with respect to a fixed base measure, VAW achieves
\[
Reg(T)\lesssim \sqrt{d\,T\,\log(T/\delta)}+\log(1/\delta),
\]
and the self-normalized martingale satisfies
\[
\|S_T\|_{V_T^\dagger}^2
\lesssim
\sigma^2\left(\sqrt{d\,T\,\log(2T/\delta)}+\log(2/\delta)\right)
\]
with high probability [2605.01628]. This suggests that scale-invariant self-normalized control in higher dimensions requires structural randomness or smoothness in the covariates, not merely algebraic normalization.

Taken together, these results delineate a broad principle. Self-normalization removes nuisance scale only when the normalizer matches the intrinsic geometry of the problem: a critical Lorentz norm in elliptic PDE, an empirical quadratic form or information divergence in probability, a variance-weighted Gram matrix in bandits, or an invariant parameterization in learning and optimization. Where such geometry is absent or incomplete, sharp counterexamples show that scale invariance can fail despite formal normalization.

Source: https://www.emergentmind.com/topics/scale-invariant-self-normalized-bounds