Papers
Topics
Authors
Recent
Search
2000 character limit reached

Diffusion Information Exponent

Updated 4 July 2026
  • Diffusion information exponent is a scalar invariant defined via Hermite expansions that marks the first statistical feature exploited by a denoiser-loss pair.
  • It governs learning progression by indicating when models recover basic pair-wise statistics before uncovering higher-order data correlations.
  • Empirical and synthetic analyses show that lower Hermite orders correlate with quasi-linear sample complexity, while higher orders require significantly more data.

The diffusion information exponent is a scalar invariant introduced in the theory of denoising diffusion learning to characterize the lowest-order statistical feature of the data that is both present in the noised distribution and exploitable by the denoiser–loss pair. Denoted kk^\star, it is defined through Hermite expansions of the noised-data likelihood ratio and of an effective nonlinearity induced by the diffusion objective. In the online high-dimensional regime, it governs the power of the ambient dimension appearing in sample complexity, and thereby formalizes an “easy-to-hard” learning progression in which diffusion models first recover pair-wise statistics and only later learn higher-order correlations (Bardone et al., 13 Mar 2026).

1. Definition and conceptual status

In the paper that introduces the term, the diffusion information exponent is defined from two Hermite expansions. The first is the likelihood ratio of the noised data,

Lt:=dPtdN(0,1d),L_t:=\frac{\mathrm d P_t}{\mathrm d \mathcal N(0,1_d)},

and the second is the effective nonlinearity

Fσ(xw):=σ(xw)σ(xw)σ(xw)σ(xw)σ(xw)xw.F_\sigma(x\cdot w):= \sigma''(x\cdot w)-\sigma'(x\cdot w)\sigma(x\cdot w)- \sigma(x\cdot w)- \sigma'(x\cdot w)x\cdot w.

If

Lt(xv)=i=0ciLhi(xv),Fσ(xw)=j=0cjFhj(xw),L_t(x\cdot v)=\sum_{i=0}^\infty c_i^L h_i(x_v), \qquad F_\sigma(x_w)=\sum_{j=0}^\infty c_j^F h_j(x_w),

then the diffusion information exponent is

k=min{k1:  ckL0 and ck1F0}.k^\star = \min\{k\ge 1:\; c_k^L\neq 0 \ \text{and}\ c_{k-1}^F\neq 0\}.

Equivalently, it is the smallest Hermite order at which the data distribution and the denoiser-loss geometry interact nontrivially (Bardone et al., 13 Mar 2026).

The same work describes kk^\star as “a scalar invariant of the loss.” More precisely, it is an invariant of the inference problem induced by diffusion training: it depends on the noised data distribution through the coefficients ckLc_k^L, and on the architecture–loss combination through the coefficients cjFc_j^F. The paper states that kk^\star “identifies the lowest-order statistical feature of the data that is both present in the distribution and exploitable by the combination of the mean-squared error loss and the nonlinearity σ\sigma” (Bardone et al., 13 Mar 2026).

The appendix places this notion in a wider family of Hermite-based invariants. It recalls the standard information exponent of a square-integrable function Lt:=dPtdN(0,1d),L_t:=\frac{\mathrm d P_t}{\mathrm d \mathcal N(0,1_d)},0, defined as the smallest Lt:=dPtdN(0,1d),L_t:=\frac{\mathrm d P_t}{\mathrm d \mathcal N(0,1_d)},1 such that the Lt:=dPtdN(0,1d),L_t:=\frac{\mathrm d P_t}{\mathrm d \mathcal N(0,1_d)},2-th Hermite coefficient of Lt:=dPtdN(0,1d),L_t:=\frac{\mathrm d P_t}{\mathrm d \mathcal N(0,1_d)},3 is nonzero. The diffusion information exponent is presented as the diffusion-score-learning analogue of that notion, and the paper explicitly relates it by analogy to the “generative exponent” and the “leap index” (Bardone et al., 13 Mar 2026).

2. Analytical setting in denoising diffusion

The definition arises in a specific denoising diffusion framework. The forward diffusion process is

Lt:=dPtdN(0,1d),L_t:=\frac{\mathrm d P_t}{\mathrm d \mathcal N(0,1_d)},4

with distributional solution

Lt:=dPtdN(0,1d),L_t:=\frac{\mathrm d P_t}{\mathrm d \mathcal N(0,1_d)},5

where Lt:=dPtdN(0,1d),L_t:=\frac{\mathrm d P_t}{\mathrm d \mathcal N(0,1_d)},6 and Lt:=dPtdN(0,1d),L_t:=\frac{\mathrm d P_t}{\mathrm d \mathcal N(0,1_d)},7. The score of the noised density Lt:=dPtdN(0,1d),L_t:=\frac{\mathrm d P_t}{\mathrm d \mathcal N(0,1_d)},8 is written via Tweedie’s formula as

Lt:=dPtdN(0,1d),L_t:=\frac{\mathrm d P_t}{\mathrm d \mathcal N(0,1_d)},9

Training uses the diffusion objective

Fσ(xw):=σ(xw)σ(xw)σ(xw)σ(xw)σ(xw)xw.F_\sigma(x\cdot w):= \sigma''(x\cdot w)-\sigma'(x\cdot w)\sigma(x\cdot w)- \sigma(x\cdot w)- \sigma'(x\cdot w)x\cdot w.0

for a simple denoiser

Fσ(xw):=σ(xw)σ(xw)σ(xw)σ(xw)σ(xw)xw.F_\sigma(x\cdot w):= \sigma''(x\cdot w)-\sigma'(x\cdot w)\sigma(x\cdot w)- \sigma(x\cdot w)- \sigma'(x\cdot w)x\cdot w.1

Projected SGD is performed on the sphere,

Fσ(xw):=σ(xw)σ(xw)σ(xw)σ(xw)σ(xw)xw.F_\sigma(x\cdot w):= \sigma''(x\cdot w)-\sigma'(x\cdot w)\sigma(x\cdot w)- \sigma(x\cdot w)- \sigma'(x\cdot w)x\cdot w.2

with spherical gradient

Fσ(xw):=σ(xw)σ(xw)σ(xw)σ(xw)σ(xw)xw.F_\sigma(x\cdot w):= \sigma''(x\cdot w)-\sigma'(x\cdot w)\sigma(x\cdot w)- \sigma(x\cdot w)- \sigma'(x\cdot w)x\cdot w.3

After integrating out the Gaussian diffusion noise with Stein’s lemma, the sample-wise spherical gradient takes the form

Fσ(xw):=σ(xw)σ(xw)σ(xw)σ(xw)σ(xw)xw.F_\sigma(x\cdot w):= \sigma''(x\cdot w)-\sigma'(x\cdot w)\sigma(x\cdot w)- \sigma(x\cdot w)- \sigma'(x\cdot w)x\cdot w.4

so the entire low-order learning analysis is reduced to the interaction between the Hermite structure of Fσ(xw):=σ(xw)σ(xw)σ(xw)σ(xw)σ(xw)xw.F_\sigma(x\cdot w):= \sigma''(x\cdot w)-\sigma'(x\cdot w)\sigma(x\cdot w)- \sigma(x\cdot w)- \sigma'(x\cdot w)x\cdot w.5 and that of Fσ(xw):=σ(xw)σ(xw)σ(xw)σ(xw)σ(xw)xw.F_\sigma(x\cdot w):= \sigma''(x\cdot w)-\sigma'(x\cdot w)\sigma(x\cdot w)- \sigma(x\cdot w)- \sigma'(x\cdot w)x\cdot w.6 (Bardone et al., 13 Mar 2026).

This setting explains why the invariant is defined through two expansions rather than one. In a diffusion objective, the data are not accessed directly; they are first convolved with Gaussian noise, producing Fσ(xw):=σ(xw)σ(xw)σ(xw)σ(xw)σ(xw)xw.F_\sigma(x\cdot w):= \sigma''(x\cdot w)-\sigma'(x\cdot w)\sigma(x\cdot w)- \sigma(x\cdot w)- \sigma'(x\cdot w)x\cdot w.7, and then filtered through the denoiser-induced effective nonlinearity. The exponent Fσ(xw):=σ(xw)σ(xw)σ(xw)σ(xw)σ(xw)xw.F_\sigma(x\cdot w):= \sigma''(x\cdot w)-\sigma'(x\cdot w)\sigma(x\cdot w)- \sigma(x\cdot w)- \sigma'(x\cdot w)x\cdot w.8 marks the first order at which these two objects align.

3. Early-time dynamics and sample-complexity scaling

The exponent is called an exponent because it determines the power of the ambient dimension Fσ(xw):=σ(xw)σ(xw)σ(xw)σ(xw)σ(xw)xw.F_\sigma(x\cdot w):= \sigma''(x\cdot w)-\sigma'(x\cdot w)\sigma(x\cdot w)- \sigma(x\cdot w)- \sigma'(x\cdot w)x\cdot w.9 appearing in the online sample complexity. In the single-spike setting analyzed in the paper, if Lt(xv)=i=0ciLhi(xv),Fσ(xw)=j=0cjFhj(xw),L_t(x\cdot v)=\sum_{i=0}^\infty c_i^L h_i(x_v), \qquad F_\sigma(x_w)=\sum_{j=0}^\infty c_j^F h_j(x_w),0 denotes the overlap with the planted direction, then early-time dynamics obey

Lt(xv)=i=0ciLhi(xv),Fσ(xw)=j=0cjFhj(xw),L_t(x\cdot v)=\sum_{i=0}^\infty c_i^L h_i(x_v), \qquad F_\sigma(x_w)=\sum_{j=0}^\infty c_j^F h_j(x_w),1

Starting from isotropic initialization, Lt(xv)=i=0ciLhi(xv),Fσ(xw)=j=0cjFhj(xw),L_t(x\cdot v)=\sum_{i=0}^\infty c_i^L h_i(x_v), \qquad F_\sigma(x_w)=\sum_{j=0}^\infty c_j^F h_j(x_w),2, so the first nonzero drift appears at order Lt(xv)=i=0ciLhi(xv),Fσ(xw)=j=0cjFhj(xw),L_t(x\cdot v)=\sum_{i=0}^\infty c_i^L h_i(x_v), \qquad F_\sigma(x_w)=\sum_{j=0}^\infty c_j^F h_j(x_w),3. The paper then shows that recovery requires on the order of

Lt(xv)=i=0ciLhi(xv),Fσ(xw)=j=0cjFhj(xw),L_t(x\cdot v)=\sum_{i=0}^\infty c_i^L h_i(x_v), \qquad F_\sigma(x_w)=\sum_{j=0}^\infty c_j^F h_j(x_w),4

samples or iterations, up to logarithmic factors (Bardone et al., 13 Mar 2026).

The positive result is formulated through a threshold Lt(xv)=i=0ciLhi(xv),Fσ(xw)=j=0cjFhj(xw),L_t(x\cdot v)=\sum_{i=0}^\infty c_i^L h_i(x_v), \qquad F_\sigma(x_w)=\sum_{j=0}^\infty c_j^F h_j(x_w),5: Lt(xv)=i=0ciLhi(xv),Fσ(xw)=j=0cjFhj(xw),L_t(x\cdot v)=\sum_{i=0}^\infty c_i^L h_i(x_v), \qquad F_\sigma(x_w)=\sum_{j=0}^\infty c_j^F h_j(x_w),6 Under suitable step-size conditions,

Lt(xv)=i=0ciLhi(xv),Fσ(xw)=j=0cjFhj(xw),L_t(x\cdot v)=\sum_{i=0}^\infty c_i^L h_i(x_v), \qquad F_\sigma(x_w)=\sum_{j=0}^\infty c_j^F h_j(x_w),7

projected gradient descent started from Lt(xv)=i=0ciLhi(xv),Fσ(xw)=j=0cjFhj(xw),L_t(x\cdot v)=\sum_{i=0}^\infty c_i^L h_i(x_v), \qquad F_\sigma(x_w)=\sum_{j=0}^\infty c_j^F h_j(x_w),8 achieves

Lt(xv)=i=0ciLhi(xv),Fσ(xw)=j=0cjFhj(xw),L_t(x\cdot v)=\sum_{i=0}^\infty c_i^L h_i(x_v), \qquad F_\sigma(x_w)=\sum_{j=0}^\infty c_j^F h_j(x_w),9

in probability and in k=min{k1:  ckL0 and ck1F0}.k^\star = \min\{k\ge 1:\; c_k^L\neq 0 \ \text{and}\ c_{k-1}^F\neq 0\}.0 for all k=min{k1:  ckL0 and ck1F0}.k^\star = \min\{k\ge 1:\; c_k^L\neq 0 \ \text{and}\ c_{k-1}^F\neq 0\}.1. A complementary negative result shows that if k=min{k1:  ckL0 and ck1F0}.k^\star = \min\{k\ge 1:\; c_k^L\neq 0 \ \text{and}\ c_{k-1}^F\neq 0\}.2, then weak recovery fails: k=min{k1:  ckL0 and ck1F0}.k^\star = \min\{k\ge 1:\; c_k^L\neq 0 \ \text{and}\ c_{k-1}^F\neq 0\}.3 in probability and in k=min{k1:  ckL0 and ck1F0}.k^\star = \min\{k\ge 1:\; c_k^L\neq 0 \ \text{and}\ c_{k-1}^F\neq 0\}.4 for every k=min{k1:  ckL0 and ck1F0}.k^\star = \min\{k\ge 1:\; c_k^L\neq 0 \ \text{and}\ c_{k-1}^F\neq 0\}.5 (Bardone et al., 13 Mar 2026).

These theorems make k=min{k1:  ckL0 and ck1F0}.k^\star = \min\{k\ge 1:\; c_k^L\neq 0 \ \text{and}\ c_{k-1}^F\neq 0\}.6 an explicit hardness index. A smaller exponent corresponds to lower-order statistical information that is visible to diffusion training at shorter sample scales. A larger exponent corresponds to information that is present only in higher Hermite modes and therefore becomes algorithmically visible later.

4. Mixed cumulant model and the “easy-to-hard” mechanism

The main solvable data model used to exhibit the exponent is the mixed cumulant model

k=min{k1:  ckL0 and ck1F0}.k^\star = \min\{k\ge 1:\; c_k^L\neq 0 \ \text{and}\ c_{k-1}^F\neq 0\}.7

where k=min{k1:  ckL0 and ck1F0}.k^\star = \min\{k\ge 1:\; c_k^L\neq 0 \ \text{and}\ c_{k-1}^F\neq 0\}.8 are planted unit vectors, k=min{k1:  ckL0 and ck1F0}.k^\star = \min\{k\ge 1:\; c_k^L\neq 0 \ \text{and}\ c_{k-1}^F\neq 0\}.9, kk^\star0, and

kk^\star1

In this model, kk^\star2 is a covariance spike and kk^\star3 is a cumulant spike. The construction is designed so that kk^\star4 is detectable from pair-wise statistics, whereas kk^\star5 is invisible to covariance and appears only through higher-order non-Gaussian structure (Bardone et al., 13 Mar 2026).

This separation yields the paper’s principal difficulty gap. In the pure cumulant-spike setting kk^\star6, the paper finds

kk^\star7

so the first informative Hermite order is the fourth one. It then states that recovery of the cumulant spike takes “a number of samples larger than cubic in the dimension” (Bardone et al., 13 Mar 2026).

By contrast, pair-wise statistics are learned at quasi-linear complexity. In Proposition kk^\star8, the covariance spike kk^\star9 is weakly recovered with

ckLc_k^L0

provided ckLc_k^L1 grows at most polynomially in ckLc_k^L2. This is the rigorous version of the statement that diffusion models learn simple pair-wise structure at linear sample complexity (Bardone et al., 13 Mar 2026).

The same proposition identifies an important exception to cubic hardness. If the latent variables are positively correlated,

ckLc_k^L3

and initialization satisfies

ckLc_k^L4

then weak recovery is achieved also for the cumulant spike ckLc_k^L5 within the same quasi-linear sample budget ckLc_k^L6. The paper presents this as a mechanism by which higher-order statistics can become easy when they share latent structure with pair-wise statistics (Bardone et al., 13 Mar 2026).

A concise summary of the model-specific regimes is useful:

Statistical feature Representative condition Sample-complexity statement
Pair-wise statistics Covariance spike ckLc_k^L7 present ckLc_k^L8
Isolated higher-order statistics ckLc_k^L9 Larger than cubic in cjFc_j^F0
Coupled higher-order statistics cjFc_j^F1 and matching-sign initialization Quasi-linear in cjFc_j^F2

The model thereby turns the diffusion information exponent into an explicit explanation of staged statistical learning: pair-wise structure corresponds to lower Hermite order and appears early, while uncoupled fourth-order structure appears late.

5. Empirical evidence in synthetic and natural-image diffusion models

The paper combines the solvable theory with empirical evidence from larger diffusion models. It reports that standard diffusion models trained on natural images exhibit a distributional simplicity bias, learning “simple, pair-wise input statistics before specializing to higher-order correlations” (Bardone et al., 13 Mar 2026).

On natural-image benchmarks, U-Net diffusion models are trained on grayscale CIFAR-10 and CelebA. Test loss is then compared across three data families: real data, Gaussian “clone” datasets with matching mean, and Gaussian clones with matching mean and covariance. The reported pattern is sequential specialization: early in training, performance on real images and covariance-matched Gaussian clones is similar, whereas only later does the model substantially outperform the covariance clone. The paper interprets this as evidence that covariance structure is learned before genuinely higher-order image statistics (Bardone et al., 13 Mar 2026).

The same pattern is reproduced in the mixed cumulant model. There, the covariance spike is learned first, and the cumulant spike is learned later unless latent correlations reduce its effective difficulty. The synthetic model therefore supplies a controlled realization of the same easy-to-hard phenomenon observed in natural-image training (Bardone et al., 13 Mar 2026).

The empirical role of the diffusion information exponent is therefore organizational rather than merely definitional. The experiments do not estimate cjFc_j^F3 directly, but they display the ordering of learning events that cjFc_j^F4 is designed to predict.

6. Relation to other uses of “diffusion exponent”

The exact phrase diffusion information exponent is specific to denoising diffusion learning. Other arXiv literatures use closely related words—diffusion exponent, information rate, variable exponent diffusion—for mathematically different objects.

Paper Exponent or rate Meaning
(Bardone et al., 13 Mar 2026) cjFc_j^F5 Hermite-order invariant controlling learning sample complexity
(Itto, 2017) cjFc_j^F6 Local anomalous-diffusion exponent in cjFc_j^F7
(Lanoiselée et al., 2024) cjFc_j^F8 Anomalous exponent estimated jointly with cjFc_j^F9 from TAMSD statistics
(Woszczek et al., 2024) kk^\star0 Random anomalous exponent in scaled Brownian motion
(Baranau, 8 Oct 2025) kk^\star1 or kk^\star2 KL-rate / entropy-growth or variance-growth rate in optimal nonlocal diffusion
(Avci, 23 Jul 2025) kk^\star3 State-dependent diffusion elasticity in a generalized CIR model

In anomalous-transport papers, the exponent usually controls mean-squared-displacement growth. For example, virus capsid motion is described by

kk^\star4

with local fluctuations of kk^\star5 treated statistically and shown to follow a Gaussian law under a maximum-entropy argument (Itto, 2017). In a recent estimator study, the same kk^\star6 is the anomalous diffusion exponent in

kk^\star7

and the central problem is the conditional joint law of kk^\star8 for finite trajectories (Lanoiselée et al., 2024). In scaled Brownian motion with random exponent, the trajectory-wise exponent becomes a random variable kk^\star9, and the ensemble MSD is no longer a pure power law but

σ\sigma0

(Woszczek et al., 2024).

A different information-theoretic usage appears in a nonlocal entropy-maximizing diffusion process, where the primitive control parameter is a fixed infinitesimal KL-rate σ\sigma1, yielding exact solutions with

σ\sigma2

That paper states that the phrase “Diffusion Information Exponent” is not defined there as a formal term; the closest quantities are σ\sigma3 for entropy or width growth and σ\sigma4 for variance growth (Baranau, 8 Oct 2025).

In stochastic-volatility theory, a state-dependent exponent σ\sigma5 in

σ\sigma6

controls the local elasticity of the diffusion coefficient with respect to the state, and thereby local volatility intensity, heteroscedasticity, and boundary behavior (Avci, 23 Jul 2025). This is again a distinct notion: the exponent carries information about the stochastic law itself, not about the learning complexity of a diffusion model.

The comparison shows that the learning-theoretic diffusion information exponent occupies a specific niche. It is neither a transport exponent nor a fractional order nor a state-dependent volatility elasticity. It is a Hermite-order invariant for score-based learning, introduced to quantify when a diffusion model can first access a given statistical feature of the data (Bardone et al., 13 Mar 2026).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Diffusion Information Exponent.