- The paper identifies spectral collapse as a failure mode in deterministic diffusion inversion, where spectrally sparse inputs produce correlated, low-dimensional latents that yield oversmoothed outputs.
- The paper shows that epsilon-prediction, network spectral bias, and low-SNR inversion dynamics drive collapse, while x0-prediction improves high-frequency recovery but can reduce structural fidelity.
- The paper introduces Orthogonal Variance Guidance, which restores latent variance while preserving low-frequency structure; EDM+OVG achieves LPIPS 0.381 and high-frequency fidelity of 0.90 on Edges2Shoes.
Overview
This paper identifies and characterizes a failure mode of deterministic diffusion inversion in unpaired image-to-image translation tasks where the source domain is spectrally sparse relative to the target domain, such as super-resolution and sketch-to-image translation. The authors show that when a source image lacks high-frequency content, the latent recovered by inverting the Probability Flow ODE (PF-ODE) does not approximate the isotropic Gaussian prior. Instead, it retains low-frequency structural imprints of the input with deficient high-frequency energy — a phenomenon the authors term spectral collapse. Because this collapsed latent lacks the stochastic variance needed to seed texture synthesis, conditional sampling from it produces oversmoothed, texture-poor outputs. The paper's central contribution is twofold: an empirical and theoretical account of why spectral collapse occurs, and an inference-time correction method, Orthogonal Variance Guidance (OVG), that restores Gaussian noise statistics without sacrificing structural fidelity to the input.
Background: inversion on the PF-ODE
The paper works within the standard framework of conditional diffusion models defined by a forward SDE whose marginals are shared by a deterministic PF-ODE. Inversion integrates this ODE forward in time to map an input x0 under source conditioning yˉ to a terminal latent xT, which is then sampled backward under target conditioning y. A wide class of deterministic samplers admits an affine update xt−1=a(t)xt+b(t)Fθ(xt,y,t), whose algebraic inversion yields the forward-time step. This inversion is exact only in the infinitesimal-step limit; finite steps implicitly assume the vector field is spatially constant across a step (the "linear assumption"). Violations of this assumption accumulate as trajectory drift, pushing xT away from the Gaussian prior.
Two parameterizations are contrasted: DDIM (ϵ-prediction under a VP schedule) and EDM (x0-prediction with preconditioning under a VE-style schedule). The choice of parameterization determines the schedules a(t) and b(t) and, critically, how errors amplify near yˉ0, where division by small signal scales magnifies local violations of the linear assumption.
Spectral collapse: empirical characterization
Using two unpaired benchmarks — Edges2Shoes (49k strictly unpaired images) and BBBC021 adapted for aggressive nearest-neighbor super-resolution at factors of ×8, ×16, and ×32 — the authors introduce three diagnostic metrics: a Gaussianity score yˉ1 based on a Kolmogorov–Smirnov test of squared latent norms against the yˉ2 distribution; a decorrelation score yˉ3 measuring the fraction of anomalously smooth patches in the latent relative to a Monte Carlo-calibrated null; and spectral fidelity scores for generated images, yˉ4 (wavelet-based high-frequency energy matched to dataset reference statistics) and yˉ5 (SSIM between Fourier low-passed output and condition).
The key observations are:
- Inverted latents from spectrally sparse inputs exhibit strong spatial patch correlations and occupy a markedly lower-dimensional subspace than isotropic Gaussians, as shown by PCA cumulative explained variance over 2000 latents; the effect intensifies with sparser inputs (×32 versus ×8).
- yˉ6 correlates strongly with both non-Gaussianity (yˉ7) and degraded high-frequency output (yˉ8), establishing a causal chain from latent statistics to texture quality.
- Stochastic inversion methods (TABA, ReNoise) restore latent independence (yˉ9) and texture but degrade structural fidelity, quantifying a structure-texture trade-off: stochastic methods improve xT0 over deterministic baselines in 93.8% of comparisons but improve xT1 in only 37.5%.
A notable finding concerns the training objective. Switching from xT2-prediction to xT3-prediction "effectively eliminates spectral collapse" — improving xT4 in 96.6% of trials and xT5 in 89.7% — but introduces structural drift, with xT6 and reconstruction metrics improving in only roughly 17–35% of trials. This contradicts the attribution in prior work (TABA), which located the problem in early inversion-step errors rather than the prediction target itself.
Theoretical analysis
The appendix provides a formal derivation grounded in the spectral bias of deep ReLU networks (Rahaman et al., 2018). Under xT7-prediction, the regression target for a spectrally sparse input is essentially white noise, whose flat spectrum the network cannot fit given its polynomial spectral decay xT8. The optimal estimator collapses to the conditional mean xT9, so the score vanishes and the PF-ODE reduces to linear decay, exponentially attenuating any microscopic high-frequency perturbation y0. Under y1-prediction, by contrast, the target aligns with the network's spectral bias; the score computed via the residual y2 recovers y3 analytically, producing a repulsive vector field that amplifies high-frequency variance toward full-rank Gaussianity.
Two corollaries sharpen the picture. First, spectral collapse does not preclude exact reconstruction: the collapsed map remains a bijective scalar contraction, so round-tripping succeeds even though the latent violates the prior hypothesis — explaining why reconstruction-focused fixes such as EDICT or BDIA cannot address the problem. Second, stability is asymmetric: inversion is stable if and only if the input already contains high-frequency variance, since only then is the noise target correlated with an input component the network can represent.
Orthogonal Variance Guidance
OVG augments each inversion step with two control drifts. A high-frequency loss y4 exploits concentration of measure to constrain the latent to the correct energy shell of radius y5, restoring the variance required for texture. A low-frequency loss penalizes drift of the denoised estimate y6 from a low-frequency reference derived from the conditioning input, preserving structure. When the corresponding gradients conflict (negative cosine similarity), each update is projected onto the normal plane of the other, in the spirit of PCGrad gradient surgery (Yu et al., 2020); when they agree, standard updates are retained. The corrected update adds these orthogonalized drifts to the base PF-ODE discretization. Step sizes y7 and y8 calibrate guidance strength per dataset and are selected by grid search; sweeping them expands the Pareto frontier of the y9–xt−1=a(t)xt+b(t)Fθ(xt,y,t)0 trade-off beyond what baselines achieve.
Experimental results
Across both datasets, deterministic baselines (DDIM, Null-Class, DirectInversion) exhibit collapsed latents (xt−1=a(t)xt+b(t)Fθ(xt,y,t)1–xt−1=a(t)xt+b(t)Fθ(xt,y,t)2) and poor texture (xt−1=a(t)xt+b(t)Fθ(xt,y,t)3 on Edges2Shoes). Stochastic methods restore latent statistics but sacrifice structure. The headline result is that EDM+OVG achieves the best joint trade-off: on Edges2Shoes it attains LPIPS 0.381 (best among all methods) with xt−1=a(t)xt+b(t)Fθ(xt,y,t)4 0.60 versus TABA's 0.46, while maintaining xt−1=a(t)xt+b(t)Fθ(xt,y,t)5 0.90; on BBBC021 ×16 it reaches MS-SSIM 0.665 and LPIPS 0.281 with FID 18.62, comparable realism to TABA (FID 12.06) at higher structural fidelity.
Ablations localize the phenomenon decisively. Replicating BBBC021 with a DiT-L/2 backbone reproduces identical collapse patterns under standard inversion and identical recovery under OVG, ruling out convolutional inductive bias. Controlled pixel-space experiments (removing the VAE bottleneck entirely) reproduce the same extreme spatial correlations, ruling out autoencoder compression. Stress-testing across downsampling factors shows OVG dominates at ×8 and ×16, while at the extreme ×32 regime all variance-restoring methods maintain realism (FID ≈ 15, xt−1=a(t)xt+b(t)Fθ(xt,y,t)6) at the cost of distortion metrics — standard inversion minimizes PSNR-style distortion but fails perceptually (FID 99.14). Velocity heatmaps of the update spectrum corroborate the mechanism: standard DDIM exhibits vanishing high-frequency update energy as SNR decays, whereas EDM maintains spectral continuity and OVG injects broadband orthogonal energy throughout the trajectory.
Limitations and open questions
Several caveats bear directly on the results. The theoretical analysis relies on the spectral decay bounds of ReLU networks trained by gradient descent, an idealization that may not strictly hold for pretrained, heavily regularized diffusion backbones at scale. The guidance step sizes xt−1=a(t)xt+b(t)Fθ(xt,y,t)7 are tuned per dataset via grid search rather than set adaptively, leaving open whether a principled, automatic calibration exists — particularly since spectrally sparser inputs appear to require stronger xt−1=a(t)xt+b(t)Fθ(xt,y,t)8. At the ×32 extreme, OVG trades distortion for realism (PSNR drops below standard inversion), and the paper does not establish a criterion for selecting operating points when no ground truth is available. Finally, evaluation is confined to two datasets and 50-step inversions; whether spectral collapse manifests identically for text-conditioned large-scale models, or under fewer/more sampling steps, remains untested.
Conclusion
This paper establishes spectral collapse as a failure mode intrinsic to deterministic diffusion inversion dynamics — universal across U-Net and DiT architectures and across pixel and latent spaces — driven by the interaction between the xt−1=a(t)xt+b(t)Fθ(xt,y,t)9-prediction objective, network spectral bias, and low-SNR regimes near the end of the inversion trajectory. It further shows that the prediction target itself, not merely solver error, governs the phenomenon, and that stochastic remedies resolve it only by breaking the structural link to the input. OVG resolves this trade-off at inference time by injecting variance orthogonally to the structural gradient, achieving state-of-the-art joint texture-fidelity performance on unpaired super-resolution and sketch-to-image translation.