Papers
Topics
Authors
Recent
Search
2000 character limit reached

Posterior Collapse as Automatic Spectral Pruning

Published 21 May 2026 in cs.LG and cond-mat.stat-mech | (2605.22691v1)

Abstract: We show that posterior collapse in ββ-VAEs implements automatic spectral pruning. A latent mode collapses if its contribution to reconstruction is below the cutoff set by ββ. Equilibrium solutions with different ββ thus reveal a cascade of collapses as latent modes decouple from least to most useful. We derive this as a consequence of the loss via a Landau stability analysis. We define a latent-rescaling-invariant order parameter that ranks active latent modes and whose collapse thresholds identify which effective variables to inspect first. In the linear Gaussian case, the collapse spectrum, utility spectrum, and normalized PCA spectrum coincide, and each collapse follows a mean-field law. We test these predictions on the WorldClim dataset.

Authors (1)

Summary

  • The paper shows that posterior collapse is an equilibrium property of the β-VAE objective, with each latent mode becoming inactive when its marginal reconstruction utility falls below a regularization-controlled threshold.
  • In the linear Gaussian setting, collapse thresholds, distortion reductions, and normalized PCA eigenvalues coincide exactly, establishing a utility–threshold duality for ranking and pruning latent dimensions.
  • WorldClim experiments validate the scale-invariant posterior signal fraction and the predicted duality through approximately rank 16, while revealing optimization gaps and weaker threshold estimates for higher-rank modes.

Overview

This paper, "Posterior Collapse as Automatic Spectral Pruning" (2605.22691), argues that posterior collapse in β\beta-VAEs is not a training pathology but an equilibrium property of the loss itself: it implements automatic spectral pruning of latent modes. A latent mode collapses when its marginal contribution to reconstruction falls below a cutoff set by the regularization strength. The author derives this from the β\beta-VAE objective via a Landau stability analysis, defines a scale-invariant order parameter that ranks active latent modes, and shows that in the linear Gaussian case the collapse spectrum, the utility spectrum, and the normalized PCA spectrum coincide exactly. Predictions are tested empirically on the WorldClim climatology dataset.

The paper's central claim is deliberately stronger than prior rate–distortion treatments: rather than describing a smooth global D(R)D(R) curve as β\beta varies, it resolves that curve into mode-wise collapse events occurring at ordered thresholds. The claim is that the set of these thresholds (the collapse spectrum) and the set of mode-wise distortion reductions (the utility spectrum) coincide mode by mode — a "utility–threshold duality."

Collapse spectroscopy

The analysis uses the standard β\beta-VAE objective L=D+βR\mathcal{L} = D + \beta R with a Gaussian decoder of fixed isotropic variance σdec2\sigma_\text{dec}^2. Distortions are normalized by the total data variance V=∑kλkV = \sum_k \lambda_k, where λk\lambda_k are PCA eigenvalues of the centered data, yielding a normalized control parameter

T≡βσdec2V,T \equiv \frac{\beta \sigma_\text{dec}^2}{V},

which the author calls an effective temperature. This normalization is presented as practically important: raw β\beta0 values depend on input scaling, loss normalization, and reconstruction convention, so statements such as "β\beta1 gives β\beta2 active dimensions" are not intrinsic to the data or model family. In normalized units, thresholds become comparable across datasets; for β\beta3, all normalized PCA thresholds lie below the information price and the fully collapsed distortion is exactly one.

For each latent coordinate, four observables are monitored on held-out data: posterior mean-square, posterior variance, mean log-variance, and per-coordinate KL rate. The default ranking observable is the posterior signal fraction

β\beta4

a monotone function of the posterior SNR. Two spectra are then defined: the collapse spectrum β\beta5 extracted from fits of the one-mode law β\beta6, and the utility spectrum β\beta7, the marginal reduction in normalized distortion from adding the β\beta8-th ranked mode in truncated reconstructions.

Single-mode Landau derivation

For a one-dimensional linear Gaussian VAE with encoder β\beta9 and linear decoder D(R)D(R)0, the author eliminates the decoder analytically at fixed encoder statistics. The optimized reconstruction depends only on the signal fraction D(R)D(R)1, not on the absolute latent scale D(R)D(R)2: under a latent rescaling D(R)D(R)3, both D(R)D(R)4 and the optimized reconstruction are invariant. This is why the signal fraction, rather than raw posterior moments, is the robust order parameter — raw observables drift with latent-scale changes while the collapse law for D(R)D(R)5 remains stable.

Expanding the decoder-eliminated loss around D(R)D(R)6 yields an exact Landau form,

D(R)D(R)7

with local reduced temperature D(R)D(R)8. The coefficient of D(R)D(R)9 changes sign at β\beta0: above threshold the collapsed branch (β\beta1) is stable; below it, an active branch appears continuously with β\beta2. Notably, this Landau structure follows from the exact loss after decoder minimization, not from a phenomenological free-energy ansatz. The author is explicit that the thermodynamic analogy is structural rather than literal: the objective is optimized over variational parameters within a chosen Gaussian family, not obtained by marginalizing over microstates, so no claim of a genuine second-order phase transition or critical behavior is made.

On the canonical branch, the theory also predicts the posterior log-variance linear in β\beta3 with slope β\beta4, the rate linear in β\beta5 with slope β\beta6, and active-branch distortion β\beta7. The constant-variance posterior used here is acknowledged as an ansatz, tested via the Jensen gap diagnostic.

PCA calibration and empirical validation

In the linear Gaussian case, rotating inputs by the PCA basis diagonalizes the quadratic data-fit term, factorizing the full loss into independent one-mode problems. Each eigendirection has local coordinate β\beta8, so the universal one-mode threshold β\beta9 maps to a global threshold

β\beta0

The same eigenvalue governs the utility: the normalized distortion reduction from adding mode β\beta1 is also β\beta2. Hence the paper's main result:

β\beta3

i.e., the collapse threshold, reconstruction utility, and PCA explained-variance ratio are the same spectral weight for each ranked mode.

Empirically, the author trains 32-latent linear VAEs on WorldClim (19 bioclimatic variables at 10 arc-minute resolution, ~808k valid samples), using spatial-block splits (500 km equal-area blocks) to reduce leakage, with equilibrium-mode training per scan point (Adam, learning rate β\beta4, up to ~256 epochs). Three findings stand out:

  • Order-parameter scans: fixed-exponent fits of β\beta5 extract collapse thresholds cleanly; the raw posterior mean-square exhibits level crossings due to latent-scale drift, confirming why the scale-invariant signal fraction is preferred.
  • Truncated distortion: the linear VAE's normalized distortion plateaus near β\beta6 after rank 18, whereas PCA at rank 18 reaches below β\beta7 — a measurable optimization-limited gap between the trained VAE and the analytic optimum.
  • Utility–threshold duality: utilities versus fitted thresholds lie close to the diagonal β\beta8, validating the duality through roughly rank 16; the author concedes that threshold extraction becomes inaccurate at rank 17 and higher, where β\beta9 signals are small.

Interpretation as spectral pruning

The scan coordinate acts as a moving marginal utility cutoff: mode L=D+βR\mathcal{L} = D + \beta R0 stays active iff L=D+βR\mathcal{L} = D + \beta R1. Crucially, the criterion is mode-wise and marginal, not based on cumulative explained variance — analogous to truncating PCA by individual eigenvalues rather than cumulative variance. The collapse cascade thus orders learned coordinates by reconstruction relevance and counts how many are needed for a prescribed accuracy, providing a way to test low-dimensionality assumptions rather than assume them.

A practical protocol follows directly: choose a target normalized information price L=D+βR\mathcal{L} = D + \beta R2 (e.g., L=D+βR\mathcal{L} = D + \beta R3 retains modes with marginal normalized utility above 1%), convert back to L=D+βR\mathcal{L} = D + \beta R4 for the chosen convention, and use the signal fraction to count and rank active latents. The author notes that in the common L=D+βR\mathcal{L} = D + \beta R5 convention, nominal L=D+βR\mathcal{L} = D + \beta R6 corresponds to L=D+βR\mathcal{L} = D + \beta R7, which for standardized high-dimensional data can be far below 1, leaving many latents active — explaining the common observation of over-complete uncollapsed representations.

Limitations and open questions

Several limitations are stated plainly. The exact duality result holds only for the linear Gaussian baseline; beyond it, nonlinear decoders, non-Gaussian likelihoods, finite-training effects, and near-degenerate regimes can rotate the latent basis, mix modes, or renormalize the mean-field collapse law. The one-mode derivation applies locally on the active side of collapse even in nonlinear VAEs, but the relevant quadratic operator there is learned, branch-dependent, and changes with the active set, so finite utilities and thresholds must be measured from the scan rather than predicted. Empirically, the constant-variance ansatz is only approximate: measured posterior scales L=D+βR\mathcal{L} = D + \beta R8 deviate measurably from one over the scan, and the linear VAE's achievable distortion (L=D+βR\mathcal{L} = D + \beta R9 at rank 18) falls well short of the PCA reference. Threshold extraction degrades at high rank. Whether the utility–threshold duality survives quantitatively in nonlinear models — and whether leading nonlinear variables become more interpretable or disentangled at particular normalized information prices — is left open.

Conclusion

The paper reframes posterior collapse as an objective-level pruning mechanism whose ordered thresholds form a measurable spectrum. In the calibrated linear Gaussian setting, collapse spectrum, utility spectrum, and normalized PCA spectrum coincide, turning the linear VAE into a null model against which departures in nonlinear VAEs become quantitative signals of nonlinear representation learning rather than artifacts of loss normalization. The scale-invariant signal fraction provides a robust order parameter for ranking latents, and the normalized temperature provides a dataset-independent way to set σdec2\sigma_\text{dec}^20. The framework positions the σdec2\sigma_\text{dec}^21-scan as a probe for identifying which effective variables deserve inspection first, with nonlinear extensions as the natural next step.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.