Manifestation of spectral bias in latent-space training

Determine whether, and in what manner, the spectral bias addressed in pixel-space flow-matching models manifests in compressed latent-space models.

Background

The paper characterizes the investigated spectral bias primarily as a pixel-space phenomenon. Compressed latent representations produced by a variational autoencoder discard substantial high-frequency pixel detail, so it is unresolved whether the same frequency imbalance should appear in latent-space training or how it would be expressed there.

An SiT-B experiment using x-prediction finds some benefit from frequency-domain supervision, but neither the frequency-domain nor pixel-space loss matches the performance of the original v-prediction setup. The authors therefore do not establish whether the observed pixel-space mechanism transfers to latent-space generative models.

References

The spectral bias we address in this work is predominantly a pixel-space phenomenon. When operating in a compressed latent space, the data has already passed through a VAE, which by construction discards much of the high-frequency detail present in pixel space, so it is unclear whether -- or how -- the same bias manifests in latent space.

Balancing Frequencies and Pixels in Flow Matching  (2609.02748 - Degeorge et al., 2 Sep 2026) in Appendix, Section "Generalization to Latent-Space Models" (Section \ref{app:latent})