Foveated Diffusion: Efficient Adaptive Generation
- Foveated Diffusion is a computational framework that exploits gaze-based, mixed-resolution tokenization to optimize image and video synthesis.
- It employs a foveation mask and a token-density function to allocate high detail in a gaze-centered region while reducing peripheral resolution, thereby cutting computational costs.
- Empirical evaluations indicate up to 2× acceleration for images and 3.5–4× for videos, achieving perceptual quality on par with full-resolution outputs.
Foveated Diffusion is a computational framework for spatially adaptive image and video generation based on diffusion and flow-matching models that exploits the eccentricity-dependent acuity of human vision. By allocating higher token density in a small, gaze-centered "foveal" region and lower density in the visual periphery, Foveated Diffusion achieves perceptually equivalent generation quality to full-resolution models while significantly reducing token count and computational complexity. The approach includes a principled mixed-resolution tokenization scheme, compatibility with standard VAE encoders, and a fine-tuning procedure applicable to pretrained diffusion or flow-matching models (Chao et al., 24 Mar 2026).
1. Foveation Mask and Token-Density Function
Foveated Diffusion relies on a spatial mask designating the foveal region (typically a disk of radius centered at gaze location ). Pixels within the disk satisfy , indicating retention of high-acuity detail, while the periphery () is treated at reduced resolution: To manage token allocation, a token-density function is introduced: where is the base patch size and is a periphery downsampling factor (commonly 0). Thus, regions near the gaze receive standard token density 1, while the periphery is sparsified by 2.
2. Mixed-Resolution Token Construction
The tokenization process begins with a high-resolution image 3. A standard VAE encoder 4 generates a uniform grid of 5 tokens: 6 A low-resolution version 7 is produced by bicubic downsampling by 8, then encoded: 9 A merge operation constructs the mixed-resolution latent sequence: 0 where, for each patch 1, a high- or low-resolution token is selected based on 2. The resulting sequence length is
3
with 4. The merged sequence corresponds to the foveated image,
5
where 6 denotes upsampling by 7.
3. Model Post-Training via Foveated Fine-Tuning
Foveated Diffusion extends any pretrained diffusion or flow-matching model by fine-tuning on mixed-resolution latents. The flow-matching procedure operates as follows:
- Sample a Gaussian target 8.
- For a uniform 9, create a noisy latent: 0
- The velocity field is 1.
- The model 2 is trained to predict this velocity, with loss: 3 No foveation-specific regularization is imposed; only standard weight decay or low-rank LoRA is applied to 4. All learning occurs directly on the 5-length foveated sequence.
4. Computational Complexity and Acceleration
The number of tokens for full-resolution generation is 6, while for foveated generation, it is
7
Attention-based models scale as 8. When 9, the peripheral term dominates and the complexity for foveated diffusion is approximated as
0
implying a theoretical speedup factor of 1 over full-resolution (e.g., 2 for 3) up to corrections from the fovea itself.
| Generation Type | Token Ratio to Full | Effective Speedup |
|---|---|---|
| Image (1024×1024) | 43%, 30%, 26% | 1.61×, 1.98×, 2.08× |
| Video (480p) | ≈38% | ≈3.5× |
5. Empirical Evaluation and Perceptual Quality
Image (1024×1024) and video (480p) generation experiments demonstrate that Foveated Diffusion achieves speedups of up to 2× for images and 3.5× for video compared to full-resolution, with mixed-resolution token ratios in the range 26–43% for images and 38% for video. Key findings include:
- Human-Preference Score (HPSv2.1) matches full-resolution performance (e.g., 0.280 vs. 0.279).
- CLIP-based metrics and precision are matched or surpassed relative to the full-resolution baseline.
- FID is not conclusive for foveation tasks.
- Qualitatively, naïve mixed-resolution approaches cause scale or structure artifacts at resolution boundaries, whereas Foveated Diffusion produces visuals indistinguishable from full-res both inside and outside the fovea.
A 2AFC user study with 11 participants (60 trials each, gaze-fixed) reveals:
- Foveated vs. full-res: 47.3% preference for foveated (p=0.48), indicating perceptual indistinguishability.
- Foveated vs. naïve: 87.4% preference for foveated (p<0.0001).
- Full-res vs. naïve: 90.8% preference for full-res (p<0.0001).
Video generation evaluations (VBench metrics) confirm that Foveated Diffusion matches or exceeds full-res performance for subject consistency, background consistency, and motion smoothness, while the naïve baseline exhibits duplicated/misaligned objects at resolution transitions.
6. Practical Implications and Scope
Foveated Diffusion requires access to either real-time gaze location (e.g., via eye tracking) or a specified fixation point, making it directly applicable to interactive and streaming scenarios where user gaze can be estimated. By aligning computational resources with the effective acuity of human vision, the framework achieves up to 2× image and 4× video synthesis acceleration, with outputs that are, per both human preference scores and user studies, perceptually indistinguishable from uniformly high-resolution synthesis (Chao et al., 24 Mar 2026).