Extension beyond DiT-based super-resolution

Extend SPARK, the input-conditioned sparse activation modulation framework for frozen Diffusion Transformer-based super-resolution, to convolutional and U-Net-based backbones, other image-restoration tasks, and a broader range of real-world degradations.

Background

SPARK is evaluated only on DiT-based super-resolution models and three datasets, namely DIV2K, RealSR, and DRealSR. Its mechanism depends on identifying a sparse set of magnitude-dominant activation channels and learning an input-conditioned affine modulation predictor while keeping the backbone frozen.

The paper explicitly identifies extension to convolutional or U-Net-based architectures, other restoration problems, and more diverse real-world degradations as unresolved. Such extensions would test whether sparse activation-space adaptation generalizes beyond the architectural and task setting used in the experiments.

References

Extending the approach to convolutional or U-Net-based backbones, to other restoration tasks, and to a broader range of real-world degradations remains an open direction.

SPARK: Input-Conditioned Sparse Activation Modulation for Frozen DiT-based Super-Resolution  (2609.03813 - Putamorsi et al., 3 Sep 2026) in Appendix, Section “Limitations,” subsection “Scope and computational cost”