Extending DLIG beyond a single discrete masked-diffusion model

Extend Diffusion Layer Integrated Gradients from the 355-million-parameter DiffuGPT-M discrete masked-diffusion model to continuous-space diffusion language models, hybrid autoregressive–diffusion models, and larger natively trained diffusion language models.

Background

The empirical evaluation instantiates DLIG only on DiffuGPT-M, a 355-million-parameter discrete masked-diffusion model adapted from GPT-2-medium. Although the method is described as broadly applicable to iterative diffusion processes, the paper does not establish its behavior or validity across other diffusion-language-model families or substantially larger native architectures.

The authors explicitly leave extension to continuous-space and hybrid autoregressive–diffusion models, as well as larger natively trained diffusion LLMs, for future work. These extensions would test whether the attribution framework transfers beyond the specific discrete masked-diffusion setting used in the experiments.

References

We also instantiate DLIG on a single 355M discrete masked-diffusion model (DiffuGPT-M); extending to continuous-space and hybrid AR--diffusion families (\S~\ref{sec:background}), and to larger, natively-trained DLMs, is left for future work.

— Temporally-Resolved Token Attribution Reveals the Generation Dynamics of Diffusion Language Models  (2610.01177 - Aswal et al., 1 Oct 2026) in Section 6, “Limitations and Future Work”