Temporally-Resolved Token Attribution Reveals the Generation Dynamics of Diffusion Language Models
Abstract: This work presents Diffusion Layer Integrated Gradients (DLIG), a token attribution method for diffusion LLMs (DLMs) that extends Integrated Gradients (IG~\cite{sundararajan2017axiomatic}) to arbitrary layers and denoising steps. DLIG attributes a DLM's progressive commitment to a self-generated or fixed completion for an input prompt. We establish direct correspondences between DLIG and the IG axioms of completeness, implementation invariance, linearity, and symmetry preservation. As a lightweight complement to interventional analysis, DLIG provides an inexpensive first check of mechanistic hypotheses across the denoising trajectory. We demonstrate this on word-sense disambiguation, multi-hop graph reasoning, and sentence infilling, revealing how DLMs draw on inputs across positions, layers, and denoising steps.
Paper Prompts
Sign up for free to create and run prompts on this paper.