Develop efficient and accurate causal attribution methods
Develop attribution methods for deep neural networks that efficiently and accurately measure the causal importance of inputs or upstream components on downstream activations and predictions, overcoming limitations of first‑order gradient approximations and distribution‑shifting perturbations.
References
Developing efficient and accurate attribution methods thus remains an open problem.
Extending the framework in this direction while keeping it parameter-free is open, and is the most obvious next step.
These estimates are known to be noisy and expensive at scale, and their fidelity is contested \citep{li2025influence}, leaving their usefulness as a filtering signal open.
Consequently, it remains unclear why a particular perturbation succeeds.
We will also examine whether highly attended regions exert greater influence on model predictions under controlled perturbations.
In particular, token-level attribution across layers and denoising steps remains open.