Develop efficient and accurate causal attribution methods

Develop attribution methods for deep neural networks that efficiently and accurately measure the causal importance of inputs or upstream components on downstream activations and predictions, overcoming limitations of first‑order gradient approximations and distribution‑shifting perturbations.

Background

Attribution techniques aim to quantify the causal contribution of inputs or internal components to outputs. Many current approaches rely on gradients (first‑order approximations) or perturbations that may take models off-distribution, each with known theoretical and practical pitfalls.

Given these limitations and adversarial vulnerabilities, the authors conclude that current methods are insufficiently reliable or efficient. A robust solution would be broadly useful for interpretation, validation, and downstream applications such as debugging and safety auditing.

References

Developing efficient and accurate attribution methods thus remains an open problem.

— Open Problems in Mechanistic Interpretability  (2501.16496 - Sharkey et al., 27 Jan 2025) in Reverse engineering step 2: Describing the functional role of components — Attribution methods (Section 2.1.2, parasection “Attribution methods are necessary for causal explanations but are often difficult to interpret”)

Extending the framework in this direction while keeping it parameter-free is open, and is the most obvious next step.

— Evaluating Explanation Methods by the Predictors They Induce  (2609.20058 - Selbæk et al., 17 Sep 2026) in Section 'The additive form is a choice, and so is the one-dimensional curve' in the Discussion

These estimates are known to be noisy and expensive at scale, and their fidelity is contested \citep{li2025influence}, leaving their usefulness as a filtering signal open.

— Can Data Attribution Filter Out Subliminal Learning? Not Reliably  (2609.20027 - Weckbecker et al., 17 Sep 2026) in Section 1, Introduction

Consequently, it remains unclear why a particular perturbation succeeds.

— Generating Medical Image Counterfactuals using Causal Explanations  (2609.02697 - Kelly et al., 2 Sep 2026) in Section Related Work

We will also examine whether highly attended regions exert greater influence on model predictions under controlled perturbations.

— Semantic RGB--Depth Based Surgical Skill Assessment in Microscopic Stereo Videos  (2610.01205 - Mao et al., 1 Oct 2026) in Section Conclusions

In particular, token-level attribution across layers and denoising steps remains open.

— Temporally-Resolved Token Attribution Reveals the Generation Dynamics of Diffusion Language Models  (2610.01177 - Aswal et al., 1 Oct 2026) in Section 2, paragraph “Explainable AI (XAI), Attribution Methods and Diffusion Explainability”