Quantitative validation of attribution faithfulness and stability
Develop formal quantitative tests of Integrated Gradients attribution faithfulness and stability, together with analyses distinguishing token identity from laboratory-value contributions in Explainable Transformer Models for Clinical Prediction Tasks on Structured Electronic Health Records.
References
Integrated Gradients itself can distribute importance across correlated tokens and does not make feature interactions explicit; more broadly, our validation is qualitative rather than a formal quantitative test of attribution faithfulness or stability, which, along with identity-versus-value analyses, we leave to future work.
— Explainable Transformer Models for Clinical Prediction Tasks on Structured Electronic Health Records
(2608.20315 - Du et al., 20 Aug 2026) in Discussion, Limitations paragraph