Quantitative validation of attribution faithfulness and stability

Develop formal quantitative tests of Integrated Gradients attribution faithfulness and stability, together with analyses distinguishing token identity from laboratory-value contributions in Explainable Transformer Models for Clinical Prediction Tasks on Structured Electronic Health Records.

Background

The paper uses Integrated Gradients to attribute BERT-LER predictions to coded EHR events and laboratory measurements. However, the discussion acknowledges that Integrated Gradients can distribute importance across correlated tokens and does not explicitly represent feature interactions. The reported validation is therefore qualitative, based on clinical plausibility and comparison with other models, rather than a formal assessment of whether the attributions faithfully and stably reflect the model’s decision process.

The authors explicitly leave formal quantitative evaluation of attribution faithfulness and stability, as well as analyses separating the effects of laboratory-test identity from laboratory-value information, for future work. These analyses would address important concerns about the reliability and interpretability of explanations generated by BERT-LER.

References

Integrated Gradients itself can distribute importance across correlated tokens and does not make feature interactions explicit; more broadly, our validation is qualitative rather than a formal quantitative test of attribution faithfulness or stability, which, along with identity-versus-value analyses, we leave to future work.

Explainable Transformer Models for Clinical Prediction Tasks on Structured Electronic Health Records  (2608.20315 - Du et al., 20 Aug 2026) in Discussion, Limitations paragraph