Reliable explainability for current AI models

Develop robust, reliable methods to properly explain the decisions of current artificial intelligence systems—especially deep learning and large language models—beyond existing post-hoc interpretability techniques whose outputs can be inconsistent or misleading.

Background

After reviewing the limitations of existing explainability approaches, the text highlights persistent opacity in deep learning–based systems and the inadequacy of current post-hoc methods for trustworthy explanations in critical domains.

The authors explicitly state that we still do not know how to properly explain the decisions of contemporary AI models, framing this as an ongoing, unresolved challenge for interpretability and trust.

References

In practice, this research confirms that we do not know how to properly explain the decisions of current AIs .

— The Impact of Artificial Intelligence on Human Thought  (2508.16628 - Gesnot, 15 Aug 2025) in Chapter 6: "Black Box" AI and the Hypothesis of an Orchestrating Consciousness, Explainability and Trust in AI

Explainability, rare-variant interpretation, interoperability, data governance, and clinical validation are open problems, and progress on any one does not guarantee progress on the others.

— Responsible Integration of AI in Cancer Genomics: Barriers, Risks, and Pathways to Trustworthy Clinical Translation  (2608.30912 - İlgen et al., 31 Aug 2026) in Section 1, Introduction

Thus, the open questions are: How can the trustworthiness of an autonomous refactoring system be measured in practice? What evidence and explanation mechanisms are needed for developers to trust autonomous refactoring decisions?

— Continuous Autonomous Refactoring: A Research Roadmap for AI-Driven Code Quality Maintenance  (2609.01236 - Sun et al., 1 Sep 2026) in Section 3.5, The Trust Problem

As a result, we still do not know how much progress either approach---interpretable-by-design or post hoc interpretability---has made toward the original goal, nor which aspects of the problem remain unsolved.

— From Interpretability Methods to Interpretable Models  (2609.05399 - Colin et al., 4 Sep 2026) in Section 4, paragraph “Where do models actually stand?”

How to reliably interpret these features remains an open problem.

— From Input to Output: A Flexible Agent for Dual-End Interpretation of Sparse Autoencoder Features  (2609.35367 - Liu et al., 28 Sep 2026) in Section 1, Introduction