Reliable explainability for current AI models
Develop robust, reliable methods to properly explain the decisions of current artificial intelligence systems—especially deep learning and large language models—beyond existing post-hoc interpretability techniques whose outputs can be inconsistent or misleading.
References
In practice, this research confirms that we do not know how to properly explain the decisions of current AIs .
Explainability, rare-variant interpretation, interoperability, data governance, and clinical validation are open problems, and progress on any one does not guarantee progress on the others.
Thus, the open questions are: How can the trustworthiness of an autonomous refactoring system be measured in practice? What evidence and explanation mechanisms are needed for developers to trust autonomous refactoring decisions?
As a result, we still do not know how much progress either approach---interpretable-by-design or post hoc interpretability---has made toward the original goal, nor which aspects of the problem remain unsolved.
How to reliably interpret these features remains an open problem.