Papers
Topics
Authors
Recent
Search
2000 character limit reached

Interpretability Guarantees with Merlin-Arthur Classifiers

Published 1 Jun 2022 in cs.LG and cs.AI | (2206.00759v3)

Abstract: We propose an interactive multi-agent classifier that provides provable interpretability guarantees even for complex agents such as neural networks. These guarantees consist of lower bounds on the mutual information between selected features and the classification decision. Our results are inspired by the Merlin-Arthur protocol from Interactive Proof Systems and express these bounds in terms of measurable metrics such as soundness and completeness. Compared to existing interactive setups, we rely neither on optimal agents nor on the assumption that features are distributed independently. Instead, we use the relative strength of the agents as well as the new concept of Asymmetric Feature Correlation which captures the precise kind of correlations that make interpretability guarantees difficult. We evaluate our results on two small-scale datasets where high mutual information can be verified explicitly.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (61)
  1. Explaining individual predictions when features are dependent: More accurate approximations to Shapley values. Artificial Intelligence, 298:103502, 2021.
  2. C. Agarwal and A. Nguyen. Explaining image classifiers by removing input features using generative models. In Computer Vision – ACCV 2020, pages 101–118. Springer International Publishing, 2021.
  3. D. Alvarez-Melis and T. S. Jaakkola. Towards robust interpretability with self-explaining neural networks. In Proceedings of the 32nd International Conference on Neural Information Processing Systems, NIPS’18, page 7786–7795. Curran Associates Inc., 2018.
  4. Fairwashing explanations with off-manifold detergent. In International Conference on Machine Learning, pages 314–323. PMLR, 2020.
  5. Learning to give checkable answers with prover-verifier games. arXiv preprint arXiv:2108.12099, 2021.
  6. S. Arora and B. Barak. Computational complexity: a modern approach. Cambridge University Press, 2009.
  7. Interpretable neural predictions with differentiable binary variables. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 2963–2977. Association for Computational Linguistics, 2019.
  8. Provably efficient, succinct, and precise explanations. Advances in Neural Information Processing Systems, 34:6129–6141, 2021.
  9. J. S. Bridle. Training stochastic model recognition algorithms as networks can lead to maximum mutual information estimation of parameters. In Proceedings of the 2nd International Conference on Neural Information Processing Systems, NIPS’89, page 211–217. MIT Press, 1989.
  10. Explaining image classifiers by counterfactual generation. arXiv preprint arXiv:1807.08024, 2018.
  11. A game theoretic approach to class-wise selective rationalization. Advances in neural information processing systems, 32, 2019.
  12. Invariant rationalization. In International Conference on Machine Learning, pages 1448–1458. PMLR, 2020.
  13. Interpretable by design: Learning predictors by composing interpretable queries. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(6):7430–7443, 2022.
  14. Learning to explain: An information-theoretic perspective on model interpretation. In International Conference on Machine Learning, pages 883–892. PMLR, 2018.
  15. P. Dabkowski and Y. Gal. Real time image saliency for black box classifiers. In Proceedings of the 31st International Conference on Neural Information Processing Systems, NIPS’17, page 6970–6979, Red Hook, NY, USA, 2017. Curran Associates Inc. ISBN 9781510860964.
  16. You shouldn’t trust me: Learning models which conceal unfairness from multiple explanation methods. In SafeAI@ AAAI, volume 2560 of CEUR Workshop Proceedings, pages 63–73. CEUR-WS.org, 2020.
  17. Explanations can be manipulated and geometry is to blame. Advances in neural information processing systems, 32:13589–13600, 2019.
  18. D. Dua and C. Graff. UCI machine learning repository, 2017. URL http://archive.ics.uci.edu/ml.
  19. R. M. Fano. Transmission of information: A statistical theory of communications. American Journal of Physics, 29(11):793–794, 1961.
  20. R. C. Fong and A. Vedaldi. Interpretable explanations of black boxes by meaningful perturbation. In Proceedings of the IEEE international conference on computer vision, pages 3429–3437, 2017.
  21. Shapley explainability on the data manifold. arXiv preprint arXiv:2006.01272, 2020.
  22. Interactive proofs for verifying machine learning. In 12th Innovations in Theoretical Computer Science Conference (ITCS 2021), volume 185 of LIPIcs, pages 1–19. Schloss Dagstuhl-Leibniz-Zentrum für Informatik, 2021.
  23. Explaining and harnessing adversarial examples. In 3rd International Conference on Learning Representations, ICLR 2015, 2015.
  24. B. Goodman and S. Flaxman. European Union regulations on algorithmic decision-making and a “right to explanation”. AI magazine, 38(3):50–57, 2017.
  25. Fooling neural network interpretations via adversarial model manipulation. Advances in Neural Information Processing Systems, 32:2925–2936, 2019.
  26. Abduction-based explanations for machine learning models. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 33, pages 1511–1519, 2019.
  27. AI Safety via debate. arXiv preprint arXiv:1805.00899, 2018.
  28. On explaining decision trees. arXiv preprint arXiv:2010.11034, 2020.
  29. M. Jaggi. Revisiting Frank-Wolfe: Projection-free sparse convex optimization. In International Conference on Machine Learning, pages 427–435. PMLR, 2013.
  30. J. Kleinberg and E. Tardos. Algorithm design. Pearson Education India, 2006.
  31. Captum: A unified and generic model interpretability library for PyTorch, 2020.
  32. Unmasking clever hans predictors and assessing what machines really learn. Nature communications, 10(1):1–8, 2019.
  33. Rationalizing neural predictions. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 107–117. Association for Computational Linguistics, 2016.
  34. Generative counterfactual introspection for explainable deep learning. In 2019 IEEE Global Conference on Signal and Information Processing (GlobalSIP), pages 1–5. IEEE, 2019.
  35. A unified approach to interpreting model predictions. In Proceedings of the 31st international conference on neural information processing systems, pages 4768–4777, 2017.
  36. J. Macdonald and S. Wäldchen. A complete characterisation of ReLu-Invariant Distributions. In International Conference on Artificial Intelligence and Statistics, pages 1457–1484. PMLR, 2022.
  37. A rate-distortion framework for explaining neural network decisions. arXiv preprint arXiv:1905.11092, 2019.
  38. Explaining neural network decisions is hard. In XXAI Workshop, 37th ICML, 2020.
  39. Interpretable neural networks with Frank-Wolfe: Sparse relevance maps and relevance orderings. In International Conference on Machine Learning, pages 14699–14716. PMLR, 2022.
  40. Explanations for monotonic classifiers. In International Conference on Machine Learning, pages 7469–7479. PMLR, 2021.
  41. A survey on bias and fairness in machine learning. ACM Computing Surveys (CSUR), 54(6):1–35, 2021.
  42. This is not the texture you are looking for! introducing novel counterfactual explanations for non-experts using generative adversarial learning. arXiv preprint arXiv:2012.11905, 2020.
  43. A multidisciplinary survey and framework for design and evaluation of explainable ai systems. ACM Transactions on Interactive Intelligent Systems (TiiS), 11(3-4):1–45, 2021.
  44. Assessing heuristic machine learning explanations with model counting. In International Conference on Theory and Applications of Satisfiability Testing, pages 267–278. Springer, 2019.
  45. The building blocks of interpretability. Distill, 3(3):e10, 2018.
  46. Deep neural network training with Frank-Wolfe. arXiv preprint arXiv:2010.07243, 2020.
  47. S. Poulis and S. Dasgupta. Learning with feature feedback: from theory to practice. In Artificial Intelligence and Statistics, pages 1104–1113. PMLR, 2017.
  48. ”Why should I trust you?” Explaining the predictions of any classifier. In Proceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining, pages 1135–1144, 2016.
  49. Anchors: High-precision model-agnostic explanations. In Proceedings of the AAAI conference on artificial intelligence, volume 32 of AAAI’18, pages 1527–1535. AAAI Press, 2018.
  50. U-net: Convolutional networks for biomedical image segmentation. In International Conference on Medical image computing and computer-assisted intervention, pages 234–241. Springer, 2015.
  51. Stabilizing training of generative adversarial networks through regularization. In Advances in neural information processing systems, volume 30 of NIPS’17, page 2015–2025, 2017.
  52. L. S. Shapley. 17. A value for n-person games. Princeton University Press, 2016.
  53. A symbolic approach to explaining bayesian network classifiers. In Proceedings of the 27th International Joint Conference on Artificial Intelligence, page 5103–5111, 2018.
  54. Fooling lime and shap: Adversarial attacks on post hoc explanation methods. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, pages 180–186, 2020.
  55. Counterfactual explanations can be manipulated. Advances in neural information processing systems, 34:62–75, 2021.
  56. The computational complexity of understanding binary classifier decisions. Journal of Artificial Intelligence Research, 70:351–387, 2021.
  57. S. Wäldchen. Hardness of deceptive certificate selection. In L. Longo, editor, Explainable Artificial Intelligence, pages 415–427. Springer Nature Switzerland, 2023.
  58. Training characteristic functions with reinforcement learning: XAI-methods play Connect Four. In International Conference on Machine Learning, pages 22457–22474. PMLR, 2022.
  59. Stabilizing generative adversarial networks: A survey. arXiv preprint arXiv:1910.00927, 2019.
  60. A learning-theoretic U-Net framework for certified auditing of machine learning models. arXiv preprint arXiv:2206.04740, 2022.
  61. Rethinking cooperative rationalization: Introspective extraction and complement control. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 4094–4103. Association for Computational Linguistics, 2019.
Citations (4)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.