Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
129 tokens/sec
GPT-4o
28 tokens/sec
Gemini 2.5 Pro Pro
42 tokens/sec
o3 Pro
4 tokens/sec
GPT-4.1 Pro
38 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

Clustered Policy Decision Ranking (2311.12970v2)

Published 21 Nov 2023 in cs.LG and cs.AI

Abstract: Policies trained via reinforcement learning (RL) are often very complex even for simple tasks. In an episode with n time steps, a policy will make n decisions on actions to take, many of which may appear non-intuitive to the observer. Moreover, it is not clear which of these decisions directly contribute towards achieving the reward and how significant their contribution is. Given a trained policy, we propose a black-box method based on statistical covariance estimation that clusters the states of the environment and ranks each cluster according to the importance of decisions made in its states. We compare our measure against a previous statistical fault localization based ranking procedure.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (11)
  1. OpenAI Gym. CoRR, abs/1606.01540.
  2. Minigrid. https://github.com/Farama-Foundation/Minigrid.
  3. DARPA’s explainable artificial intelligence program. AI Magazine, 40(2): 44–58.
  4. Hotelling, H. 1936. Relations Between Two Sets of Variates. Biometrika, 28(3/4): 321–377.
  5. Jones, K. S. 1972. A statistical interpretation of term specificity and its application in retrieval. Journal of Documentation, 28(1): 11–21.
  6. Deep Learning, Transparency and Trust in Human Robot Teamwork. Preprint.
  7. Causal Policy Ranking. In ICLR2022 Workshop on the Elements of Reasoning: Objects, Structure and Causality.
  8. Ranking Policy Decisions. In Proceedings of Annual Conference on Neural Information Processing Systems (NeurIPS), 8702–8713.
  9. Learning Transferable Visual Models From Natural Language Supervision. arXiv:2103.00020.
  10. Reinforcement Learning: An Introduction. MIT Press.
  11. Between MDPs and Semi-MDPs: A Framework for Temporal Abstraction in Reinforcement Learning. Artificial Intelligence, 112: 181 – 211.

Summary

We haven't generated a summary for this paper yet.