Human-transfer predictiveness of AI-AI Partner Loss

Determine whether AI-AI Partner Loss predicts Human-AI Partner Loss for deep reinforcement-learning systems such as BAD, SAD, and Other-Play, and for large-language-model agents paired with human players.

Background

The paper compares partner failures in AI-AI and human-AI Hanabi games and finds that AI-AI Partner Loss does not reliably predict Human-AI Partner Loss for the two agents available in both datasets, Outer and Intentional. The authors note that the deep reinforcement-learning systems BAD, SAD, and Other-Play, as well as LLM-based agents playing with humans, remain insufficiently tested on this transfer question.

Resolving this problem would establish whether AI-AI evaluation can serve as a useful proxy for human-AI cooperation for broader classes of agents, or whether direct human-partner evaluation is necessary.

References

For deep RL systems, the convention-gap side of the question is now partially closed---the off-belief-learning family is measured directly in \S\ref{sec:obl_results}---while whether AI-AI Partner Loss predicts Human-AI Partner Loss for such systems (BAD, SAD, Other-Play) and for LLM-vs-human play remains open.

The Convention Gap: Towards Measuring Implicit Communication in Cooperative AI Evaluation  (2609.11489 - Fukushima et al., 10 Sep 2026) in Section Discussion, paragraph “Limitations”