Explaining the Full–Intentional partner-loss difference

Determine why the Full intent-based Hanabi agent produces lower Human-AI Partner Loss than the Intentional agent despite their nearly identical hint playability rates, through controlled comparisons that resolve the candidate explanations for their 3.2-percentage-point difference.

Background

Full and Intentional have similar hint playability rates, yet human partners fail less often with Full than with Intentional. Full also exhibits a positive own-play convention gap, whereas Intentional’s own-play gap is near zero, suggesting that differences in intent-based interpretation or play style may be relevant.

The between-participant difference does not replicate within participants: the paired Full-minus-Intentional contrast is not statistically separable from zero. Consequently, the paper treats the candidate mechanisms for this micro-ordering as unresolved and calls for controlled comparisons.

References

These differences motivate the intent-based interpretation below, but one caution applies to the Full-vs-Intentional micro-ordering specifically: unlike the Full/Intentional-vs-Outer contrast, which replicates within participants, the Full$-$Intentional difference does not replicate within participants ($-$3.4~pp paired mean, 95\% CI $[-7.8, +0.9]$, sign-test $p = 0.63$; Appendix~\ref{app:within}), so the between-participant 3.2~pp difference and its candidate explanations should be treated as unresolved pending controlled comparisons.

The Convention Gap: Towards Measuring Implicit Communication in Cooperative AI Evaluation  (2609.11489 - Fukushima et al., 10 Sep 2026) in Section Discussion, paragraph “Beyond hint-level playability”; see also Appendix “Within-Participant Partner Analysis”