Interpretability of zero-shot LLM rapport estimates

Identify the linguistic and interactional cues underlying Gemini 2.5 Flash’s rapport estimates in real-world human-robot interaction.

Background

Gemini 2.5 Flash achieves strong rapport-estimation performance from text-only input and improves further when fused with audio-visual predictors. However, the study evaluates predictive accuracy rather than explaining which observable cues drive the model’s judgments. The authors explicitly state that the linguistic and interactional basis of the model’s estimates is unresolved, making interpretability an open direction for understanding and validating automated rapport assessment.

References

Finally, although Gemini 2.5 Flash performed strongly, the linguistic and interactional cues underlying its rapport estimates remain unclear and should be examined in future work.

Multimodal Rapport Estimation in Real-World HRI  (2608.18401 - Sakuramoto et al., 19 Aug 2026) in Section 5.4, “Limitations and Future Work”