Validity of LLM-simulated users as proxies for real users

Determine whether interactions between large language model agents and LLM-simulated users accurately reflect and predict interactions between those agents and real human users in agentic evaluations, in order to establish the validity of using LLM-simulated users for benchmarking multi-turn, tool-using conversational agents.

Background

Agentic benchmarks often replace human participants with LLM-simulated users to enable scalable, automated evaluation of multi-turn, tool-using conversational agents. However, without direct validation against human interactions, there is a risk that outcomes measured with simulated users do not generalize to real users, potentially leading to miscalibrated assessments of agent capabilities.

This paper studies this concern using τ-Bench retail tasks and a cross-national user study, introducing a Human–LLM calibration metric and documenting systematic miscalibration across difficulty levels and demographic groups—motivating the need to rigorously determine whether simulated interactions reliably predict real human outcomes.

References

Second, without validation with actual users \citep{salaudeen2025measurementmeaningvaliditycenteredframework}, it remains unclear whether interactions between agents and LLM-simulated users accurately reflect and predict interactions between agents and real people (validity, Figure \ref{fig:fig1}).

Lost in Simulation: LLM-Simulated Users are Unreliable Proxies for Human Users in Agentic Evaluations  (2601.17087 - Seshadri et al., 23 Jan 2026) in Introduction

Third, what endpoint signature is sufficient to certify an agent group as a proxy for human collective cognition: full consensus alone is not, and whether jointly matching minority survival, participation, and correct-versus-wrong consensus out of sample suffices is an open empirical question.

Language-model groups overstate consensus when replaying human deliberation on a reasoning task  (2609.20543 - Shao, 17 Sep 2026) in Discussion, final paragraph

Whether the proposed method can help discover issues across various dialogue systems remains a topic for future work.

Generating Diverse Personas for User Simulators to Test Interview Dialogue Systems  (2608.19549 - Nakano et al., 20 Aug 2026) in Section Discussion, immediately before the paragraph beginning “In the future, we aim to integrate this method...”

The validity of LLM-based student simulation has emerged as an open question.

StudentSim: Training LLM-based Student Simulators  (2609.01591 - Yang et al., 1 Sep 2026) in Appendix, Section Student-Simulation Validity (Appendix A, subsection app:related_work_validity)

Whether real users would react to the replies in the same way cannot be concluded from this study alone; a comparison of the evaluator's scores with human ratings of the same exchanges is the natural next step.

How a Chatbot's Response Style Shapes a Classroom: A Multi-Agent Simulation of Students Consulting AI  (2609.05018 - Tamai et al., 4 Sep 2026) in Section 4.1, Limitations

External validity to human-authored interactions is therefore an open question that future work should address with real-world dialogue corpora.

CC-Mediation: Evaluating Large Language Models for Cross-Cultural Conflict Mediation  (2609.04855 - Lee et al., 4 Sep 2026) in Limitations, subsection “Synthetic data and LLM-mediated supervision”