Validity of LLM-simulated users as proxies for real users
Determine whether interactions between large language model agents and LLM-simulated users accurately reflect and predict interactions between those agents and real human users in agentic evaluations, in order to establish the validity of using LLM-simulated users for benchmarking multi-turn, tool-using conversational agents.
References
Second, without validation with actual users \citep{salaudeen2025measurementmeaningvaliditycenteredframework}, it remains unclear whether interactions between agents and LLM-simulated users accurately reflect and predict interactions between agents and real people (validity, Figure \ref{fig:fig1}).
Third, what endpoint signature is sufficient to certify an agent group as a proxy for human collective cognition: full consensus alone is not, and whether jointly matching minority survival, participation, and correct-versus-wrong consensus out of sample suffices is an open empirical question.
Whether the proposed method can help discover issues across various dialogue systems remains a topic for future work.
The validity of LLM-based student simulation has emerged as an open question.
Whether real users would react to the replies in the same way cannot be concluded from this study alone; a comparison of the evaluator's scores with human ratings of the same exchanges is the natural next step.
External validity to human-authored interactions is therefore an open question that future work should address with real-world dialogue corpora.