Turning video games into effective LLM benchmarks
Determine whether existing video game environments can be transformed into effective, reliable benchmarks for evaluating large language models, such that benchmark performance is discriminative and robust despite brittle vision perception, prompt sensitivity, and potential training-data contamination.
References
This leaves an open question: can we turn games into more effective benchmarks for evaluating LLMs?
— lmgame-Bench: How Good are LLMs at Playing Games?
(2505.15146 - Hu et al., 21 May 2025) in Section 1 (Introduction)
We do not lead with in-simulator goal-condition success: it is low across all methods with differences that are not statistically reliable, so the available checkers cannot separate planners --- an open measurement finding (\S\ref{sec:limitations}).
— SAGE: Symbolic Action-Gating and Editing for LLM Task Planners
(2609.34268 - Bui et al., 28 Sep 2026) in Section 4.4, subsection “Grounded execution in AI2-THOR,” paragraph “On goal-condition success”; also discussed in Section 4.11 (Limitations) and the Conclusion