Evaluating Agents That Autonomously Generate Their Own Tasks
Establish reliable evaluation protocols and criteria for large language model agents that autonomously generate their own tasks in open-ended environments, enabling systematic assessment of their capabilities, behaviors, and progress without relying solely on predefined, single-run task performance.
References
Evaluating agents that generate their own tasks remains an open challenge; here we summarize qualitative observations of our system.
— LLM Agents Beyond Utility: An Open-Ended Perspective
(2510.14548 - Nachkov et al., 16 Oct 2025) in Section 3 (Qualitative Results), opening paragraph
Together, these efforts shift evaluation from static code correctness toward the quality of executable, interactive game artifacts. However, these works score a single submitted build, leaving open how an agent should keep improving a game across rounds of development, which is the setting we study.
— RSIGame: Autonomous Agentic Game Development with Recursive Self-improvement
(2609.39045 - Wu et al., 30 Sep 2026) in Section 4, Related Work, paragraph ‘Game Generation and Development Benchmarks’