Autonomy of LLMs for production-scale, specification-driven software construction
Determine whether large language models can autonomously build production-scale software systems from explicit specifications.
References
Although LLMs have demonstrated impressive coding capabilities, their ability to autonomously build production-scale software from explicit specifications remains an open question.
— SWE-AGI: Benchmarking Specification-Driven Software Construction with MoonBit in the Era of Autonomous Agents
(2602.09447 - Zhang et al., 10 Feb 2026) in Abstract (Page 1)
Third, a comparative evaluation against four installable frameworks in which agents write the code, GitHub Spec Kit, obra/superpowers, BMAD, and GSD, and a pre-registered hypothesis for the output-quality claim we cannot yet prove.
— Introducing Consort: A Spec-First Agent Framework for Enforced, Test-Driven Development on Live Database Branches
(2609.09671 - Hartman, 9 Sep 2026) in Section 1, Introduction; Section 5, “A controlled evaluation (pre-registered, future work)”
These challenges indicate that, although the underlying models are rapidly improving, fully autonomous software development remains an open research problem.
— Generation of Web Apps with Agentic IDEs: An Empirical Assessment
(2608.20903 - Marceca et al., 21 Aug 2026) in Section 2, subsection “LLM-Based Tools”