Autonomy of LLMs for production-scale, specification-driven software construction

Determine whether large language models can autonomously build production-scale software systems from explicit specifications.

Background

The paper introduces SWE-AGI, a benchmark designed to evaluate whether LLM-based agents can construct complex software systems from authoritative specifications under a fixed scaffold. This directly targets the gap between demonstrated coding skills and full, end-to-end system implementation from requirements.

By leveraging MoonBit’s nascent ecosystem and spec-first workflow, the benchmark minimizes retrieval-based shortcuts and emphasizes long-horizon reasoning and architectural consistency. The opening of the abstract frames the overarching, field-level uncertainty about the capability of LLMs to meet this standard of autonomy.

References

Although LLMs have demonstrated impressive coding capabilities, their ability to autonomously build production-scale software from explicit specifications remains an open question.

Third, a comparative evaluation against four installable frameworks in which agents write the code, GitHub Spec Kit, obra/superpowers, BMAD, and GSD, and a pre-registered hypothesis for the output-quality claim we cannot yet prove.

— Introducing Consort: A Spec-First Agent Framework for Enforced, Test-Driven Development on Live Database Branches  (2609.09671 - Hartman, 9 Sep 2026) in Section 1, Introduction; Section 5, “A controlled evaluation (pre-registered, future work)”

These challenges indicate that, although the underlying models are rapidly improving, fully autonomous software development remains an open research problem.

— Generation of Web Apps with Agentic IDEs: An Empirical Assessment  (2608.20903 - Marceca et al., 21 Aug 2026) in Section 2, subsection “LLM-Based Tools”