DeNovoSWE: Automated Repo Generation
- DeNovoSWE is a large-scale dataset designed for complete software repository generation from documentation, leveraging LLM-based agents in a fully automated, sandboxed pipeline.
- It employs a divide-and-conquer strategy along with a critic-repair loop to ensure detailed, structurally coherent documentation and high-quality generation tasks.
- Supervised fine-tuning on DeNovoSWE significantly boosts performance on repository-generation benchmarks, improving test pass rates by over 40 percentage points.
Searching arXiv for the DeNovoSWE paper and closely related benchmark context. DeNovoSWE is a large-scale dataset for whole-repository generation introduced to support long-horizon software engineering with LLM-based code agents. It comprises 4,818 high-quality document-to-repository instances, each requiring generation of a complete repository from documentation, and is constructed through a fully automated, sandboxed agentic workflow without human annotation. The dataset is designed around “divide and conquer” and critic-repair philosophy, and it is paired with a difficulty-aware trajectory filtering strategy intended to balance data quality and diversity. In the associated study, supervised fine-tuning of Qwen3-30B-A3B on DeNovoSWE substantially improved performance on repository-generation benchmarks, including an increase on BeyondSWE-Doc2Repo from 5.8% to 47.2% (Zhao et al., 9 Jun 2026).
1. Definition and scope
DeNovoSWE targets a setting in which agents move beyond localized bug fixing in existing codebases toward architecting and implementing complete software repositories from high-level specifications. The central problem identified in the work is the scarcity of large-scale, verifiable whole-repository generation data for training such systems. DeNovoSWE is presented as a response to that data bottleneck through automated construction of end-to-end document-to-repository tasks (Zhao et al., 9 Jun 2026).
Each instance includes ground-truth documentation in the form of an overview together with capability chapters, a clean Docker image, unit tests, test patches, binary fixtures, and a precomputed difficulty score . The source task is not partial code completion or repair, but generation of an entire repository from documentation. This distinguishes the dataset from task formulations centered on local edits or limited-context synthesis. A plausible implication is that DeNovoSWE is intended not merely as a benchmark corpus but as training infrastructure for agents expected to sustain planning and execution across many interdependent implementation steps.
2. Automated construction pipeline
At the core of DeNovoSWE is a fully automated, sandboxed “agentic” pipeline that constructs document-to-repository generation tasks without human annotation. Candidate GitHub projects are first screened for permissive licenses, reproducible unit tests with at least a 90% pass rate and at least 50% coverage, and stable Docker environment builds (Zhao et al., 9 Jun 2026).
The pipeline then proceeds in three stages. In the Divide Phase, two concurrent agents decompose the repository into a high-level overview and a set of capabilities by analyzing code structure, metadata such as pyproject.toml, and runtime traces from unit tests. A small LLM-based judge reconciles these two tracks to map each capability to the precise set of functions and classes it encompasses. This decomposition is explicitly meant to localize documentation effort around coherent functional modules rather than treating the codebase as a monolithic object.
In the Conquer phase, documentation is generated through an iterative Draft–Critic–Repair process applied to each extracted capability . A Draft Agent produces an initial documentation fragment , a Critic Agent identifies missing or under-specified components and emits structured criticism , and a Repair Agent refines the documentation into using the criticism together with code inspection and test feedback. The paper formalizes this loop as
$\begin{aligned} D_i^{(0)} & = f_{\mathrm{draft}\bigl(a_i,\mathcal{A},O,M_i,\mathcal{E}\bigr),\ C_i^{(t)} & = f_{\mathrm{critic}\bigl(D_i^{(t)},M_i^{\mathrm{miss},O,\mathcal{A},a_i,\mathcal{E}\bigr),\ D_i^{(t+1)} & = f_{\mathrm{repair}\bigl(D_i^{(t)},C_i^{(t)},a_i,\mathcal{A},M_i,M_i^{\mathrm{miss},O,\mathcal{E}\bigr). \end{aligned}$
These agents iterate until no further omissions are detected or a maximum iteration budget is reached. The stated function of the critic is to inspect coverage of direct and core-indirect components, structural coherence, and alignment with the test suite, while repair revisits code and tests to fill omissions such as API signatures, default values, or usage examples without leaking test internals (Zhao et al., 9 Jun 2026).
The final stage is Evaluation & Leak Prevention. Before release, all original code, tests, and Git history are stripped; network and package-install commands are blocked; and an LLM-judge audits container logs to ensure that no reference-implementation leakage remains. This leak-prevention stage is central to the claim that each instance remains a genuine generation problem rather than a recovery task from residual artifacts.
3. Dataset composition and statistics
DeNovoSWE contains 4,818 document-to-repository instances, described as approximately 0 NL2Repo and 1 BeyondSWE (Zhao et al., 9 Jun 2026). The dataset is therefore positioned as substantially larger than prior whole-repository resources named in the paper. This suggests an attempt to shift repository generation from a low-data evaluation regime toward a scale more compatible with supervised fine-tuning.
The reported summary statistics over the 4,818 instances are as follows:
| Metric | Mean | Max |
|---|---|---|
| Unit-test Count | 205.0 | 8903 |
| Test Files Count | 31.5 | 8807 |
| Coverage (%) | 85.5 | 100.0 |
| Measured Source Files | 19.3 | 1995 |
The paper also reports percentile values: for Unit-test Count, 2, 3, 4; for Test Files Count, 5, 6, 7; for Coverage, 8, 9, 0; and for Measured Source Files, 1, 2, 3 (Zhao et al., 9 Jun 2026).
In addition, the distribution of repository size, measured by lines exercised by tests, is described as broad enough to include both easy and extremely complex instances. Because the dataset retains unit tests and execution environments while removing source code, the verification target is operational: success is measured by whether generated repositories satisfy the supplied tests in the provided Docker environment.
4. Divide-and-conquer and critic-repair rationale
The divide-and-conquer design is motivated by the claim that asking an agent to document an entire codebase in one pass is brittle. By first extracting a small set of capabilities 4 and profiling the code units 5 relevant to each capability 6, the draft agent can focus on a single feature at a time, reducing instruction complexity (Zhao et al., 9 Jun 2026).
The critic-repair loop is intended to ensure both coverage and restraint. According to the paper, critic passes check direct and core-indirect component coverage, structural coherence such as consistent naming, and alignment with the test suite. Repair then fills omissions while avoiding over-specification and avoiding leakage of test internals. The resulting documentation is characterized as richly detailed and minimally prescriptive. This balance is methodologically important: overly sparse documentation would make the task underdetermined, whereas overly specific documentation could collapse repository generation into template reconstruction.
A common misconception is that automated documentation extraction for training data can simply mirror code structure and therefore produce faithful generation tasks. The DeNovoSWE design explicitly rejects that assumption by inserting adversarial quality control through criticism and repair. Another misconception is that end-to-end repository generation data can be curated by exporting existing repositories and deleting selected files; the leak-prevention stage indicates that the authors regard residual artifacts, package installation, and network access as sources of contamination that must be blocked at construction time (Zhao et al., 9 Jun 2026).
5. Difficulty-aware trajectory filtering
Because whole-repository generation trajectories seldom achieve 100% test pass, DeNovoSWE introduces a per-instance difficulty score 7 used to adaptively filter agent rollouts. The score fuses a structural signal 8, two LLM qualitative signals 9 and 0 each taking values in 1, and an empirical agent pass rate 2, with the pass rate used only for weight fitting (Zhao et al., 9 Jun 2026).
After normalization, the structural and qualitative components are defined as
3
and the final difficulty score is
4
The weights are selected by maximizing the absolute Pearson correlation with pass rate:
5
Once 6 is computed, a rollout for instance 7 is retained only if its pass ratio
8
exceeds a threshold 9 that decreases with difficulty. The reported schedule is:
| Difficulty bin | Threshold 0 |
|---|---|
| 0.0–0.2 | 0.90 |
| 0.2–0.4 | 0.85 |
| 0.4–0.6 | 0.80 |
| 0.6–0.8 | 0.70 |
| 0.8–1.0 | 0.60 |
The stated rationale is that this stratified filtering preserves informative partial successes on hard tasks while enforcing high fidelity on easy ones. An ablation result in the paper supports this claim: a uniform 1 filter yields 48.8% on Doc2Repo, compared with 50.0% for the schedule above (Zhao et al., 9 Jun 2026). This suggests that the filtering strategy is not merely a curation heuristic but an integral part of the training distribution design.
6. Fine-tuning protocol and empirical results
The associated agent model, DeNovoSWE-Agent, is obtained by supervised fine-tuning Qwen3-30B-A3B-Instruct on approximately 11K high-quality trajectories distilled from DeepSeek-V4-Pro rollouts. The reported hyperparameters are a learning rate of 2, batch size 128, warmup ratio 0.05, maximum context length 131,072 tokens, and cosine decay learning-rate scheduling. Training is reported to have been conducted on A100-based GPU clusters over roughly 1B total tokens (Zhao et al., 9 Jun 2026).
During supervised fine-tuning, assistant outputs corresponding to failed tool calls or heredoc fragments are masked out of the loss. This design choice indicates that the target supervision is restricted to valid, task-relevant action traces rather than all emitted text.
Evaluation is performed on BeyondSWE-Doc2Repo, with 50 instances, and NL2Repo-Bench, with 104 instances, using OpenHands/AweAgent in locked Docker environments with no internet or package installation. Performance is measured by full-repo test pass rate averaged over three trials. The reported results are:
| Model | Doc2Repo | NL2Repo |
|---|---|---|
| Qwen3-30B-A3B-Instruct | 5.8 % | 4.3 % |
| Scale-SWE-Agent | 29.2 % | 18.3 % |
| DeNovoSWE-Agent-30A3B | 47.2 % | 23.0 % |
| Qwen3.5-35B-A3B | 43.8 % | 23.5 % |
| DeNovoSWE-Agent-35A3B | 50.0 % | 27.1 % |
The paper states that fine-tuning on DeNovoSWE bridges over 40 percentage points on Qwen3-30B and brings open-weight performance to within approximately 2 percentage points of top proprietary models on BeyondSWE (Zhao et al., 9 Jun 2026). Within the reported setup, the main empirical conclusion is that whole-repository generation improves substantially when the training data are both verifiable and structured around long-horizon execution traces.
7. Limitations, interpretation, and research significance
The paper identifies three explicit limitations. First, language scope is restricted to Python and to repositories with robust unit-test suites. Second, reliance on LLM-based critic and annotators may introduce bias or hallucinations in documentation quality. Third, the pipeline is compute-intensive, which may limit immediate extensibility to other ecosystems (Zhao et al., 9 Jun 2026).
The authors propose several future directions: extending to additional languages such as Java and JavaScript and to multi-language repositories; incorporating human-in-the-loop validation to improve documentation fidelity; leveraging reinforcement learning or self-play to refine agent workflows beyond supervised trajectories; and expanding benchmarks to cover deployment, performance tuning, and CI/CD pipelines for end-to-end SWE agent evaluation (Zhao et al., 9 Jun 2026).
In methodological context, DeNovoSWE can be understood as an attempt to turn repository generation into a verifiable data-generation problem rather than solely a benchmark-design problem. Its emphasis on sandboxing, leak prevention, capability decomposition, critic-repair iteration, and adaptive filtering suggests a broader design principle: long-horizon software engineering data may require explicit control over task decomposition, supervision fidelity, and verification environment to remain trainable at scale. A plausible implication is that the contribution is as much about data-engineering protocol as about dataset size alone.
DeNovoSWE is therefore notable within LLM-based software engineering for defining a concrete operational regime for document-to-repository generation: complete repositories, stripped reference implementations, executable tests, locked environments, and trajectory selection conditioned on estimated task difficulty. Under that regime, the dataset serves both as a training resource and as a template for constructing future long-horizon SWE environments (Zhao et al., 9 Jun 2026).