Does extended DGM runtime surpass closed-source SWE-bench systems?
Determine whether increasing the number of iterations and compute allocated to the Darwin Gödel Machine—an open-ended, self-improving system that iteratively modifies its own code to design LLM-based coding agents—continues to yield further performance gains on SWE-bench and can exceed the performance of closed-source state-of-the-art SWE-bench systems.
References
However, it still falls short of closed-source SoTA SWE-bench solutions. An open question is whether running the DGM for longer would continue to yield performance gains and eventually surpass closed-source solutions.
The results also reflect a single pass of the evaluation--selection--update loop, so whether the gains a model realizes from being used in the harness can continue to accumulate across successive iterations remains to be tested.
The development environment $D$ is held fixed across both stages; whether an evolved harness can itself serve as the development environment for further evolution is left to future work.