Papers
Topics
Authors
Recent
Search
2000 character limit reached

SEMAG: Self-Evolutionary Multi-Agent Code Generation

Published 16 Mar 2026 in cs.SE and cs.AI | (2603.15707v1)

Abstract: LLMs have made significant progress in handling complex programming tasks. However, current methods rely on manual model selection and fixed workflows, which limit their ability to adapt to changing task complexities. To address this, we propose SEMAG, a Self-Evolutionary Multi-Agent code Generation framework that mimics human coding practices. It decomposes programming tasks into stages, including planning, coding, debugging, and discussion, while adapting workflows to task difficulty. Its self-evolutionary agents can access the latest models in real time and automatically upgrade the backbone model. SEMAG sets new state-of-the-art Pass@1 accuracy across benchmarks. Using identical backbone models, SEMAG outperforms prior methods by 3.3% on CodeContests. When augmented with self-evolutionary model selection that automatically identifies optimal backbones, SEMAG reaches 52.6%, showcasing both framework effectiveness and adaptability to evolving LLM capabilities.

Summary

  • The paper introduces SEMAG, a hierarchical multi-agent framework that escalates planning, trace-guided debugging, and debate based on task difficulty rather than using a fixed workflow.
  • SEMAG raises GPT-4o Pass@1 over LPW from 34.7% to 38.0% on CodeContests and achieves leading results across seven benchmarks while reducing average token use by up to 19.3%.
  • The paper’s self-evolution module selects stronger backbones through web retrieval and performance-weighted voting, reaching 52.6% on CodeContests with Claude-3.7-Sonnet, though retrieval depth, latency, and generalization remain limitations.

Overview

SEMAG is a multi-agent framework for text-to-code generation that addresses three structural weaknesses of prior agentic systems: fixed reasoning depth that wastes compute on easy tasks and under-solves hard ones, single-pass debugging loops prone to local optima, and tight coupling to one static backbone LLM (2603.15707). The framework combines a four-level hierarchical workflow whose depth scales with measured task difficulty, a discussion–decision phase that aggregates multiple reasoning trajectories, and a "self-evolution" module in which selector agents crawl recent web information to autonomously pick the best backbone model. Under a controlled backbone comparison against GPT-4o, SEMAG improves Pass@1 on CodeContests from 34.7% to 38.0% (+3.3 points) over LPW, the prior state-of-the-art planning-driven workflow; with self-evolutionary model selection identifying Claude-3.7-Sonnet as the backbone, it reaches 52.6% on CodeContests.

Motivation and positioning

The paper identifies three concrete failure modes in existing pipelines. Fixed-depth debuggers such as Self-Debugging and LDB apply the same workflow regardless of difficulty, producing redundant token expenditure on trivial problems and insufficient refinement on hard ones (Chen et al., 2023, Lei et al., 2024). Single-iteration debugging degrades sharply when initial outputs diverge substantially from intent — a regime where Chain-of-Thought or Tree-of-Thoughts-style methods lack an explicit aggregation phase over competing solutions (Kojima et al., 2022, Yao et al., 2023). Finally, frameworks built on a single commercial model cannot exploit newly released backbones without manual re-engineering. SEMAG is positioned relative to Self-Planning, MapCoder, LDB, and LPW, with LPW serving as the strongest baseline throughout (Lei et al., 2024).

Method

The framework formalizes code synthesis as operations over problem descriptions PP, sample I/O pairs SS, plan space Π\Pi, and program space C\mathcal{C}, instantiated as specialized agents: PLANNER, VERIFIER, CODER, DEBUGGER, EMBEDTRACE, EXPLAINER, SUGGESTOR, plus DEBATER/DISCIMINATOR agents at the top level.

Hierarchical escalation. Level 1 attempts direct generation with minimal prompting. On failure, Level 2 produces a structured plan π\pi refined through up to MplanM_{\text{plan}} verifier rounds that simulate execution on the provided examples before any code is written. Level 3 performs trace-guided debugging across up to five attempts (Mtry=5M_{\text{try}}=5) and four debugging iterations (Mdebug=4M_{\text{debug}}=4): EMBEDTRACE instruments runtime variable states, EXPLAINER supplies semantic analysis, SUGGESTOR synthesizes targeted fixes, and DEBUGGER applies them. Level 4 activates only when debugging stalls: NdebaterN_{\text{debater}} agents propose alternative solutions conditioned on accumulated discussion history, and a weighted-consensus discriminator selects among them using softmax weights over historical agent performance.

Adaptive transition. Rather than fixed iteration counts, level transitions are gated by similarity between successive execution traces, measured as normalized edit distance:

ρ(τt,τt−1)=1−EditDist(τt,τt−1)max⁡(∣τt∣,∣τt−1∣)\rho(\tau_t, \tau_{t-1}) = 1 - \frac{\text{EditDist}(\tau_t, \tau_{t-1})}{\max(|\tau_t|, |\tau_{t-1}|)}

against a threshold SS0 with SS1, SS2. Stagnant traces trigger escalation earlier on harder tasks; improving traces keep the system at its current level. The paper does not validate this threshold design independently of end-to-end accuracy, so its contribution versus simpler heuristics remains unquantified.

Self-evolution. Parallel selector agents generate task-specific keywords, retrieve and filter recent web pages (30-day window), summarize evidence, and vote — weighted by sampled performance on a small task subset — over candidate models. This module is optional: all headline comparisons hold with a fixed backbone.

Empirical results

Across seven benchmarks (HumanEval, MBPP, their edge-test ET variants, APPS, LiveCodeBench, CodeContests), SEMAG sets new best Pass@1 results under GPT-4o (2024-05-13):

Benchmark Direct LDB LPW SEMAG
HumanEval 91.5 92.1 98.2 98.8
MBPP 62.8 82.4 84.8 87.6
HumanEval-ET 79.3 81.7 84.8 86.6
MBPP-ET 51.0 65.4 65.8 71.8
APPS 47.5 53.2 62.6 67.6
LiveCode 46.4 54.3 59.3 65.0
CodeContests 24.6 29.3 34.7 38.0

With GPT-3.5, gains over LPW are smaller on HumanEval/MBPP (+2.5/+0.2) but larger on edge-case variants (+2.5/+6.8), indicating the framework's advantage concentrates on problems requiring robust handling of hidden tests. Per-difficulty appendix analysis shows one notable reversal: on APPS Competition-level problems, SEMAG scores 32.6% versus LPW's 34.8% — the authors attribute this to hierarchical prompting succeeding on visible tests while failing hidden ones — so the aggregate superiority on APPS does not extend uniformly to its hardest tier. On LiveCode, SEMAG leads at every difficulty tier, with a 12.7-point margin over LPW at Medium.

Self-evolution outcome. Deployed on CodeContests, the selectors autonomously identified Claude-3.7-Sonnet, GPT-4.1, and DeepSeek-v3 as candidates; Claude-3.7-Sonnet achieved 52.6% Pass@1, well above GPT-4o's 38.0%, with GPT-4.1 and DeepSeek-v3 both at 48.7%. Crawl-depth ablation shows the mechanism's sensitivity: with fewer than 20 retrieved pages, the probability of surfacing Claude-3.7-Sonnet in the Top-3 drops to 40–60%, risking suboptimal selection; SS3 saturates discovery probability at 80% for roughly 46k tokens and six minutes, while deeper crawls add no benefit and inflate cost by 30–55%.

Efficiency and ablations. The hierarchical controller reduces average tokens per problem relative to LPW by 19.3% (HumanEval) and 15.5% (MBPP), shrinking to 9.3%/5.1% on APPS/CodeContests where Level 4 dominates. Component ablation on HumanEval with GPT-3.5 shows strong synergy effects: individual components yield 77.4–81.7% versus the 71.9% baseline, pairs reach 82.9–83.5%, but the full system reaches 91.5% — no pair configuration comes within 8 points, implying non-additive interactions. Tool use during planning contributes 3.7 points (91.5% vs. 87.8%). A grid search over SS4 shows monotone improvement to a plateau near SS5; the maximum observed accuracy (92.1% at SS6) differs marginally, and the chosen operating point is justified on cost grounds. Temperature sweeps favor SS7 for variance control.

Limitations and open questions

The authors are explicit about several constraints. The hyperparameters SS8 were tuned via grid search on HumanEval and fixed globally; adaptive tuning across benchmarks is left open. The iterative refinement pipeline increases latency, which is problematic for time-sensitive deployment despite the token savings on simple tasks. The self-evolution module depends on live web retrieval, inheriting incompleteness and bias from search ranking and recency effects — evidenced by up to 60% missed detection of the optimal backbone at shallow crawl depths — and offline recommendation alternatives remain unexplored. The adaptive-threshold mechanism itself is never ablated in isolation. Finally, the paper notes the standard caveat that machine-generated code should be executed in sandboxes.

Conclusion

SEMAG demonstrates that difficulty-adaptive escalation, explicit debate-and-decide aggregation, and automated backbone switching can be composed into a code-generation system that simultaneously improves accuracy and reduces token overhead relative to fixed-workflow baselines. Its strongest claims rest on consistent state-of-the-art Pass@1 across seven benchmarks under a controlled backbone, and on the observation that autonomous model selection alone contributed +14.6 points on CodeContests — more than the entire architectural contribution over LPW. Whether the self-evolution mechanism remains reliable when web coverage of new models is sparse, and whether the adaptive transition thresholds generalize beyond the tested settings, are questions the paper leaves open.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 1 tweet with 2 likes about this paper.