- The paper introduces SEMAG, a hierarchical multi-agent framework that escalates planning, trace-guided debugging, and debate based on task difficulty rather than using a fixed workflow.
- SEMAG raises GPT-4o Pass@1 over LPW from 34.7% to 38.0% on CodeContests and achieves leading results across seven benchmarks while reducing average token use by up to 19.3%.
- The paper’s self-evolution module selects stronger backbones through web retrieval and performance-weighted voting, reaching 52.6% on CodeContests with Claude-3.7-Sonnet, though retrieval depth, latency, and generalization remain limitations.
Overview
SEMAG is a multi-agent framework for text-to-code generation that addresses three structural weaknesses of prior agentic systems: fixed reasoning depth that wastes compute on easy tasks and under-solves hard ones, single-pass debugging loops prone to local optima, and tight coupling to one static backbone LLM (2603.15707). The framework combines a four-level hierarchical workflow whose depth scales with measured task difficulty, a discussion–decision phase that aggregates multiple reasoning trajectories, and a "self-evolution" module in which selector agents crawl recent web information to autonomously pick the best backbone model. Under a controlled backbone comparison against GPT-4o, SEMAG improves Pass@1 on CodeContests from 34.7% to 38.0% (+3.3 points) over LPW, the prior state-of-the-art planning-driven workflow; with self-evolutionary model selection identifying Claude-3.7-Sonnet as the backbone, it reaches 52.6% on CodeContests.
Motivation and positioning
The paper identifies three concrete failure modes in existing pipelines. Fixed-depth debuggers such as Self-Debugging and LDB apply the same workflow regardless of difficulty, producing redundant token expenditure on trivial problems and insufficient refinement on hard ones (Chen et al., 2023, Lei et al., 2024). Single-iteration debugging degrades sharply when initial outputs diverge substantially from intent — a regime where Chain-of-Thought or Tree-of-Thoughts-style methods lack an explicit aggregation phase over competing solutions (Kojima et al., 2022, Yao et al., 2023). Finally, frameworks built on a single commercial model cannot exploit newly released backbones without manual re-engineering. SEMAG is positioned relative to Self-Planning, MapCoder, LDB, and LPW, with LPW serving as the strongest baseline throughout (Lei et al., 2024).
Method
The framework formalizes code synthesis as operations over problem descriptions P, sample I/O pairs S, plan space Π, and program space C, instantiated as specialized agents: PLANNER, VERIFIER, CODER, DEBUGGER, EMBEDTRACE, EXPLAINER, SUGGESTOR, plus DEBATER/DISCIMINATOR agents at the top level.
Hierarchical escalation. Level 1 attempts direct generation with minimal prompting. On failure, Level 2 produces a structured plan π refined through up to Mplan verifier rounds that simulate execution on the provided examples before any code is written. Level 3 performs trace-guided debugging across up to five attempts (Mtry=5) and four debugging iterations (Mdebug=4): EMBEDTRACE instruments runtime variable states, EXPLAINER supplies semantic analysis, SUGGESTOR synthesizes targeted fixes, and DEBUGGER applies them. Level 4 activates only when debugging stalls: Ndebater agents propose alternative solutions conditioned on accumulated discussion history, and a weighted-consensus discriminator selects among them using softmax weights over historical agent performance.
Adaptive transition. Rather than fixed iteration counts, level transitions are gated by similarity between successive execution traces, measured as normalized edit distance:
ρ(τt,τt−1)=1−max(∣τt∣,∣τt−1∣)EditDist(τt,τt−1)
against a threshold S0 with S1, S2. Stagnant traces trigger escalation earlier on harder tasks; improving traces keep the system at its current level. The paper does not validate this threshold design independently of end-to-end accuracy, so its contribution versus simpler heuristics remains unquantified.
Self-evolution. Parallel selector agents generate task-specific keywords, retrieve and filter recent web pages (30-day window), summarize evidence, and vote — weighted by sampled performance on a small task subset — over candidate models. This module is optional: all headline comparisons hold with a fixed backbone.
Empirical results
Across seven benchmarks (HumanEval, MBPP, their edge-test ET variants, APPS, LiveCodeBench, CodeContests), SEMAG sets new best Pass@1 results under GPT-4o (2024-05-13):
| Benchmark |
Direct |
LDB |
LPW |
SEMAG |
| HumanEval |
91.5 |
92.1 |
98.2 |
98.8 |
| MBPP |
62.8 |
82.4 |
84.8 |
87.6 |
| HumanEval-ET |
79.3 |
81.7 |
84.8 |
86.6 |
| MBPP-ET |
51.0 |
65.4 |
65.8 |
71.8 |
| APPS |
47.5 |
53.2 |
62.6 |
67.6 |
| LiveCode |
46.4 |
54.3 |
59.3 |
65.0 |
| CodeContests |
24.6 |
29.3 |
34.7 |
38.0 |
With GPT-3.5, gains over LPW are smaller on HumanEval/MBPP (+2.5/+0.2) but larger on edge-case variants (+2.5/+6.8), indicating the framework's advantage concentrates on problems requiring robust handling of hidden tests. Per-difficulty appendix analysis shows one notable reversal: on APPS Competition-level problems, SEMAG scores 32.6% versus LPW's 34.8% — the authors attribute this to hierarchical prompting succeeding on visible tests while failing hidden ones — so the aggregate superiority on APPS does not extend uniformly to its hardest tier. On LiveCode, SEMAG leads at every difficulty tier, with a 12.7-point margin over LPW at Medium.
Self-evolution outcome. Deployed on CodeContests, the selectors autonomously identified Claude-3.7-Sonnet, GPT-4.1, and DeepSeek-v3 as candidates; Claude-3.7-Sonnet achieved 52.6% Pass@1, well above GPT-4o's 38.0%, with GPT-4.1 and DeepSeek-v3 both at 48.7%. Crawl-depth ablation shows the mechanism's sensitivity: with fewer than 20 retrieved pages, the probability of surfacing Claude-3.7-Sonnet in the Top-3 drops to 40–60%, risking suboptimal selection; S3 saturates discovery probability at 80% for roughly 46k tokens and six minutes, while deeper crawls add no benefit and inflate cost by 30–55%.
Efficiency and ablations. The hierarchical controller reduces average tokens per problem relative to LPW by 19.3% (HumanEval) and 15.5% (MBPP), shrinking to 9.3%/5.1% on APPS/CodeContests where Level 4 dominates. Component ablation on HumanEval with GPT-3.5 shows strong synergy effects: individual components yield 77.4–81.7% versus the 71.9% baseline, pairs reach 82.9–83.5%, but the full system reaches 91.5% — no pair configuration comes within 8 points, implying non-additive interactions. Tool use during planning contributes 3.7 points (91.5% vs. 87.8%). A grid search over S4 shows monotone improvement to a plateau near S5; the maximum observed accuracy (92.1% at S6) differs marginally, and the chosen operating point is justified on cost grounds. Temperature sweeps favor S7 for variance control.
Limitations and open questions
The authors are explicit about several constraints. The hyperparameters S8 were tuned via grid search on HumanEval and fixed globally; adaptive tuning across benchmarks is left open. The iterative refinement pipeline increases latency, which is problematic for time-sensitive deployment despite the token savings on simple tasks. The self-evolution module depends on live web retrieval, inheriting incompleteness and bias from search ranking and recency effects — evidenced by up to 60% missed detection of the optimal backbone at shallow crawl depths — and offline recommendation alternatives remain unexplored. The adaptive-threshold mechanism itself is never ablated in isolation. Finally, the paper notes the standard caveat that machine-generated code should be executed in sandboxes.
Conclusion
SEMAG demonstrates that difficulty-adaptive escalation, explicit debate-and-decide aggregation, and automated backbone switching can be composed into a code-generation system that simultaneously improves accuracy and reduces token overhead relative to fixed-workflow baselines. Its strongest claims rest on consistent state-of-the-art Pass@1 across seven benchmarks under a controlled backbone, and on the observation that autonomous model selection alone contributed +14.6 points on CodeContests — more than the entire architectural contribution over LPW. Whether the self-evolution mechanism remains reliable when web coverage of new models is sparse, and whether the adaptive transition thresholds generalize beyond the tested settings, are questions the paper leaves open.