When multi-agent coordination outperforms single-agent tool use

Ascertain the conditions under which language-model-based multi-agent coordination provides value over single strong language models equipped with tool use, identifying the task properties and architectural configurations that yield multi-agent advantages relative to single-agent baselines.

Background

The authors survey prior claims that multi-agent collaboration universally improves performance and note mixed findings, including reports that benefits diminish with stronger base models. They highlight the lack of a principled framework for predicting when multi-agent coordination is advantageous.

They explicitly state that determining when multi-agent coordination provides value over single strong models with tool use remains empirically open, motivating their controlled evaluation and scaling principles.

References

The question of when multi-agent coordination provides value over single strong models with tool use remains empirically open, with \citet{qian2024scaling}'s proposed scaling laws showing no significant universal pattern \citep{wang2024survey}, motivating our systematic evaluation.

— Towards a Science of Scaling Agent Systems  (2512.08296 - Kim et al., 9 Dec 2025) in Related Work, Multi-Agent Systems (MAS) versus Single-Agent Systems (SAS)

The sample is too small to rule out substantively meaningful effects; an advantage of a third of a round is inside that interval.

— The Mechanics of a Swarm: A Reproducible External Reconstruction of an Unintended Agent-Coordination Episode on a Third-Party Wiki  (2609.12748 - Lütje, 11 Sep 2026) in Section 3.8, “The trade, and its bounded outcome evidence,” paragraph “No robust positive association with documented progress”

The multi-agent design draws on cognitive science's searcher-evaluator-generator model; the swarm was not probed at inference in this work, so whether joint training unlocks a multi-agent advantage is an open question.

— Training-Free Inference-Time Self-Reflection and Cost-Bounded Early Stopping for Large Language Models  (2608.18884 - Yu et al., 19 Aug 2026) in Section 6, Discussion, subsection “Implications”

The two profiles are close to complementary, and since the score is identical across modes, whether a human--agent pairing beats either alone, and which capability the gain comes from, is measurable. We leave that measurement to future work.

— FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents  (2608.18423 - Wang et al., 19 Aug 2026) in Section 3.1, subsection “Human study”

Whether communication produces similar gains when feedback is sparse, noisy, or subjective remains an important open question.

— Scaling Discovery through Test-Time Communication  (2609.21032 - Park et al., 17 Sep 2026) in Section Conclusion and Future Directions