When multi-agent coordination outperforms single-agent tool use
Ascertain the conditions under which language-model-based multi-agent coordination provides value over single strong language models equipped with tool use, identifying the task properties and architectural configurations that yield multi-agent advantages relative to single-agent baselines.
References
The question of when multi-agent coordination provides value over single strong models with tool use remains empirically open, with \citet{qian2024scaling}'s proposed scaling laws showing no significant universal pattern \citep{wang2024survey}, motivating our systematic evaluation.
The sample is too small to rule out substantively meaningful effects; an advantage of a third of a round is inside that interval.
The multi-agent design draws on cognitive science's searcher-evaluator-generator model; the swarm was not probed at inference in this work, so whether joint training unlocks a multi-agent advantage is an open question.
The two profiles are close to complementary, and since the score is identical across modes, whether a human--agent pairing beats either alone, and which capability the gain comes from, is measurable. We leave that measurement to future work.
Whether communication produces similar gains when feedback is sparse, noisy, or subjective remains an important open question.