---
title: 'MAD-Spear: Attack on Multi-Agent Debate Systems'
url: https://www.emergentmind.com/topics/mad-spear
type: topic
---

# MAD-Spear: Attack on Multi-Agent Debate Systems

Searching arXiv for the MAD-Spear paper and referenced related work.
Multi-agent debate (MAD) systems are LLM-based collective reasoning architectures in which multiple agents iteratively exchange answers and revise their outputs before a final consensus is selected. MAD-Spear is a targeted prompt injection attack against such systems that compromises only a small subset of agents yet substantially degrades collective performance by exploiting conformity in inter-agent deliberation. The attack works by inducing compromised agents to emit plausible but incorrect reasoning traces and to simulate additional “Sybil” outputs, thereby shifting the apparent majority and increasing the likelihood that benign agents adopt false answers. The framework is introduced together with a formal notion of MAD fault-tolerance and an evaluation methodology spanning accuracy, consensus efficiency, and scalability [2507.13038].

## 1. Multi-agent debate as a conformity-sensitive reasoning architecture

A MAD system consists of $N$ LLM-based agents $AS=\{a_0,a_1,\ldots,a_{N-1}\}$. For each query $q$, all agents first generate independent initial answers $o_i$ in Round 0. In each subsequent debate round $r=1\ldots \Delta R$, every agent $a_i$ receives peer outputs $\{o_j\}_{j\ne i}$, incorporates them into its prompt, and produces a revised answer. After $\Delta R$ rounds, final consensus is reached through a voting or confidence-weighted selection mechanism, with SoM’s assessment algorithm cited as an example in the source formulation [2507.13038].

The interaction protocol may be homogeneous or heterogeneous. In the homogeneous case, agents share a common system prompt to enforce aligned behavior; in the heterogeneous case, they may differ in LLM choice or configuration. Communication topology also varies. In synchronous MAD, each agent observes all $N-1$ peer messages in a round, whereas in sparse MAD, it observes only $N-u$ messages. These design choices affect information flow, convergence behavior, and the exposure surface for adversarial manipulation [2507.13038].

A central property of the architecture is conformity, described as a function of peer pressure and interaction time. In benign settings, conformity can improve aggregate accuracy by driving weaker agents toward the majority. The same mechanism also creates an attack surface: once an adversary can alter the apparent majority, benign agents may propagate misinformation through ordinary update dynamics rather than through explicit model compromise. This suggests that MAD robustness depends not only on the quality of individual agents but also on the stability of the social dynamics induced by debate [2507.13038].

## 2. Threat model and attack construction

The threat model assumes that an adversary can inject malicious content into the external data $D_i$ of up to $t \le \lfloor (N-1)/P \rfloor$ agents, where $P \ge 3$. The attacker’s objective is to minimize $t$ while converting a finite MAD process—one guaranteed to converge within $\Delta R$ rounds—into either an infinite one or one that converges incorrectly. The attack therefore targets system-level consensus rather than only local answer corruption [2507.13038].

The adversary selects a subset of high-capability agents $a_{x_1},\ldots,a_{x_t}$ for compromise. For each selected agent, a template prompt $p(L)$ is injected. This template has three specified effects: it instructs the compromised agent to ignore honest peers, forces it to output a predefined incorrect reasoning trace, and spawns $L$ “Sybil” pseudo-agents by repeating the malicious output under different “One agent solution:” prefixes. The use of high-capability agents is significant because their outputs are more likely to appear credible to other participants in the debate [2507.13038].

The resulting manipulated agents broadcast plausible but incorrect answers with high confidence. Benign agents are modeled as exhibiting conformity probability $\lambda$; under this dynamic they defect by selecting one of the majority outputs with probability at least $\lambda$. MAD-Spear therefore does not merely inject false content into the transcript. It weaponizes the endogenous consensus mechanism of MAD, causing false answers to diffuse because they appear socially validated within the debate [2507.13038].

## 3. Algorithmic mechanism of MAD-Spear

The attack algorithm is specified in four steps. First, the attacker selects $t \le \lfloor (N-1)/3 \rfloor$ agents to compromise. Second, for each compromised agent $a_x$, the adversary injects external data via
$$
D_x^s \leftarrow \operatorname{Inject}(D_x,p(L)).
$$
Third, compromised agents generate expanded malicious output
$$
o_1^s\|\cdots\|o_L^s\|\delta \leftarrow f_x(q\|D_x^s,\{o_i\}),
$$
where $\delta$ instructs peers to adopt the Sybil outputs. Fourth, each honest agent $a_y$ receives the augmented set $\{o_i\}\cup\{o_{N+1}^s,\ldots,o_{N+L}^s\}\|\delta$ and produces an updated answer [2507.13038].

The key analytical device is the transformation of the balance between correct and anomalous agents. Let $AS_m$ and $AS_a$ denote the correct and anomalous agent sets after Round 0, with $|AS_m|>|AS_a|$. By appending $L$ Sybil agents,
$$
|AS_a|' = |AS_a| + L,\quad e = |AS_m| - |AS_a|' < 0.
$$
The attacker chooses $L$ so as to maximize the probability of incorrect convergence under conformity constraints:
$$
\underset{x\in\mathcal{C}}{\operatorname{arg\,max}\;\Pr(\text{incorrect}\mid \text{conformity constraints})}
=
\underset{x\in\mathcal{C}}{\operatorname{arg\,max}\;\lambda\cdot \mathbf{1}_{\{x\in \text{Sybil outputs}\}}.
$$
In effect, the Sybil mechanism converts a minority of compromised agents into an apparent majority at the discourse level, even when the underlying number of physically compromised agents remains small [2507.13038].

A plausible implication is that MAD-Spear is best understood as a discourse-layer amplification attack. The direct compromise budget is bounded, but the message budget is not: by multiplying outputs and attaching an instruction $\delta$, the attack alters both the count and the salience of candidate answers presented to benign agents.

## 4. Fault-tolerance formalization

To assess robustness, the framework defines MAD fault-tolerance in terms of the partition of agents into correct and faulty subsets after initial answering. For a problem $q$ and agent set $AS$ with $N=|AS|$, let
$$
AS = AS_m \cup AS_a,
$$
where $AS_m$ gives the correct answer and $AS_a$ is faulty. The tolerance factor is
$$
e = |AS_m| - |AS_a| \ge 0.
$$
This quantity measures the margin by which correct agents outnumber faulty ones at the outset [2507.13038].

A MAD system $\mathcal{M}$ is defined to be $\alpha$-fault-tolerant if it can tolerate up to $\alpha$ faulty agents and still converge correctly. Formally,
$$
FT_{\alpha}(\mathcal{M})
=
\min\{\,k : \forall\,A\subseteq AS,\ |A|\le k \implies \text{consensus}_{\mathcal{M}}(q,A)\text{ is correct}\}.
$$
The interpretation given is that $FT_\alpha(\mathcal{M})$ is the smallest number $k$ of simultaneous faulty agents that breaks correct convergence. Larger $FT_\alpha$ indicates stronger robustness; equivalently, larger $e$ and fewer rounds $\Delta R$ improve fault tolerance [2507.13038].

This formalization situates MAD-Spear within a systems perspective rather than a prompt-security perspective alone. The relevant question is not merely whether an individual agent can be induced to misbehave, but how many faulty agents—or effective faulty outputs—the debate protocol can absorb before consensus is destabilized. Because MAD-Spear targets the sign and magnitude of $e$, it directly attacks the condition underlying correct convergence [2507.13038].

## 5. Evaluation methodology and empirical findings

The evaluation framework jointly measures accuracy, scalability, and consensus efficiency. Accuracy is operationalized via Attack Success Rate (ASR), defined as $1 - (\text{Accuracy of MAD under attack})$, using whether the final MAD answer matches ground truth. Scalability is measured through token consumption,
$$
\text{TC}=\sum_{r=0}^{\Delta R-1}\sum_{i=0}^{N-1} OT_i^r,
$$
where $OT_i^r$ denotes the number of output tokens of agent $a_i$ in round $r$. Consensus efficiency is measured by rounds to convergence $\Delta R$, with the attacker seeking to drive $\Delta R \to \infty$. An optional metric, ConsensusEfficiency, is also defined as
$$
\text{ConsensusEfficiency}
=
\frac{T_{\text{baseline}}-T_{\text{attack}}}{T_{\text{baseline}}}.
$$
The benchmark suite includes GSM-Ranges Levels 3–6 and the Logical Fallacies subset of MMLU [2507.13038].

The reported experiments show that MAD-Spear consistently exceeds the baseline attack in degrading system performance. On Levels 3–4, average ASR rises from 6.67% for the baseline to 56.66% for MAD-Spear, while average token consumption increases from 26,959 tokens to 85,101.5 tokens. Accuracy under MAD-Spear drops to 26.67% on Level 4 from 100%, and token consumption exhibits a $3\times$ overhead on the hardest tasks. Convergence probability declines as the number of rounds increases, indicating that larger $\Delta R$ can push a finite MAD process toward an infinite one [2507.13038].

| Measure | Baseline | MAD-Spear |
|---|---:|---:|
| Avg ASR on Levels 3–4 | 6.67% | 56.66% |
| Avg TC on Levels 3–4 | 26,959 tokens | 85,101.5 tokens |

These results indicate a threefold degradation pattern: correctness falls, computation cost rises, and termination reliability weakens. The combination is notable because attacks on reasoning systems are often evaluated only through final-answer accuracy; here the framework treats scalability and convergence as first-class security properties [2507.13038].

## 6. Agent diversity, attack composition, and operational implications

The experiments also report that agent diversity substantially improves performance in mathematical reasoning tasks. On GSM-Ranges Level 3, heterogeneous MAD attains 93.33% accuracy versus 60.00% for homogeneous MAD, a 56% increase, with statistical significance reported as $p<0.01$ via paired $t$-test. This finding is presented as challenging prior work that suggested agent diversity has minimal impact on performance [2507.13038].

MAD-Spear is further described as composable with communication attacks. In the composite setting, message-loss $C$ reduces the count of correct messages:
$$
|AS_m|'=|AS_m|-C,\quad e = |AS_m|' - |AS_a|' \ll 0.
$$
Sybil messages exactly replace lost honest messages, maximizing confusion. In this formulation, communication disruption and prompt injection act on the same robustness margin $e$ from opposite directions: the former subtracts correct influence, while the latter adds anomalous influence [2507.13038].

A plausible implication is that diversity may provide partial resilience by reducing correlated conformity behavior, although the source material states the performance effect rather than an explicit causal defense mechanism. By contrast, the compositional attack analysis makes clear that robustness cannot be reduced to one layer of protection; both message transport and prompt integrity shape whether the debate process remains fault-tolerant [2507.13038].

## 7. Security interpretation and mitigation directions

The security significance of MAD-Spear lies in its exploitation of ordinary debate dynamics rather than an exotic failure mode. The attack compromises few agents, but by simulating additional voices and leveraging conformity probability $\lambda$, it can drastically reduce fault tolerance, lower accuracy, increase token consumption, and prevent timely convergence. The paper therefore frames security in MAD design as an urgent systems problem rather than a peripheral concern [2507.13038].

The recommended defenses include log-analysis and automated failure attribution to detect compromised agents, G-Safeguard to filter out Sybil-pattern outputs, limiting peer message format variance such as “One agent solution:” prefixes, and enforcing cryptographic signatures per agent. These measures target different stages of the attack pipeline: post hoc diagnosis, message-level filtering, prompt-format hardening, and provenance assurance [2507.13038].

One common misconception is that the security of MAD can be inferred directly from the security of its constituent LLMs. The MAD-Spear analysis indicates otherwise. Even when only a bounded subset of agents is compromised, the consensus process can amplify local faults into global failure through conformity-driven propagation. Another misconception is that more debate rounds necessarily improve deliberation; the reported convergence results indicate that increased $\Delta R$ can instead reduce correct consensus under attack and drive finite MAD toward infinite behavior [2507.13038].

Source: https://www.emergentmind.com/topics/mad-spear