- The paper introduces MINT, a dataset of 10,278 intent-based misinformation texts aligned with 1,142 reasoning, knowledge, and ethics benchmark samples, and uses controlled experiments to measure misinformation effects in single-agent and multi-agent LLM systems.
- The paper finds that relevant misinformation can reduce single-agent accuracy by up to 26.75%, while three-agent debate lowers overall performance degradation to roughly 2.2%–10.3% but does not eliminate misinformation persistence.
- The paper shows that robustness depends on the model, group composition, and decision protocol, with self-correction rising from 8.0% to 20.5% when uninformed agents become the majority and voting becoming more vulnerable than consensus in some settings.
This paper investigates how intent-based misinformation, injected into the local context of individual agents, affects single-agent prompting and fully benign multi-agent debate (MAD) built on LLMs. Unlike prior work on adversarial persuasion or manipulated agents, all agents here follow the prescribed debate protocol; misinformation enters only through asymmetric context exposure. The authors release MINT (Misinformation INTents), an LLM-generated dataset of 10,278 intent-based misinformation texts aligned to 1,142 samples from three benchmarks, and systematically vary misinformation relevance, group composition, and decision protocol. Their central finding is that misinformation degrades single-agent performance, persists during debate, yet is partially mitigated by multi-agent interaction — with robustness depending jointly on the underlying model, group composition, and the decision-making protocol (2606.16710).
Motivation and positioning
LLM-based multi-agent systems (MAS) are increasingly proposed for high-stakes settings — medical decision support, legal analysis, forensic investigation — where agents exchange intermediate reasoning as evidence during deliberation (2606.16710). Misinformation can enter such systems through retrieval-augmented generation, web search, model hallucinations, or noisy external sources, without any agent acting maliciously. Prior studies of misinformation in LLMs have largely examined isolated models (2606.16710) or adversarial settings in which a manipulated agent deliberately spreads counterfactual or toxic knowledge (2606.16710). Closest to this work, ARGUS and MisinfoTask frame misinformation as an attack on agents' information flows and propose detection-and-correction defenses (2606.16710). This paper instead isolates a cleaner phenomenon: what happens when a misinformed model is one participant among benign peers, and its claims are observed, challenged, or adopted during collective reasoning. This design choice is the paper's main methodological contribution relative to the adversarial literature.
The MINT dataset
MINT augments WinoGrande (reasoning), Complex Web Questions (knowledge), and the commonsense subset of the Ethics benchmark (alignment) with sample-specific misinformation generated by Llama-3.3-70B-Instruct across nine intent-based categories: neutral, clickbait, hoax, rumor, satire, propaganda, framing, conspiracy, and an unconstrained other category, following established misinformation typologies (2606.16710). A two-stage pipeline first generates a neutral "false fact" per sample, then rewrites it into each strategy. An "irrelevant true information" control — Wikipedia passages of comparable token length — distinguishes misinformation effects from generic prompt-length effects.
A three-annotator human audit of 385 texts reports majority-vote faithfulness of 75.8% to the intended false fact and 79.2% to the intended category, but inter-annotator agreement is only fair (Fleiss' κ = 0.24 and 0.26; Krippendorff's α = 0.07). The authors acknowledge that intent-based categories are sometimes ambiguous, particularly when texts combine rhetorical cues from multiple strategies. This is a genuine caveat: category-level analyses (e.g., which strategies persist most) rest on labels with substantial item-level disagreement.
Single-agent vulnerability is task- and relevance-dependent
Relevant misinformation reliably degrades single-agent accuracy, with large relative drops: Llama-3.3 loses 26.75% on CWQ (0.49 to 0.36), 25.71% on Ethics, and 19.51% on WinoGrande. GLM-4.7 shows 16.16% degradation on CWQ and 25.56% on Ethics, but — notably — a 4.61% improvement on WinoGrande under relevant misinformation. Irrelevant misinformation has a substantially weaker effect (e.g., -12.96%, -0.45%, +0.68% for Llama-3.3), and irrelevant true information generally improves performance (up to +27.5% on CWQ for GLM-4.7). The authors conclude that degradation cannot be attributed to longer prompts or generic distraction; harm arises specifically when misinformation is semantically aligned with the task and framed with intent. A practical implication is that open-ended knowledge-intensive QA and ethical judgment are more sensitive to misleading context than commonsense pronoun resolution, so mitigation efforts should prioritize knowledge- and alignment-critical deployments.
In three-agent, five-turn MAD, misinformation-induced degradation is smaller than in the single-agent setting (-2.2% to -10.3% versus -12.9% to -17.2%), which the authors attribute to self-refinement and divergent thinking in iterative debate. Even on CWQ, where MAD does not improve uninformed performance, it still mitigates misinformation effects — a result with direct relevance for system design, since it suggests MAD can serve as a more resilient reasoning interface even where it adds no baseline capability.
However, mitigation at the system level does not mean misinformation disappears at the agent level. Measuring opinion persistence (the probability that an answer proposed at turn t is repeated at turn t+1), the authors find that Llama-3.3 agents retain answers from misinformed peers more often than answers from uninformed peers, with average deltas of -10.4% on CWQ and -7.7% on WinoGrande. Persistence varies by strategy: hoaxes and unconstrained misinformation are among the most persistent; framing and rumors less so. On Ethics, agents retain correct judgments more often (+1.3% delta). Critically, this strong persistence is largely absent for GLM-4.7 (deltas between -0.6% and +0.3%). Two confounds are acknowledged: the two models come from different training and alignment regimes, and Llama-3.3 both generated and consumed the misinformation — a deliberately included same-model condition, motivated by evidence that LLM evaluators favor their own outputs (2606.16710). The stronger Llama-3.3 persistence should therefore be read as a plausible same-model scenario, not a model-independent estimate of misinformation persuasiveness. This is an honest but important qualification on the paper's headline persistence numbers.
Group composition and decision protocol determine robustness
Scaling to five agents, the paper contrasts voting (majority over final answers) with consensus (the last agent decides based on the transcript). For Llama-3.3 on WinoGrande, voting achieves higher absolute accuracy (0.938 with no misinformed agents) but degrades as misinformation spreads (0.857 with five misinformed agents, a 0.081 drop), while consensus remains nearly flat (0.729 to 0.758). The voting advantage over consensus consequently shrinks from 0.208 to 0.099. For GLM-4.7, the trade-off largely disappears: both protocols remain stable, with gaps of only 0.008 to 0.025. The authors conclude that robustness to misinformed peer pressure is a property of the model–protocol combination, not of the protocol alone. They further note that this pattern differs from human peer-pressure findings, including Asch-style conformity with robots, and speculate that agents' awareness that peers are LLMs may weaken conformity effects — a hypothesis the paper does not test, since all agents know they are interacting with LLMs.
Self-correction under peer pressure shows a sharp, non-linear threshold: the rate at which an initially misinformed agent revises to the correct answer by turn 5 jumps from 8.0% with two uninformed agents to 20.5% with three — that is, correction emerges only once uninformed agents form a majority. The design implication is a compute-for-robustness trade-off: adding agents raises cost but reduces the chance that misinformed agents dominate. Supplementary results reinforce this: even fully misinformed MAD on WinoGrande degrades less (-8.7%) than the misinformed single-agent setup (-17.2%), suggesting debate provides some mitigation even when every agent receives misleading context, though performance declines as misinformation becomes universal.
Limitations
The paper is explicit about the boundaries of its claims. It evaluates only two open-weight models, so the model-dependence of robustness is demonstrated but not mapped across model families. MINT's misinformation is machine-generated and may differ from human-written, retrieved, or adversarially optimized misinformation; the same-model generation condition, while realistic, may inflate model-specific persistence effects. The MAS design space is restricted to fixed debate structures, fixed turn counts, and two decision protocols — no tool use, memory, retrieval, dynamic roles, or source verification. Finally, the fair inter-annotator agreement on MINT means intent-category boundaries are themselves uncertain, which limits the resolution of strategy-level conclusions.
Conclusion
This paper provides controlled evidence that contextual misinformation harms benign LLM agents, persists through multi-agent debate at the level of individual agent answers, and yet is partially absorbed by collective reasoning — particularly when uninformed agents constitute a majority. Its most actionable findings are the voting-versus-consensus trade-off under misinformed peer pressure and the threshold effect in self-correction, both of which depend on the underlying model. The work leaves open whether these dynamics hold for human-written or adversarially optimized misinformation, for larger and heterogeneous agent networks, for protocols in which agents do not know their interlocutors are LLMs, and for comparisons against Chain-of-Thought, Self-Refine, and drift-monitoring methods.