Belief Coevolution in a Social Network of Generalist and Specialist Large Language Models
This presentation examines how beliefs diffuse and evolve when populations of language model agents interact within social networks. Using a novel simulation framework called CoevolveSim, researchers systematically isolated the effects of model specialization, social role assignment, and network structure on collective belief formation. The key finding: genuine model heterogeneity—deploying domain-specialized language models alongside generalists—is the critical driver of both individual belief revision and population-level consensus shifts, far outweighing the effects of persona prompting or network topology alone.Script
When you connect dozens of language model agents in a social network and let them exchange beliefs, what happens? This work reveals that the type of model powering each agent matters far more than how you prompt it or wire the network.
The researchers built CoevolveSim, a framework that deploys 48 language model agents across controlled network topologies. Each agent holds beliefs about factual statements and updates them based on what its neighbors say. The simulation varies three factors: whether agents are powered by generalist or domain-specialized models, whether they're assigned social role labels, and how the network is wired.
Across 1,280 simulation runs, one factor dominated all others. Introducing specialist models more than doubled the shift in collective consensus compared to populations where diversity came only from persona prompting. Social roles made individual agents more plastic and induced oscillations, but barely moved the population needle.
To understand these dynamics, the authors built surrogate models inspired by classical opinion diffusion theory. Predictive accuracy climbed as they added features: agent persistence, then population belief composition, then local neighborhood structure, and finally agent identity. In homogeneous populations, simple models sufficed. In heterogeneous ones, you needed the full picture to reproduce collective outcomes.
Here's the catch: the mechanisms that best explained individual agent transitions were not the ones that reproduced population consensus. You can't infer collective outcomes from agent-level heuristics alone. This has real consequences for any system that orchestrates multiple language models for deliberation, reasoning, or content moderation.
Prompt engineering and network design matter, but they're secondary. If you want realistic multi-agent belief dynamics, you need genuine model diversity. To explore this work in depth and create your own video summaries of cutting-edge research, visit EmergentMind.com.