Measuring and Shaping Capabilities and Dispositions Under Multi-Agent Selection Pressures

Develop methods to measure and shape the capabilities and dispositions of AI systems that specifically account for multi-agent selection pressures, including how competitive and adaptive interactions select for traits that impact cooperation, deception, and conflict.

Background

Training and deployment in multi-agent environments can select for dispositions and capabilities distinct from those arising in single-agent settings. Competitive pressures, repeated interactions, and adaptive co-learning may favor traits such as deception, aggression, or collusion that elevate systemic risk.

Existing evaluations often overlook how agents change when interacting with diverse co-players and evolving environments. New measurement tools and training schemes are needed to characterize and steer these selection dynamics toward cooperative, reliable behavior.

References

It is therefore an important open problem to develop methods for measuring and shaping the capabilities and dispositions of AI systems that account for multi-agent selection pressures.

— Multi-Agent Risks from Advanced AI  (2502.14143 - Hammond et al., 19 Feb 2025) in Section Selection Pressures, Directions

The open problem is the fitness function. In biology, fitness is imposed from outside: an organism either survives its environment or it does not, and no lineage votes on the criterion. A model population has no such guarantee, because the criterion is written by whoever runs it, and selection will improve whatever that criterion measures. Scored on agreement with its own members instead of against reality, the population here evolved to a confident consensus that was wrong, and stayed well below a population selected on true fitness (Fig. 4D--F). Verified real data therefore do two jobs that discussions of model collapse treat as one: they restore the variety lost to drift, and they supply the selective environment. Building that environment for large populations of models, through verification, replication and challenge, is the question this paper raises and does not answer.

— The evolution of sex for artificial intelligence: a population-genetic framework for multigenerational model populations  (2609.18560 - Gilestro, 16 Sep 2026) in Discussion, paragraph beginning “The open problem is the fitness function.”

Since this experiment examines one model and a fixed injection order, future work should vary condition composition and order to test whether subtler combinations amplify deception without eliciting comparable refusal.

— DecepEval: A Benchmark for Evaluating Deception in LLM Agents  (2610.07967 - Xu et al., 6 Oct 2026) in Section 3.4, Effects of Combined Conditions (p. 8)