Measuring and Shaping Capabilities and Dispositions Under Multi-Agent Selection Pressures
Develop methods to measure and shape the capabilities and dispositions of AI systems that specifically account for multi-agent selection pressures, including how competitive and adaptive interactions select for traits that impact cooperation, deception, and conflict.
References
It is therefore an important open problem to develop methods for measuring and shaping the capabilities and dispositions of AI systems that account for multi-agent selection pressures.
The open problem is the fitness function. In biology, fitness is imposed from outside: an organism either survives its environment or it does not, and no lineage votes on the criterion. A model population has no such guarantee, because the criterion is written by whoever runs it, and selection will improve whatever that criterion measures. Scored on agreement with its own members instead of against reality, the population here evolved to a confident consensus that was wrong, and stayed well below a population selected on true fitness (Fig. 4D--F). Verified real data therefore do two jobs that discussions of model collapse treat as one: they restore the variety lost to drift, and they supply the selective environment. Building that environment for large populations of models, through verification, replication and challenge, is the question this paper raises and does not answer.
Since this experiment examines one model and a fixed injection order, future work should vary condition composition and order to test whether subtler combinations amplify deception without eliciting comparable refusal.