Papers
Topics
Authors
Recent
Search
2000 character limit reached

Axelrod's Second Tournament

Updated 2 July 2026
  • Axelrod's Second Tournament is a landmark experiment that advanced studies on cooperation through large-scale, Fortran-coded iterations of the Prisoner's Dilemma.
  • The experiment employed a round-robin format with geometric match lengths and introduced noise to assess strategies under indefinite horizons.
  • Reproduction studies confirmed Tit for Tat's dominance while highlighting the influence of noise, match length variability, and population diversity on strategy performance.

Axelrod's Second Tournament of the Iterated Prisoner's Dilemma (IPD) marks a pivotal experiment in the empirical study of direct reciprocity and the evolution of cooperation. Conducted in 1980–1981 by Robert Axelrod, it expanded significantly on his initial 1980 tournament, both in field diversity and methodological sophistication. Core findings—such as the predominance of Tit for Tat (TFT) amidst a diverse field—have shaped the theoretical and computational paradigms of cooperation. Despite its historical influence, incomplete archival records have necessitated recent reconstruction and rigorous reproducibility analysis (Knight et al., 17 Oct 2025).

1. Historical Context and Objectives

Axelrod's second IPD tournament arose within a research agenda aiming to elucidate the emergence of cooperation among self-interested agents. The first tournament (1980) employed 14 strategies, fixed-length matches, and standard PD payoffs, with TFT emerging decisively as the winner. The second, conducted shortly thereafter, substantially increased experimental scale: 63 Fortran-coded submissions, each playing round-robin with all others (and self), with match lengths drawn from a geometric distribution (termination probability p=0.00346p = 0.00346 per round, yielding a median length around 200; actual sampled lengths were 63, 77, 151, 156, and 308 moves). Unlike the first, the horizon was unknown to participants, a critical shift designed to mirror indefinite social and biological interactions. The goal was multi-faceted: to probe whether TFT would maintain its success, to identify principles underpinning robust strategies (niceness, provocability, forgiveness, non-envy, simplicity), and to assess how population composition and horizon uncertainty affect outcomes (Knight et al., 17 Oct 2025).

2. Experimental Design and Technical Implementation

Payoff Structure and Match Protocols

The experiment retained the canonical PD payoff ordering (T>R>P>ST > R > P > S; $2R > T + S$), using T=5T = 5, R=3R = 3, P=1P = 1, S=0S = 0. Each of 63 strategies was tested in pairwise round-robin matches (including self-play), across the same five predefined lengths for every pairing. Outcomes of each match provided total payoffs, which were averaged by strategy, opponent, and match length, then normalized as mean per-turn payoff for final ranking.

Reviving the Tournament Code

The Fortran source (TourExec1.1.f) was recovered from archival HTML, with HTML tags stripped and the code modularized to isolate each strategy. Minimal compatibility edits were made for current Fortran compilers, and a Makefile introduced to automate building all 63 into a shared object library. To enable robust experimental control and extensibility, an “axelrod-fortran” adapter was developed: this Python-based interface within Axelrod-Python passes, at every turn, all required match state (opponent’s and own last moves, turn number, cumulative scores, random value) directly to the original Fortran logic, retrieving the resulting binary cooperate/defect action without reinterpretation. This pipeline permits exact preservation of the original strategy semantics and deterministic or stochastic behavior (Knight et al., 17 Oct 2025).

3. Principal Findings and Ranking Recovery

Reproducibility of Rankings

With the revived code, 25,000 full tournaments were simulated to average out stochastic effects and eliminate artifacts from a single realization. TFT (k92r, Anatol Rapoport) again attained rank 1, with other strategies’ ranks remaining stable: 23 of the top 25 matched the original ordering within ±2 positions, except for two strategies (due to a corrected bug in k61r, “Champion”). The lowest-ranked strategy (k36r) was unchanged.

Top 11 Strategies and Metrics

ID Author Mean Score Reproduced Rank Original Rank Coop. Rate
k92r Anatol Rapoport 2.878 1 1 0.922
k42r Otto Borufsen 2.861 2 3 0.916
k44r William Adams 2.826 3 5 0.882
k49r Rob Cave 2.826 4 4 0.892
k75r Paul Harrington 2.826 5 8 0.802
k32r Charles Kluepfel 2.817 6 10 0.873
k84r Tideman & Chieruzzi 2.812 7 9 0.882
k41r Herb Weiner 2.811 8 7 0.865
k60r Graaskamp & Katzen 2.810 9 6 0.844
k68r Fr. Leyvraz 2.808 10 12 0.917
k35r Abraham Getzler 2.801 11 11 0.888

Across 25,000 replications, TFT won 63.2% of tournaments; k42r (Borufsen) won 31.4%; the remainder was divided among k75r, k32r, k44r, k60r, and others with minor proportions.

Cooperation and Strategy Typology

Average cooperation across all matches was approximately 75%. Top strategies reciprocated cooperation above 90% of the time against each other, manifesting the targeted mechanism of “niceness, provocability, and forgiveness.” Representative-strategy regression (predicting mean score via regressors on payoff vs five Axelrod-chosen standard strategies) reproduced the original R20.978R^2 \approx 0.978 result (original: R2=0.979R^2=0.979), affirming that the primary structure of performance could be linearly decomposed (Knight et al., 17 Oct 2025).

4. Robustness, Noise, and Strategy Diversity

Field Expansion and Strategy Invasion

Addition of up to four distinct strategies (sampled from 209 additional Axelrod-Python entrants) to the original roster led to only incremental disturbance of rankings: with one new entrant, TFT retained ≈85% win probability; with four, ≈75%. The leading challenger remained Borufsen’s k42r. Some original strategies (e.g., k32r) rapidly lost prominence as diversity increased, while others remained robust.

Introduction of Noise

With a 1% error rate per move (intended cooperate/defect decisions probabilistically flipped), ranking volatility increased. TFT’s rank dropped, but it persisted near the top; mean cooperation declined from 0.75 to 0.65; top strategy cooperation rates fell further (winners averaged ≈0.61). This indicates a structural advantage for strategies slightly less cooperative than the mean under noisy conditions.

Large–Scale and Mega–Tournaments

With the inclusion of Stewart & Plotkin ZD strategies (~70 total strategies), leadership remained with TFT, k42r, and generous ZD variants; pure ZD strategies were not generally dominant. When the entire Axelrod-Python portfolio (272 strategies, including all originals) was evaluated, reinforcement-learning–optimized strategies (EvolvedLookerUp2_2_2, Evolved HMM 5, Omega TFT) led, with “Gradual” and k42r the best-performing original submissions at ranks 9 and 16, respectively. Overall cooperation in these large tournaments dropped to ≈62%. Under further increases in noise (1% and 5%), some originally engineered strategies re-emerged among the global top 5, suggesting a context-dependent re-prioritization of effective reciprocity mechanisms under execution uncertainty (Knight et al., 17 Oct 2025).

5. Theoretical Implications for Cooperation Research

The persistence of TFT as winner in the original and replicated settings is attributed to a highly cooperative tournament composition: most strategies reciprocated and forgave, privileging contingent, non-exploitative rules. The unknown horizon systematically discouraged last-move defections and favored robust, parsimonious policies. However, robustness tests reveal that TFT’s supremacy is context-dependent: in more heterogeneous, error-prone, or competitive fields, strategies with nontrivial envy or nuanced responses (including some mid-ranking originals and evolved heuristics) can dominate.

Reproducing the original results with high fidelity validates the key mechanisms identified by Axelrod—niceness, retaliation, forgiveness—as engines of cooperation, but also surfaces the field’s sensitivity to noise, population composition, and other model features. Empirically, the strong linear predictability of performance with respect to a handful of representative strategies remains stable (Knight et al., 17 Oct 2025).

6. Methodological Limitations and Future Research Directions

Methodological constraints persist: both the original tournament’s exact random seed and its final orchestration code are missing, rendering exact historical reproduction infeasible. The historical result’s dependence on a single realization of match lengths exposes rank sensitivity to sampling variation. No systematic error analysis (noise) was undertaken in the original study.

Open avenues include systematic parameter sweeps over execution error and match-length distribution; evolutionary simulations with the modernized codebase; and preservation/reconstruction of other landmark computational social science experiments. The integration of the revived Fortran strategies into modern open-source infrastructure (Axelrod-Python) ensures accessibility for future replication and extension, supporting methodological rigor and theory development in the computational study of cooperation (Knight et al., 17 Oct 2025).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Axelrod's Second Tournament.