---
title: Optional Prisoner's Dilemma Game
url: https://www.emergentmind.com/topics/optional-prisoner-s-dilemma-game
type: topic
---

# Optional Prisoner's Dilemma Game

The **Optional Prisoner’s Dilemma Game** (OPD), also called the **voluntary prisoner’s dilemma**, extends the standard Prisoner’s Dilemma by adding a third action—typically **abstain**, **loner**, or **exit**—alongside cooperation and defection. In the canonical optional formulation, if one or both players abstain, both receive a fixed loner payoff \(L\) or \(\sigma\), so the game is no longer a \(2\times 2\) dilemma but a three-strategy system. This modification changes the strategic geometry from a purely binary tension between \(C\) and \(D\) to a setting in which \(D\) beats \(C\), \(C\) beats \(A\), and \(A\) beats \(D\), thereby enabling cyclic dominance, coexistence, and structure-dependent support for cooperation [1811.10114, 1907.05482].

## 1. Formal structure and representative payoff schemes

The standard Prisoner’s Dilemma is parameterized by the reward for mutual cooperation \(R\), the punishment for mutual defection \(P\), the sucker’s payoff \(S\), and the temptation payoff \(T\). In the classical form used in the probabilistic-abstention literature, the dilemma condition is
\[
T > R > P > S,
\]
with the common normalization
\[
R = 1,\quad P = 0,\quad S = 0,\quad 1 < T < 2.
\]
The OPD adds a third option, abstention, and if either player abstains both receive the loner payoff \(L\), so the ordering becomes
\[
T > R > L > P > S
\]
in the formulation that explicitly treats abstention as a third strategic alternative [1811.10114].

A frequently used weak-OPD parametrization in spatial models is
\[
R=1,\qquad P=0,\qquad S=0,\qquad T=b,\quad 1<b<2,
\]
with abstention payoff \(L=l\) satisfying
\[
P<L<R \quad \Rightarrow \quad 0<l<1.
\]
Under this specification, the payoff matrix is
\[
\begin{array}{c|ccc}
 & C & D & A\\ \hline
C & (R,R) & (S,T) & (L,L)\\
D & (T,S) & (P,P) & (L,L)\\
A & (L,L) & (L,L) & (L,L)
\end{array}
\]
so abstention interrupts the exploitative \(C\)-\(D\) interaction and replaces it by a fixed outside option [1609.06560].

Other papers use different but equivalent normalizations. One study adopts the Axelrod values
\[
T=5,\quad R=3,\quad P=1,\quad S=0,\quad L\in]0,3[
\]
and interprets any interaction involving \(A\) as giving both players payoff \(L\) [1608.05044]. A one-shot anonymous OPD in a human–machine mixed population rescales payoffs by
\[
r=\frac{c}{b-c},
\]
yielding
\[
\begin{array}{c|ccc}
~ & C & D & L\\ \hline
C & 1 & -r & \sigma\\
D & 1+r & 0 & \sigma\\
L & \sigma & \sigma & \sigma
\end{array}
\]
with \(L\) written as the loner action and \(\sigma>0\) as the corresponding payoff [2305.15818].

| Formulation | Strategy set | Characteristic conditions |
|---|---|---|
| Weak spatial OPD | \(C,D,A\) | \(R=1,P=0,S=0,T=b,1<b<2,0<l<1\) |
| Axelrod-style optional strategy | \(C,D,A\) | \(T=5,R=3,P=1,S=0,L\in]0,3[\) |
| One-shot anonymous OPD | \(C,D,L\) | payoff matrix rescaled by \(r=\frac{c}{b-c}\) |

These formulations differ in normalization and application, but they share the same defining feature: abstention is an explicit participation decision with its own payoff consequences.

## 2. Strategic logic of optional participation

The introduction of abstention changes the strategic relation among the available actions. In the standard account of OPD, the three strategies satisfy the cyclic relation
\[
D \text{ beats } C,\qquad C \text{ beats } A,\qquad A \text{ beats } D.
\]
This rock-paper-scissors-like structure is one of the central reasons optional participation can sustain cooperation even when direct \(C\)-\(D\) competition would favor defection [1811.10114].

The logic is straightforward. Defectors exploit cooperators in direct interaction. Cooperators outperform abstainers because abstention yields only the loner payoff, whereas mutual cooperation can yield \(R>L\). Abstainers outperform defectors because opting out avoids the low-payoff environments created by defection. In evolutionary formulations, this means abstention can prevent unconditional takeover by defectors, but it does not eliminate strategic turnover; defection often persists because the game becomes cyclic rather than monotone [1907.05482].

A sharper statement appears in coevolutionary spatial OPD. When the population is reduced to only two strategies, the cycle collapses:
- \(C+A \rightarrow C\) dominates,
- \(D+A \rightarrow A\) dominates,
- \(C+D \rightarrow D\) dominates.

Thus the coexistence mechanism is genuinely three-strategy; removing any one component breaks the intransitive loop [1702.04299].

An early abstract on repeated finite Prisoner’s Dilemma also suggested a variant in which players can choose to opt out. That modification was said to enrich the game and to suggest dominance of cooperative strategies, while also linking bounded rationality, computational limits, and competitive analysis to the study of tractable but sub-optimal play [0701139]. This suggests that optionality entered the literature not only as a payoff perturbation but also as a way of altering the temporal and computational structure of the dilemma.

## 3. Evolutionary behavior in well-mixed and spatial populations

The OPD behaves differently in non-spatial and spatial environments. In a non-spatial evolutionary model with tournament selection, the threshold between defectors and abstainers is set by the comparison between the loner payoff and mutual defection payoff. For the pairwise defector–abstainer comparison, the paper derives
\[
P_D - P_A = |D-1|(P-L).
\]
Hence defectors and abstainers are tied at
\[
L=P=1,
\]
defectors dominate when \(L<P\), and abstainers dominate when \(L>P\). For the cooperator–abstainer comparison, it derives
\[
P_C - P_A = |C-1|(R-L),
\]
so cooperators always dominate abstainers because \(R>L\) in that model [1608.05044].

In the same well-mixed setting with all three strategies initially present at equal frequency, the reported outcomes are:
- for \(L<P\), defectors dominate;
- for \(L=P\), defectors still dominate on most runs, though abstainers occasionally prevail;
- for \(L>P\), abstainers become increasingly dominant, and in some runs cooperators can outperform defectors.

The same paper reports that cooperation is fragile in the non-spatial model, surviving mainly when abstainers gain an advantage over defectors and thereby indirectly protect cooperators [1608.05044].

Spatial structure changes the dynamics because it allows clustering. On a \(100\times100\) lattice with Moore neighborhoods, pairwise spatial comparisons reveal that adjacent cooperators can reinforce each other and spread against abstainers regardless of \(L\), while the \(D\)–\(A\) relation retains the threshold logic tied to \(L\) versus \(P\). With equal random initial densities of \(C\), \(D\), and \(A\):
- for \(L<P\), defectors quickly dominate, but cooperative clusters survive in about **65%** of simulations thanks to abstainers;
- for \(L\in[1.1,2.0]\), abstainers often dominate, but cooperation may persist in stable clusters;
- in about **51.5%** of simulations for \(L\in[1.1,2.0]\), a cooperative cluster of minimum size **9** forms early and persists.

The stable morphology described there is a “sandwich” configuration in which cooperators are surrounded by defectors and abstainers occupy the outer region. The same study reports “gliders” for \(L\in[1.7,1.9]\), especially at \(L=1.5\) and \(L=2.0\), where defectors and abstainers switch cyclically near the boundary [1608.05044].

A different line of work replaces pure abstention by **probabilistic abstention**. In this hybrid model, each agent is described by
\[
(s,a),
\]
where \(s\in\{0,1\}\) denotes cooperation or defection and \(a\in[0,1]\) is the probability of abstaining. The paper defines
\[
\epsilon=(1-s)(1-a),
\]
so \(\epsilon=1\) corresponds to a pure cooperator who always plays, \(\epsilon=0\) to a defector or full abstainer, and intermediate values to sporadic participation. The model reduces to the standard PD when \(a=0\) for all players, and to the OPD when \(a\in\{0,1\}\). Across the tested parameter ranges, this hybrid sustains higher cooperation than both standard PD and standard OPD under synchronous and asynchronous updating, with intermediate abstention probabilities reported as the most favorable regime for cooperation [1811.10114].

## 4. Coevolution, mobility, and adaptive interaction structure

A major development in OPD research is the move from static lattices to coevolving or diluted interaction structures. In a weighted spatial OPD, agents occupy a \(100\times100\) square lattice with Moore neighborhoods, each edge begins with weight
\[
w_{xy}=1,
\]
and utilities are computed as
\[
u_{xy}=w_{xy}P_{xy},\qquad U_x=\sum_{y\in\Omega_x}u_{xy}.
\]
Link weights are then adapted according to whether a local interaction utility is above or below the focal player’s average utility, with weights constrained by
\[
1-\delta \le w_{xy} \le 1+\delta.
\]
Strategy imitation occurs only if a random neighbor has higher utility, with probability
\[
p(s_x=s_y)=\frac{U_y-U_x}{8(T-P)}.
\]
Within this framework, abstainers are reported to protect cooperators against exploitation, especially when the link-weight amplitude is large. The paper identifies three qualitative regimes: abstainer-dominated freezing in the static or weakly adaptive case, cyclic dominance for intermediate coevolution strength, and cooperation dominance for strong coevolution combined with sufficiently favorable loner payoff [1609.06560].

A representative cyclic-dominance regime is reported at
\[
\frac{\Delta}{\delta}=0.2,\qquad b=1.9,\qquad l=0.6,\qquad \delta=0.8,
\]
where abstainers invade defectors, defectors invade cooperators, and cooperators invade abstainers. A representative cooperation-dominant regime is
\[
\frac{\Delta}{\delta}=1.0,\qquad l=0.6,\qquad b=1.9,\qquad \delta=0.8,
\]
where defectors are first suppressed by abstainers and small cooperative clusters later expand and invade abstainers [1609.06560].

A related coevolutionary model on a \(102\times102\) lattice reports that when
\[
\frac{\Delta}{\delta}=0.3,\quad \delta=0.8,\quad l=0.5,\quad b=1.9,
\]
the three strategies stabilize around
\[
\approx 33\% \pm 7\%
\]
each. The same study emphasizes that cyclic dominance breaks down under two-strategy reductions and that recovery after severe mutation depends on the continued presence of the full spatial-support chain \(C\)-near-\(A\), \(A\)-near-\(D\), and \(D\)-near-\(C\) [1702.04299].

Mobility on diluted lattices adds another mechanism. In the voluntary prisoner’s dilemma with density
\[
\rho=\frac{N}{M^2},\qquad 0<\rho<1,
\]
agents interact on a diluted square lattice with von Neumann neighborhoods and may move to neighboring empty sites according to a Fermi-like rule based on normalized utility. A key geometric threshold is the lattice percolation threshold
\[
\rho_p\approx 0.59.
\]
On a fully occupied lattice (\(\rho=1\)), cyclic dominance survives under noisy imitation, but under fully rational imitation cooperators die out and the system freezes into a \(D+A\) state. With dilution and movement, the same cooperation-supporting mechanism reappears for most \(b\)-\(\sigma\) values when
\[
1>\rho>\rho_p.
\]
At very low density, the situation reverses: for \(\rho\le 0.05\) cooperators die out and abstainers dominate, while around \(\rho\approx 0.10\) the system shows bistability, with runs ending in all-\(C\) or all-\(A\) [1907.05482].

These results collectively indicate that optionality does not operate independently of interaction structure. Its effect depends strongly on whether the environment is well mixed, spatial, weighted, diluted, or mobile.

## 5. Behavioral, environmental, and institutional extensions

The OPD has also been extended beyond standard evolutionary settings. In a one-shot anonymous human–machine mixed population, simple bots are assigned fixed strategies: always cooperate, always defect, never participate, or choose each action with probability \(1/3\). In well-mixed populations, cooperative bots are reported to facilitate the emergence of cooperation under weak imitation, while loner bots have no meaningful effect. On regular lattices, loner bots become more consequential: under strong imitation they can facilitate the dominance of cooperation, but the effect is nonmonotonic. Around \(\rho=0.4\), defectors are eliminated and cooperation can expand; around \(\rho=0.5\), loner bots surround cooperative clusters and block further spread, reducing cooperation [2305.15818].

A different institutional extension adds **pre-game commitment**. In that two-stage model, each player first chooses whether to accept commitment,
\[
A \text{ or } N,
\]
and then, in the game stage, chooses among
\[
C,\ D,\ L.
\]
A full strategy has the form
\[
XYZ,
\]
where \(X\in\{A,N\}\), \(Y\in\{C,D,L\}\) is the action if commitment is formed, and \(Z\in\{C,D,L\}\) is the action otherwise, for a total of **18 strategies**. The OPD payoff matrix is
\[
\begin{array}{c|ccc}
 & C & D & L\\ \hline
C & R & S & \sigma\\
D & T & P & \sigma\\
L & \sigma & \sigma & \sigma
\end{array}
\]
with
\[
T>R>\sigma>P>S,\qquad R=1,\ S=-1,\ T=2,\ P=0,\ \sigma\in[0,1].
\]
The main result is that optional participation boosts commitment acceptance but fails to foster cooperation, leading instead to widespread exit behavior. Two institutional reward rules are then compared. Under **STRICT-COM**, only committed players who cooperate are rewarded; under **FLEXIBLE-COM**, any committed player who does not defect is rewarded. The strict rule is reported to promote cooperation more effectively, while the flexible rule creates an opportunistic exit loophole, though it can yield higher social welfare when \(\sigma\) is high and the reward budget \(u\) is limited [2508.06702].

Another extension couples OPD to a dynamic environment. There the payoff matrix depends on an environmental variable \(n\in[0,1]\),
\[
A(n)=(1-n)\begin{bmatrix}
T & P & L\\
R & S & L\\
L & L & L
\end{bmatrix}
+n\begin{bmatrix}
R & S & L\\
T & P & L\\
L & L & L
\end{bmatrix},
\]
and the environment evolves according to
\[
\dot n=n(1-n)\big[(1+\lambda)x_1-1\big],
\]
where \(x_1\) is the cooperator fraction. In the replicator version, the paper reports **11 fixed points** and, in its illustrative example, eventual attraction to the all-abstain state \((0,0,0)\). In the pairwise-comparison version inspired by prospect theory, the paper reports **10 fixed points** and convergence to an asymptotically stable interior equilibrium with
\[
x_1^*=\frac{1}{1+\lambda},\qquad n^*=\frac12.
\]
Here optionality is not merely a static outside option; it becomes part of a closed game–environment feedback system in which abstention, cooperation, and defection coevolve with environmental quality [2106.07100].

## 6. Conceptual boundaries, related variants, and interpretive issues

Not every three-action extension of the Prisoner’s Dilemma is an OPD in the strict sense. A conceptually related but distinct model is the **generalized prisoner’s dilemma** with strategy set
\[
\{C,D,S\},
\]
where \(S\) denotes **Silence**. In that formulation, \(S\) is described as neither cooperation nor defection, an ambiguous attitude, and a special state that may correspond to either \(C\) or \(D\) but is not distinguishable to the police. The classical Prisoner’s Dilemma is recovered when the third state is not taken into consideration. However, the model provides no explicit numerical payoff entries for the \(S\)-rows or \(S\)-columns and no equilibrium analysis for the generalized case. It is therefore conceptually related to optional participation but not equivalent to the standard abstain/exit interpretation of OPD [1403.3595].

A second conceptual distinction concerns what counts as “optionality.” In standard OPD, abstention is itself a strategic action. In the mobility-based voluntary prisoner’s dilemma, this point is made explicit: abstention is modeled as a true strategy \(A\) with loner payoff \(\sigma\), not as physical movement away from an interaction. Mobility is a separate coevolutionary process that can restore cyclic dominance when strict rational updating on fully occupied graphs would otherwise create artificial frozen states [1907.05482].

A third distinction concerns whether abstention is a pure strategy or a participation propensity. In probabilistic abstention, abstention is an attribute \(a\in[0,1]\) attached to each agent rather than a separate pure strategy, and the OPD appears as the limiting case \(a\in\{0,1\}\). This suggests that “optional participation” spans a family of formal devices: pure loner strategies, probabilistic participation rates, pre-game exit contingencies, and environment-coupled outside options [1811.10114].

Taken together, the literature describes the Optional Prisoner’s Dilemma not as a single frozen model but as a research program organized around one structural innovation: agents may refuse direct participation in the dilemma. The consequences of that innovation are highly sensitive to payoff normalization, microscopic update rules, spatial organization, link adaptation, mobility, commitment institutions, and game–environment feedback. In some settings abstention mainly preserves biodiversity through cyclic dominance; in others it protects cooperative clusters, induces exit-dominated equilibria, or becomes the basis for stronger institutional design problems. This suggests that the OPD is best regarded as a family of participation-sensitive Prisoner’s Dilemma models rather than a single canonical game.

Source: https://www.emergentmind.com/topics/optional-prisoner-s-dilemma-game