Papers
Topics
Authors
Recent
Search
2000 character limit reached

Program Equilibria: Code-Based Strategic Interaction

Updated 6 December 2025
  • Program equilibria are a generalization of Nash equilibrium where strategies are computer programs that inspect opponents’ source code to decide actions.
  • Key methods include proof-based and simulation-based bots, ensuring cooperation through recursive reasoning and probabilistic decision rules.
  • SDP-based techniques and MPEC frameworks provide tractable solutions for optimizing equilibria in both two-player dilemmas and multi-agent systems.

Program equilibria generalize Nash equilibrium to settings in which strategies are themselves computer programs capable of inspecting and simulating the opponents’ source code prior to selecting actions. This paradigm models interactions among mutually transparent agents, including AI systems and contractual institutions. Major frameworks include the study of “program games” for classical dilemmas such as the Prisoner’s Dilemma, multi-agent generalizations, simulation-based approaches for robustness, and mathematical programs with equilibrium constraints (MPECs) for optimization problems. Conceptual advances focus on guaranteeing robust cooperative outcomes, characterizing equilibria under informational and computational constraints, and providing SDP-based solution methods for polynomial program games.

1. Formal Definitions and Program Game Structures

In a two-player program game for the Prisoner’s Dilemma, each player submits a program pi∈PROGip_i \in \mathit{PROG}_i, which, given read access to the opponent’s source code p−ip_{-i}, outputs either CC (cooperate) or DD (defect) (Oesterheld, 2022). These programs are assumed (almost surely) to halt; otherwise, non-termination is interpreted as DD. The payoff is given by ui(p1,p2)=ui(a1,a2)u_i(p_1, p_2) = u_i(a_1, a_2), where ai=pi(p−i)a_i = p_i(p_{-i}) and uiu_i is the standard payoff function with T>R>P>ST > R > P > S.

Generalizing to nn players, let p−ip_{-i}0 be an p−ip_{-i}1-player normal-form game with finite actions. A program p−ip_{-i}2 for player p−ip_{-i}3 is a function

p−ip_{-i}4

where p−ip_{-i}5 denotes a random bitstream (sources of randomness may be independent or shared among players for coordination) (Cooper et al., 2024). Program equilibria are profiles p−ip_{-i}6 such that no player benefits from unilaterally switching to a different halting program.

2. Types of Program Equilibria: Fair Bots and Simulation-Based Strategies

The literature distinguishes proof-based bots (which inspect opponent code for provable cooperation), simulation-based bots (which actively run opponent programs on simulated inputs), and hybrids. Two principal constructions are:

  • p−ip_{-i}7-Grounded Fair Bot (p−ip_{-i}8GFB): On input opponent p−ip_{-i}9, with probability CC0 return CC1, otherwise invoke and copy CC2GFBCC3 (Oesterheld, 2022). This introduces randomization and ensures almost sure halting.
  • Proof-Based Fair Bot (PFB): On input opponent CC4, return CC5 if and only if CC6(PFB) CC7 is provable within Peano arithmetic (CC8), else return CC9 (Oesterheld, 2022).

Simulation-based strategies (e.g., DD0-GroundedDD1Bot) generalize DD2GFB: players simulate opponents on truncated random streams, halting with probability DD3, and select actions via policies DD4 based on accumulated simulated histories (Cooper et al., 2024). These are robust to code obfuscation and enable extension to the multi-player case.

3. Existence and Characterization of Cooperative Program Equilibria

For the two-player Prisoner’s Dilemma, several robust cooperative Nash equilibria have been formally established:

  • DD5 is a Nash equilibrium yielding DD6, as the recursion ensures mutual cooperation almost surely.
  • DD7 constitutes an equilibrium yielding DD8, established via Löb’s theorem: DD9 implies DD0 (Oesterheld, 2022).
  • DD1 also achieves cooperation, as DD2 can provably infer the cooperative tendency embedded in DD3.

These constructions are compatible, supporting families of robust cooperative equilibria that tolerate syntactic and algorithmic variation.

In DD4-player settings, simulation-based program equilibria are characterized as follows:

  • With shared randomness (correlated program game), a folk theorem holds: any feasible and strictly individually rational payoff vector can be supported by a program equilibrium of correlated DD5-GroundedDD6Bots. Coordination on triggers and punishment is enabled by shared random cutoffs (Cooper et al., 2024).
  • With private randomness (uncorrelated program game), stricter constraints apply. The attainable payoffs are those strictly dominating the best mixtures under undetectable deviations, as codified by a penalty parameter DD7 that controls the detection trade-off. If utilities are additively separable, the set of feasible equilibria is widened (Cooper et al., 2024).

4. Compatibility, Generalizations, and Limits

Syntactically distinct strategies (randomized grounding vs. proof-search) are shown to be compatible and jointly support cooperative outcomes. Proof-based bots can cooperate with any DD8GFB program via Löb-style arguments, and multi-parameter mixtures of DD9GFBs are stable. PrudentBot variants using extended consistency (e.g., ui(p1,p2)=ui(a1,a2)u_i(p_1, p_2) = u_i(a_1, a_2)0), and probabilistic proof requirements, also integrate into this robust equilibrium family (Oesterheld, 2022).

Key limitations arise in multi-agent settings without shared randomness. The impossibility of full folk theorem enforcement: in the pirates’ dilemma (three players, strictly Pareto-optimal payoffs), cooperation cannot be sustained by simulationist programs because unobservable deviations elude collective punishment (Cooper et al., 2024). Coordination failure in private random times is fundamental.

5. Mathematical Programs with Equilibrium Constraints (MPECs)

A related but distinct instantiation involves optimization under equilibrium or complementarity constraints (MPECs), especially when data is polynomial. Formally, for ui(p1,p2)=ui(a1,a2)u_i(p_1, p_2) = u_i(a_1, a_2)1, one minimizes ui(p1,p2)=ui(a1,a2)u_i(p_1, p_2) = u_i(a_1, a_2)2 subject to semialgebraic constraints and complementarity (ui(p1,p2)=ui(a1,a2)u_i(p_1, p_2) = u_i(a_1, a_2)3). Equilibrium can be expressed via a lower-level value function ui(p1,p2)=ui(a1,a2)u_i(p_1, p_2) = u_i(a_1, a_2)4 (Jiao et al., 2019).

Global minimizers are found using moment–sum-of-squares (SOS) hierarchies, solved by semidefinite programming (SDP). Each relaxation ui(p1,p2)=ui(a1,a2)u_i(p_1, p_2) = u_i(a_1, a_2)5 corresponds to a truncated moment matrix and localizing matrices indexed by degree. The sequence of SDP solutions converges monotonically to the global minimum under Archimedean (compactness) conditions. Experiments demonstrate computational feasibility for small examples; matrix size is the main bottleneck (Jiao et al., 2019).

Program Type Halting Guarantee Syntactic Robustness Multi-Player Generality
Proof-Based Fair Bot (PFB) Yes Low–Medium 2-player
ui(p1,p2)=ui(a1,a2)u_i(p_1, p_2) = u_i(a_1, a_2)6-Grounded Fair Bot Yes High 2-player (original)
ui(p1,p2)=ui(a1,a2)u_i(p_1, p_2) = u_i(a_1, a_2)7-Groundedui(p1,p2)=ui(a1,a2)u_i(p_1, p_2) = u_i(a_1, a_2)8Bot Yes High ui(p1,p2)=ui(a1,a2)u_i(p_1, p_2) = u_i(a_1, a_2)9-player (Cooper et al., 2024)

6. Illustrative Examples and Applications

In the canonical one-shot Prisoner’s Dilemma (ai=pi(p−i)a_i = p_i(p_{-i})0), correlated ai=pi(p−i)a_i = p_i(p_{-i})1-Groundedai=pi(p−i)a_i = p_i(p_{-i})2Bots with grim-trigger policies support mutual cooperation for small ai=pi(p−i)a_i = p_i(p_{-i})3; a single deviation triggers punishment with probability ai=pi(p−i)a_i = p_i(p_{-i})4, implying approximately optimal payoffs ai=pi(p−i)a_i = p_i(p_{-i})5 (Cooper et al., 2024).

The trust game demonstrates mixed-strategy equilibrium with uncorrelated ai=pi(p−i)a_i = p_i(p_{-i})6-Groundedai=pi(p−i)a_i = p_i(p_{-i})7Bots: mixing greedy and charitable responses enforces a stable equilibrium indistinguishable from true mixed strategies, with detection and punishment dependent on random stopping (Cooper et al., 2024).

Polynomial MPECs are solved by SDP relaxations, as evidenced by small-scale benchmarks in (Jiao et al., 2019). Realistic use is currently limited to problems with ai=pi(p−i)a_i = p_i(p_{-i})8 and relaxation order ai=pi(p−i)a_i = p_i(p_{-i})9.

7. Implications for AI, Multi-Agent Systems, and Economic Mechanisms

The program equilibrium paradigm expands the set of attainable cooperative solutions in strategic settings once agents have transparent access to each other’s source code. Randomization, proof-based reasoning, and simulation protocols are mutually compatible, suggesting robust frameworks for multi-agent trust and commitment without requiring identical implementations. The necessity of shared randomness for full coordination in uiu_i0-player games is fundamental. SDP-based techniques for polynomial games with equilibrium and complementarity constraints provide tractable global optimization methods for small to medium scale problems. These developments underpin advances in programmatic contract design, AI alignment, and mechanism design in transparent agent environments (Oesterheld, 2022, Cooper et al., 2024, Jiao et al., 2019).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Program Equilibria.