Papers
Topics
Authors
Recent
Search
2000 character limit reached

Closed-form CE-ML Estimator

Updated 22 January 2026
  • Closed-form CE-ML is a statistical estimator that inversely learns agent payoffs in 2x2 games using observed action frequencies under correlated equilibrium assumptions.
  • It exploits the tractable structure of the CE polytope to derive closed-form solutions for both the equilibrium distribution and payoff parameters via payoff ratios α and β.
  • Empirical evaluations reveal that CE-ML delivers superior accuracy and computational efficiency compared to ICE and LBR-ML in scenarios such as coordination and traffic games.

A Closed-form Correlated Equilibrium Maximum-Likelihood Estimator (CE-ML) provides a statistically efficient and interpretable method for inverse learning of agent payoffs in 2×22\times2 games under the assumption that observed joint action frequencies are generated according to a Correlated Equilibrium (CE). The CE-ML estimator exploits the tractable combinatorial structure of the CE polytope in strict coordination games and yields parameters that are directly consistent with empirical frequencies, leveraging a closed-form solution for both the equilibrium distribution and underlying payoff parameters. This approach is specialized for scenarios where agent strategies coordinate through CE, and offers explicit trade-offs between interpretability, computational efficiency, and fidelity to observed behavior (Salazar et al., 15 Jan 2026).

1. Inverse Learning in 2×22\times2 Games and Correlated Equilibrium

Consider a two-player game with each player i∈{1,2}i \in \{1,2\} having actions A1={a11,a12}A_1 = \{a_1^1,a_1^2\}, A2={a21,a22}A_2 = \{a_2^1,a_2^2\}, leading to four possible joint action profiles A={a(1),a(2),a(3),a(4)}A = \{a(1),a(2),a(3),a(4)\}, where a(1)=(a11,a21)a(1) = (a_1^1,a_2^1), a(2)=(a11,a22)a(2) = (a_1^1,a_2^2), a(3)=(a12,a21)a(3) = (a_1^2,a_2^1), a(4)=(a12,a22)a(4) = (a_1^2,a_2^2). Each player’s utility for a profile 2×22\times20 is specified via a feature mapping and linear parameterization: 2×22\times21, with 2×22\times22. Observing 2×22\times23 i.i.d. samples 2×22\times24 drawn from an unknown equilibrium strategy 2×22\times25, inverse game-theoretic learning asks for parameters 2×22\times26 such that 2×22\times27 as a CE matches empirical action frequencies.

A joint distribution 2×22\times28 over 2×22\times29 is a correlated equilibrium for payoffs i∈{1,2}i \in \{1,2\}0 if, for each player i∈{1,2}i \in \{1,2\}1, any deviation i∈{1,2}i \in \{1,2\}2 satisfies

i∈{1,2}i \in \{1,2\}3

2. Structure of the i∈{1,2}i \in \{1,2\}4 CE Polytope

Under the no strictly dominated strategies assumption, the CE polytope of a i∈{1,2}i \in \{1,2\}5 coordination game has exactly five extreme points (vertices) i∈{1,2}i \in \{1,2\}6 [Calvo-Armengol 2003]. Each CE is a mixture over these vertices, parameterized as i∈{1,2}i \in \{1,2\}7 with i∈{1,2}i \in \{1,2\}8. Of particular importance is i∈{1,2}i \in \{1,2\}9, the unique interior CE, for which the equilibrium probabilities are determined by payoff ratios:

  • Define

A1={a11,a12}A_1 = \{a_1^1,a_1^2\}0

  • The interior CE probabilities are:

A1={a11,a12}A_1 = \{a_1^1,a_1^2\}1

The other four vertices are degenerate CEs, placing probability mass on one or two pure profiles.

3. Maximum-Likelihood Estimation: Closed-Form Solution

Given action counts A1={a11,a12}A_1 = \{a_1^1,a_1^2\}2 and empirical frequencies A1={a11,a12}A_1 = \{a_1^1,a_1^2\}3, the log-likelihood under CE parameterization is

A1={a11,a12}A_1 = \{a_1^1,a_1^2\}4

Due to non-concavity of the log-likelihood in A1={a11,a12}A_1 = \{a_1^1,a_1^2\}5, the optimum always lies at a CE vertex. The estimator proceeds by:

  • Computing A1={a11,a12}A_1 = \{a_1^1,a_1^2\}6 for A1={a11,a12}A_1 = \{a_1^1,a_1^2\}7 (log-likelihood under each vertex).
  • Selecting A1={a11,a12}A_1 = \{a_1^1,a_1^2\}8 and setting A1={a11,a12}A_1 = \{a_1^1,a_1^2\}9.
  • Optimizing A2={a21,a22}A_2 = \{a_2^1,a_2^2\}0 for the selected vertex.

When A2={a21,a22}A_2 = \{a_2^1,a_2^2\}1 (interior CE), partial derivatives with respect to A2={a21,a22}A_2 = \{a_2^1,a_2^2\}2 and stationarity yield closed-form MLEs:

A2={a21,a22}A_2 = \{a_2^1,a_2^2\}3

Substituting definitions, this translates to linear constraints on A2={a21,a22}A_2 = \{a_2^1,a_2^2\}4:

A2={a21,a22}A_2 = \{a_2^1,a_2^2\}5

With normalization (e.g., fixing A2={a21,a22}A_2 = \{a_2^1,a_2^2\}6 or specifying a component), this yields A2={a21,a22}A_2 = \{a_2^1,a_2^2\}7 explicitly. If another vertex is selected, the constraints collapse to linear equalities reflecting the induced payoff ordering for pure or edge CEs.

4. Assumptions, Regularity, and Computational Properties

CE-ML relies on:

  • Absence of strictly dominated strategies, so A2={a21,a22}A_2 = \{a_2^1,a_2^2\}8 remain finite and the polytope has five vertices.
  • Sufficient data support: A2={a21,a22}A_2 = \{a_2^1,a_2^2\}9, A={a(1),a(2),a(3),a(4)}A = \{a(1),a(2),a(3),a(4)\}0 to ensure A={a(1),a(2),a(3),a(4)}A = \{a(1),a(2),a(3),a(4)\}1 are defined.
  • Unique maximal vertex for the likelihood (ties resolved arbitrarily).

Computation is efficient: single-pass statistics and closed-form solutions yield total complexity A={a(1),a(2),a(3),a(4)}A = \{a(1),a(2),a(3),a(4)\}2. In degenerate data scenarios (e.g., samples occupy just two action cells), if denominators in the MLE vanish, any A={a(1),a(2),a(3),a(4)}A = \{a(1),a(2),a(3),a(4)\}3 consistent with the observed payoff orderings suffices.

Numerical stability is ensured by selecting vertices that allocate zero mass to unobserved profiles when A={a(1),a(2),a(3),a(4)}A = \{a(1),a(2),a(3),a(4)\}4 for some A={a(1),a(2),a(3),a(4)}A = \{a(1),a(2),a(3),a(4)\}5.

5. Illustrative Application: Chicken-Dare Game

In the "chicken-dare" case [Bestick et al. 2013], 1000 simulated samples yield frequencies A={a(1),a(2),a(3),a(4)}A = \{a(1),a(2),a(3),a(4)\}6, A={a(1),a(2),a(3),a(4)}A = \{a(1),a(2),a(3),a(4)\}7, A={a(1),a(2),a(3),a(4)}A = \{a(1),a(2),a(3),a(4)\}8, A={a(1),a(2),a(3),a(4)}A = \{a(1),a(2),a(3),a(4)\}9. Applying the closed-form:

a(1)=(a11,a21)a(1) = (a_1^1,a_2^1)0

The interior vertex likelihood dominates, so payoff ratio constraints directly recover a(1)=(a11,a21)a(1) = (a_1^1,a_2^1)1 and a(1)=(a11,a21)a(1) = (a_1^1,a_2^1)2 up to scale, matching the true payoffs within numerical tolerance. This outcome demonstrates the estimator's ability to precisely reconstruct underlying game parameters when agent behavior is CE-conforming.

6. Empirical Evaluation and Performance

CE-ML was evaluated alongside Inverse Correlated Equilibrium (ICE) and Logit Best Response ML (LBR-ML) estimators on synthetic and SUMO traffic interaction data, with four primary experiments:

  • E1 (Chicken via CE): As a(1)=(a11,a21)a(1) = (a_1^1,a_2^1)3 increases, CE-ML MAE/RMSE improves (a(1)=(a11,a21)a(1) = (a_1^1,a_2^1)4), besting ICE and LBR-ML.
  • E2 (Traffic, maximum-entropy CE): CE-ML achieves MAE a(1)=(a11,a21)a(1) = (a_1^1,a_2^1)5, ICE a(1)=(a11,a21)a(1) = (a_1^1,a_2^1)6, LBR-ML a(1)=(a11,a21)a(1) = (a_1^1,a_2^1)7, with CE-ML also giving the best KL and prediction accuracy.
  • E3 (Traffic with signaling device): CE-ML identifies the correct mixture vertex (a(1)=(a11,a21)a(1) = (a_1^1,a_2^1)8) and 86.4% decision accuracy (ICE: 41.8%).
  • E4 (No coordination): CE-ML fails (non-CE data), but LBR-ML with fitted a(1)=(a11,a21)a(1) = (a_1^1,a_2^1)9 attains 72.6% accuracy, equaling the best fixed rationality baseline.

7. Comparison to Logit Best Response ML and Practical Recommendations

Criterion CE-ML LBR-ML
Interpretability Explicit payoff-ratio (α,β); mixture vertex interpretable Includes rationality λ; models stochastic adaptation
Computational Cost a(2)=(a11,a22)a(2) = (a_1^1,a_2^2)0, closed-form solution Requires a(2)=(a11,a22)a(2) = (a_1^1,a_2^2)1 linear system, nonconvex optimization, a(2)=(a11,a22)a(2) = (a_1^1,a_2^2)2
Behavioral Assumptions One-shot correlation device, perfect regret consistency Repeated logit best responses, bounded rationality; robust to non-CE behavior

CE-ML achieves fast, closed-form inverse learning for small a(2)=(a11,a22)a(2) = (a_1^1,a_2^2)3 games when agent behavior plausibly arises from a CE, such as in coordinated, signaled, or regulated environments. In settings without a central correlating device—such as unregulated traffic or when stochastic adaptation is prominent—LBR-ML better captures bounded rationality and noisy, non-equilibrium patterns, albeit with higher computational overhead and additional parameters (Salazar et al., 15 Jan 2026).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Closed-form Correlated Equilibrium Maximum-Likelihood Estimator (CE-ML).