Game Arena: Strategic LLM Evaluation in Competitive Environments
This presentation introduces Game Arena, a benchmark that evaluates large language models through competitive gameplay rather than static tests. By measuring performance across Chess, Poker, and Werewolf, the benchmark reveals that strategic competence is not a single transferable skill: models that dominate deterministic planning fail at probabilistic reasoning, and rankings reverse completely across environments. The work demonstrates how objective game outcomes, large-scale variance control, and role-sensitive metrics can expose capability differences that traditional benchmarks miss.Script
The strongest Chess-playing language model loses money at Poker. The best Poker player struggles at social deduction. Game Arena reveals that strategic intelligence fragments across competitive environments in ways that no single benchmark can capture.
The authors built an open tournament system where models play thousands of games under a uniform text protocol. Every turn is recorded, every invalid move triggers a retry, and bootstrap resampling over complete game trajectories produces statistically grounded confidence intervals for each metric.
In Chess, models stay roughly even through the opening, but Gemini 3 Pro Preview pulls ahead in the middlegame and converts its advantage in the endgame. Weaker models like Claude Haiku require more than 10 retries per game and their win probability collapses as positions grow complex.
Poker produces a ranking inversion. GPT 5.2 earns 46.6 big blinds per hundred hands and beats every opponent, while Gemini 3 Pro Preview, the Chess leader, loses chips at negative 15.2. Aggression alone does not explain the gap: GPT 5 mini raises 98 percent of hands and loses catastrophically, while GPT 5.2 raises 92 percent and dominates.
Werewolf introduced role-sensitive evaluation through Game Theoretic Evaluation, which decomposes each model's rating into contributions from Werewolf, Seer, Doctor, and Villager roles. Gemini 3 Pro Preview leads overall, but Claude Sonnet and Grok 4.1 fail specifically as Seers, while GPT 5 mini contributes negatively in every role.
No single model dominates the cost-performance frontier across all three games, which means environment-specific benchmarks are not luxuries but necessities. You can explore the full leaderboard, replay game trajectories, and build your own competitive evaluations at EmergentMind.com.