Papers
Topics
Authors
Recent
Search
2000 character limit reached

Level-k Game Theory: Bounded Rationality

Updated 12 July 2026
  • Level-k Game Theory is a family of iterative models where players at higher levels best respond to the strategies of lower-level agents.
  • The framework incorporates quantal best responses, using a logit function to capture noisy strategic optimization with a precision parameter.
  • Empirical studies show that refining the level-0 specification improves predictive accuracy, with applications ranging from auctions to dynamic traffic models.

Searching arXiv for recent and foundational papers on level-k game theory. Level-k game theory is a family of behavioral models in which players are partitioned into discrete reasoning levels. A level-0 player follows an exogenously specified nonstrategic rule, while a level-kk player, for k1k\ge 1, chooses a best response, or in quantal variants a quantal best response, to beliefs about lower-level opponents. In unrepeated simultaneous-move games, iterative models such as level-k, cognitive hierarchy, and quantal cognitive hierarchy are described as the state of the art for predicting human play, but their predictions depend critically on how level-0 behavior is specified (Wright et al., 2016).

1. Formal structure of level-k reasoning

In the normal-form setting, a game is written as G=(N,A,u)G=(N,A,u), where NN is a finite player set, each player ii has a finite action set AiA_i, A=×iAiA=\times_i A_i is the profile space, and ui(a)Ru_i(a)\in\mathbb{R} is player ii's payoff at profile aa. A behavioral model assigns to each player k1k\ge 10 a mixed strategy k1k\ge 11. Iterative models then define a hierarchy of types k1k\ge 12: level-0 players use an exogenous rule k1k\ge 13, and higher levels reason about lower ones. In the canonical level-k recursion, one fixes a naïve rule k1k\ge 14, and then for k1k\ge 15,

k1k\ge 16

This formulation makes level-k a non-equilibrium model of bounded strategic depth rather than a fixed-point concept (Wright et al., 2016).

A central refinement replaces exact best response with quantal best response. In Wright and Leyton-Brown’s formulation, the logit quantal best response of player k1k\ge 17 to opponents’ mixed strategy k1k\ge 18 is

k1k\ge 19

where G=(N,A,u)G=(N,A,u)0 is a precision parameter; G=(N,A,u)G=(N,A,u)1 yields exact best response and G=(N,A,u)G=(N,A,u)2 yields uniform randomization. This quantal form underlies quantal cognitive hierarchy models and gives a continuous way to represent noisy strategic optimization (Wright et al., 2016).

Standard level-k also imposes a specific epistemic asymmetry: a level-G=(N,A,u)G=(N,A,u)3 player assumes opponents are exactly level-G=(N,A,u)G=(N,A,u)4. The "Bridging Level-K to Nash Equilibrium" model characterizes this as a substantive restriction, because standard level-k does not allow a player to put weight on equally sophisticated opponents or on Nash beliefs (Levin et al., 2022).

2. The level-0 problem

The specification of level-0 behavior is foundational because all higher-level strategies are grounded in responses to it. Wright and Leyton-Brown argue that almost all existing work specifies level-0 as a uniform distribution over actions, even though in most games it is not plausible that even nonstrategic agents would choose an action uniformly at random, nor that other agents would expect them to do so. Their paper therefore replaces uniform level-0 with a payoff-based feature model that can be computed from any normal-form game (Wright et al., 2016).

The candidate features are six general-purpose payoff summaries, each considered in binary and real-valued forms: maxmin payoff, maxmax payoff, minimax regret, minmin unfairness, max symmetric in symmetric games, and maxmax total welfare. After normalizing feature values to G=(N,A,u)G=(N,A,u)5 and dropping any feature constant across a player’s action set, the level-0 distribution is defined by a linear softmax: G=(N,A,u)G=(N,A,u)6 Here each feature weight satisfies G=(N,A,u)G=(N,A,u)7, and the remaining mass is assigned to a uniform-noise weight G=(N,A,u)G=(N,A,u)8 (Wright et al., 2016).

Weights and model configurations were learned with SMAC by maximizing cross-validated log-likelihood on the “All10” dataset under a G=(N,A,u)G=(N,A,u)9-fold stratified cross-validation protocol. The data consist of 142 distinct two-player normal-form games from 10 published experiments and 13,863 human-player observations, with payoffs normalized to expected cents. After approximately NN0 hours on NN1 machines, the best configurations were a “4-feature” model using binary maxmax, maxmin, fairness, and symmetry features, and an “8-feature” model augmenting those four with their real-valued counterparts. Moving from uniform level-0 to the 4-feature model yielded an approximately NN2–NN3 log-likelihood improvement overall, with an additional modest boost from the 8-feature variant; all gains were statistically significant under NN4 confidence intervals from NN5 cross-validation (Wright et al., 2016).

Level-k belongs to a wider family of iterative models that differ mainly in their treatment of opponents’ sophistication and in the status of equilibrium reasoning.

Concept Core belief about opponents Relation highlighted in the literature
Level-k Opponents are exactly level-NN6 Non-equilibrium recursive best response
Cognitive hierarchy / QCH Opponents are a mixture of lower levels Smooths lower-level beliefs; QCH uses quantal response
NLK Opponent is naïve with probability NN7, equally sophisticated with probability NN8 Bridges Level-1 and Nash
Strong level-k At each information set, assign opponents the highest lower level consistent with arrival there Extensive-form refinement of normal-form level-k outcomes

Cognitive hierarchy differs from level-k by replacing the degenerate belief on level-NN9 with a distribution over lower levels. In the auction analysis of Rasooly, cognitive hierarchy is described as assuming each player’s level is Poisson-distributed and that a player best-responds to a “smoothed” mix of levels ii0, rather than to a single lower type. In Wright and Leyton-Brown’s Spike-Poisson QCH, the level distribution has a Poisson-plus-spike form with mass ii1 at ii2, and ii3 for ii4; higher-level strategies are then generated by repeated quantal best responses to truncated mixtures of lower levels (Wright et al., 2016).

The NLK model introduces a different modification. A profile ii5 is a ii6-NLK equilibrium if each player best-responds to a virtual mixture in which the opponent is naïve with probability ii7 and another NLK player with probability ii8. When ii9, the model collapses to Nash equilibrium; when AiA_i0, it collapses to Level-1; and the generalized NLKAiA_i1 nests Level-AiA_i2 at AiA_i3 and Nash at AiA_i4. This makes NLK a one-parameter bridge between the two theories rather than a simple additional level in the standard hierarchy (Levin et al., 2022).

In extensive-form games, Schipper and Zhou define “strong level-k thinking.” A level-AiA_i5 player, at any reached information set, assigns to opponents the highest level AiA_i6 consistent with that information set, reverting to the first-level belief system if no lower level can reach it. This concept refines normal-form level-k outcomes in the sense that

AiA_i7

but it has no general inclusion relation with strong rationalizability, AiA_i8-rationalizability, backward rationalizability, or backward level-k thinking. The comparison is important because it separates forward-induction-style updating from both dominance-based and equilibrium-based reasoning in trees (Schipper et al., 2023).

4. Empirical performance and domain limits

The empirical case for level-k reasoning is strongest in unrepeated normal-form games when level-0 is modeled carefully. Wright and Leyton-Brown evaluate level-k, Poisson cognitive hierarchy, and Spike-Poisson QCH on 142 games and 13,863 observations. For each model they compare uniform, 4-feature, and 8-feature level-0 specifications using cross-validated log-likelihood per observation. The 4-feature level-0 model improves predictive accuracy by approximately AiA_i9–A=×iAiA=\times_i A_i0 overall, and the weaker iterative models, namely level-k and Poisson-CH, gain even more, closing most of their gap to QCH. The result is not merely a better calibration of level-0 mass: because higher-level play is built recursively from level-0, changing the anchor alters the entire hierarchy (Wright et al., 2016).

Auction environments provide a sharply different assessment. Rasooly studies two settings designed to disentangle level-k from Bayes–Nash equilibrium: a discrete all-pay auction with A=×iAiA=\times_i A_i1, values uniformly distributed on A=×iAiA=\times_i A_i2, and a discrete first-price auction with bid cancellation, A=×iAiA=\times_i A_i3, A=×iAiA=\times_i A_i4, values uniformly distributed on A=×iAiA=\times_i A_i5. Under the standard level-k specification with uniform level-0, low-A=×iAiA=\times_i A_i6 bidding functions are step-like and extremely low: in the all-pay environment, level-1 bids A=×iAiA=\times_i A_i7 for all values, level-2 overcuts by bidding A=×iAiA=\times_i A_i8 when A=×iAiA=\times_i A_i9, and more generally ui(a)Ru_i(a)\in\mathbb{R}0 if ui(a)Ru_i(a)\in\mathbb{R}1 and ui(a)Ru_i(a)\in\mathbb{R}2 if ui(a)Ru_i(a)\in\mathbb{R}3, for sufficiently small ui(a)Ru_i(a)\in\mathbb{R}4. Empirically, those predictions fail badly. Level-k with ui(a)Ru_i(a)\in\mathbb{R}5 predicts bids nearly zero on average, root-mean-square errors are approximately ui(a)Ru_i(a)\in\mathbb{R}6–ui(a)Ru_i(a)\in\mathbb{R}7 versus approximately ui(a)Ru_i(a)\in\mathbb{R}8–ui(a)Ru_i(a)\in\mathbb{R}9 for equilibrium, equilibrium-only models have better log-likelihood and BIC than any ii0 model, and forcing level-k to fit requires implausibly high estimated levels around ii1–ii2. Hybrid models assign approximately ii3–ii4 of mass to equilibrium types, and estimated levels in auctions are essentially uncorrelated with levels elicited from a modified 11–20 game. Rasooly concludes that, despite notable success in other strategic settings, level-k and cognitive hierarchy cannot explain behavior in these auctions, and suggests that the lack of a psychologically plausible level-0 anchor is likely a major reason (Rasooly, 2021).

This contrast matters methodologically. It indicates that level-k is not a domain-invariant substitute for equilibrium, and that its success depends jointly on the informational structure of the game, the plausibility of the level-0 anchor, and the availability of a psychologically meaningful recursive decomposition.

5. Dynamic and extensive-form generalizations

Level-k reasoning has been extended beyond one-shot normal-form interaction. In the traffic-control model of Wang, Li, and Zhao, the environment is an infinite-horizon, complete-information, sequential game at an unsignalized intersection. The player set is the set of vehicles ii5, the state ii6 records the system at decision epoch ii7, each vehicle has a finite action set such as ii8, and discounted payoff is

ii9

with aa0 for a stage cost combining delay, collision risk, and speed penalty. Level-0 is an instinctive rule that ignores interactions, while level-aa1 chooses a best response to the level-aa2 strategy in the current state: aa3 The solution procedure recursively computes policies for aa4 over reachable states and can terminate when aa5. In simulations with up to aa6 vehicles, level-1 reduces mean delay by approximately aa7 and increases throughput by approximately aa8 relative to level-0, while level-2 and level-3 differ by less than aa9 from level-1; assigning emergency vehicles higher queuing priority cuts their mean delay by up to k1k\ge 100 without significant penalty to overall flow (Wang et al., 2020).

Extensive-form generalizations pursue a different issue: how a player should update beliefs about opponents’ sophistication when play itself reveals that some lower levels are impossible. Strong level-k thinking addresses precisely this question. In a two-step centipede example with uniform first-level beliefs, strong level-1 yields k1k\ge 101, but strong level-2 and all higher levels yield k1k\ge 102, because reaching the second player’s information set causes the first player to attribute a lower level consistent with that arrival. Schipper and Zhou use this apparatus to reanalyze Battle-of-the-Sexes-with-Outside-Option experiments. In Cooper et al. (1993), strong level-3 with uniform first-level beliefs predicts k1k\ge 103, and in the last 11 rounds k1k\ge 104 of row-player In choices and k1k\ge 105 of column-player “1” choices match that prediction. In other datasets, player 1 conforms more closely than player 2, indicating that the uniform level-0 premise may be role-dependent even when the extensive-form refinement is behaviorally relevant (Schipper et al., 2023).

Taken together, these developments show that “level” can index either bounded depth of strategic anticipation in dynamic control problems or bounded depth combined with path-dependent belief revision in trees. The two uses are compatible, but they solve different problems.

6. Level-k reasoning and LLMs

Recent work has used level-k both as a diagnostic language for LLM behavior and as an explicit prompting framework for strategic inference. In "K-Level Reasoning: Establishing Higher Order Beliefs in LLMs for Strategic Reasoning" (Zhang et al., 2024), the game unfolds over rounds k1k\ge 106, the environment updates as k1k\ge 107, and the level-k1k\ge 108 decision is defined recursively by

k1k\ge 109

The recursion is implemented by prompt orchestration rather than a new neural architecture: the model is repeatedly queried to simulate opponents’ lower-level moves and then to best-respond. Tested on a ten-round beauty contest and a ten-day survival auction, K-Level with k1k\ge 110 outperforms direct prompting, chain-of-thought, persona, reflection, refinement, and prediction-CoT baselines. In the beauty contest, its win-rate is k1k\ge 111 versus k1k\ge 112 for Direct and k1k\ge 113 for PCoT, and its adaptation index is k1k\ge 114 versus k1k\ge 115 for PCoT. In the survival auction, average survival round is k1k\ge 116 for K-Level versus k1k\ge 117 for Direct. Depth sensitivity is also consistent with classic level-k intuition: being one level ahead helps, while excessive depth can hurt against simpler opponents (Zhang et al., 2024).

A distinct line of work asks whether LLMs naturally behave like human level-k players. "LLM Agents as Static Level-k Players in Behavioural Games" (Teo, 26 Jun 2026) studies a 360-cell factorial over model scale, temperature, quantisation, alignment, and framing in a p-beauty contest and a public-goods game. The key result is not merely that LLM outputs can be mapped to level-k anchors, but that each fixed deployment setting produces a largely static reasoning depth. In the beauty contest, k1k\ge 118B models peak near k1k\ge 119 (level-0), k1k\ge 120B near k1k\ge 121 (level-1), k1k\ge 122B near k1k\ge 123 (level-2), and k1k\ge 124B–k1k\ge 125B near zero; raising temperature broadens the spike slightly but does not create additional modes. Pooling across heterogeneous settings can recover human-like dispersion, but no single cell matches the full human distribution, with Kolmogorov–Smirnov distances around k1k\ge 126–k1k\ge 127. In the public-goods game, humans begin around k1k\ge 128–k1k\ge 129 and then contributions decay by approximately k1k\ge 130 tokens per round, whereas LLM contributions stay flat or rise slightly; when the final round is explicit, no last-round defection appears. The authors interpret this as evidence that current LLMs act as static, category-retrieved level-k players rather than as agents performing within-game belief updating or backward induction (Teo, 26 Jun 2026).

These computational uses reinforce a longstanding theoretical point. Level-k is not only a descriptive hierarchy over human strategic sophistication; it is also a modular template for constructing bounded recursive reasoners. The open question, highlighted in both the behavioral and LLM literatures, is how to augment that template with endogenous belief revision, richer level-0 anchors, and game-dependent heterogeneity without losing its tractability.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Level-k Game Theory.