Level-k Game Theory: Bounded Rationality
- Level-k Game Theory is a family of iterative models where players at higher levels best respond to the strategies of lower-level agents.
- The framework incorporates quantal best responses, using a logit function to capture noisy strategic optimization with a precision parameter.
- Empirical studies show that refining the level-0 specification improves predictive accuracy, with applications ranging from auctions to dynamic traffic models.
Searching arXiv for recent and foundational papers on level-k game theory. Level-k game theory is a family of behavioral models in which players are partitioned into discrete reasoning levels. A level-0 player follows an exogenously specified nonstrategic rule, while a level- player, for , chooses a best response, or in quantal variants a quantal best response, to beliefs about lower-level opponents. In unrepeated simultaneous-move games, iterative models such as level-k, cognitive hierarchy, and quantal cognitive hierarchy are described as the state of the art for predicting human play, but their predictions depend critically on how level-0 behavior is specified (Wright et al., 2016).
1. Formal structure of level-k reasoning
In the normal-form setting, a game is written as , where is a finite player set, each player has a finite action set , is the profile space, and is player 's payoff at profile . A behavioral model assigns to each player 0 a mixed strategy 1. Iterative models then define a hierarchy of types 2: level-0 players use an exogenous rule 3, and higher levels reason about lower ones. In the canonical level-k recursion, one fixes a naïve rule 4, and then for 5,
6
This formulation makes level-k a non-equilibrium model of bounded strategic depth rather than a fixed-point concept (Wright et al., 2016).
A central refinement replaces exact best response with quantal best response. In Wright and Leyton-Brown’s formulation, the logit quantal best response of player 7 to opponents’ mixed strategy 8 is
9
where 0 is a precision parameter; 1 yields exact best response and 2 yields uniform randomization. This quantal form underlies quantal cognitive hierarchy models and gives a continuous way to represent noisy strategic optimization (Wright et al., 2016).
Standard level-k also imposes a specific epistemic asymmetry: a level-3 player assumes opponents are exactly level-4. The "Bridging Level-K to Nash Equilibrium" model characterizes this as a substantive restriction, because standard level-k does not allow a player to put weight on equally sophisticated opponents or on Nash beliefs (Levin et al., 2022).
2. The level-0 problem
The specification of level-0 behavior is foundational because all higher-level strategies are grounded in responses to it. Wright and Leyton-Brown argue that almost all existing work specifies level-0 as a uniform distribution over actions, even though in most games it is not plausible that even nonstrategic agents would choose an action uniformly at random, nor that other agents would expect them to do so. Their paper therefore replaces uniform level-0 with a payoff-based feature model that can be computed from any normal-form game (Wright et al., 2016).
The candidate features are six general-purpose payoff summaries, each considered in binary and real-valued forms: maxmin payoff, maxmax payoff, minimax regret, minmin unfairness, max symmetric in symmetric games, and maxmax total welfare. After normalizing feature values to 5 and dropping any feature constant across a player’s action set, the level-0 distribution is defined by a linear softmax: 6 Here each feature weight satisfies 7, and the remaining mass is assigned to a uniform-noise weight 8 (Wright et al., 2016).
Weights and model configurations were learned with SMAC by maximizing cross-validated log-likelihood on the “All10” dataset under a 9-fold stratified cross-validation protocol. The data consist of 142 distinct two-player normal-form games from 10 published experiments and 13,863 human-player observations, with payoffs normalized to expected cents. After approximately 0 hours on 1 machines, the best configurations were a “4-feature” model using binary maxmax, maxmin, fairness, and symmetry features, and an “8-feature” model augmenting those four with their real-valued counterparts. Moving from uniform level-0 to the 4-feature model yielded an approximately 2–3 log-likelihood improvement overall, with an additional modest boost from the 8-feature variant; all gains were statistically significant under 4 confidence intervals from 5 cross-validation (Wright et al., 2016).
3. Variants and related solution concepts
Level-k belongs to a wider family of iterative models that differ mainly in their treatment of opponents’ sophistication and in the status of equilibrium reasoning.
| Concept | Core belief about opponents | Relation highlighted in the literature |
|---|---|---|
| Level-k | Opponents are exactly level-6 | Non-equilibrium recursive best response |
| Cognitive hierarchy / QCH | Opponents are a mixture of lower levels | Smooths lower-level beliefs; QCH uses quantal response |
| NLK | Opponent is naïve with probability 7, equally sophisticated with probability 8 | Bridges Level-1 and Nash |
| Strong level-k | At each information set, assign opponents the highest lower level consistent with arrival there | Extensive-form refinement of normal-form level-k outcomes |
Cognitive hierarchy differs from level-k by replacing the degenerate belief on level-9 with a distribution over lower levels. In the auction analysis of Rasooly, cognitive hierarchy is described as assuming each player’s level is Poisson-distributed and that a player best-responds to a “smoothed” mix of levels 0, rather than to a single lower type. In Wright and Leyton-Brown’s Spike-Poisson QCH, the level distribution has a Poisson-plus-spike form with mass 1 at 2, and 3 for 4; higher-level strategies are then generated by repeated quantal best responses to truncated mixtures of lower levels (Wright et al., 2016).
The NLK model introduces a different modification. A profile 5 is a 6-NLK equilibrium if each player best-responds to a virtual mixture in which the opponent is naïve with probability 7 and another NLK player with probability 8. When 9, the model collapses to Nash equilibrium; when 0, it collapses to Level-1; and the generalized NLK1 nests Level-2 at 3 and Nash at 4. This makes NLK a one-parameter bridge between the two theories rather than a simple additional level in the standard hierarchy (Levin et al., 2022).
In extensive-form games, Schipper and Zhou define “strong level-k thinking.” A level-5 player, at any reached information set, assigns to opponents the highest level 6 consistent with that information set, reverting to the first-level belief system if no lower level can reach it. This concept refines normal-form level-k outcomes in the sense that
7
but it has no general inclusion relation with strong rationalizability, 8-rationalizability, backward rationalizability, or backward level-k thinking. The comparison is important because it separates forward-induction-style updating from both dominance-based and equilibrium-based reasoning in trees (Schipper et al., 2023).
4. Empirical performance and domain limits
The empirical case for level-k reasoning is strongest in unrepeated normal-form games when level-0 is modeled carefully. Wright and Leyton-Brown evaluate level-k, Poisson cognitive hierarchy, and Spike-Poisson QCH on 142 games and 13,863 observations. For each model they compare uniform, 4-feature, and 8-feature level-0 specifications using cross-validated log-likelihood per observation. The 4-feature level-0 model improves predictive accuracy by approximately 9–0 overall, and the weaker iterative models, namely level-k and Poisson-CH, gain even more, closing most of their gap to QCH. The result is not merely a better calibration of level-0 mass: because higher-level play is built recursively from level-0, changing the anchor alters the entire hierarchy (Wright et al., 2016).
Auction environments provide a sharply different assessment. Rasooly studies two settings designed to disentangle level-k from Bayes–Nash equilibrium: a discrete all-pay auction with 1, values uniformly distributed on 2, and a discrete first-price auction with bid cancellation, 3, 4, values uniformly distributed on 5. Under the standard level-k specification with uniform level-0, low-6 bidding functions are step-like and extremely low: in the all-pay environment, level-1 bids 7 for all values, level-2 overcuts by bidding 8 when 9, and more generally 0 if 1 and 2 if 3, for sufficiently small 4. Empirically, those predictions fail badly. Level-k with 5 predicts bids nearly zero on average, root-mean-square errors are approximately 6–7 versus approximately 8–9 for equilibrium, equilibrium-only models have better log-likelihood and BIC than any 0 model, and forcing level-k to fit requires implausibly high estimated levels around 1–2. Hybrid models assign approximately 3–4 of mass to equilibrium types, and estimated levels in auctions are essentially uncorrelated with levels elicited from a modified 11–20 game. Rasooly concludes that, despite notable success in other strategic settings, level-k and cognitive hierarchy cannot explain behavior in these auctions, and suggests that the lack of a psychologically plausible level-0 anchor is likely a major reason (Rasooly, 2021).
This contrast matters methodologically. It indicates that level-k is not a domain-invariant substitute for equilibrium, and that its success depends jointly on the informational structure of the game, the plausibility of the level-0 anchor, and the availability of a psychologically meaningful recursive decomposition.
5. Dynamic and extensive-form generalizations
Level-k reasoning has been extended beyond one-shot normal-form interaction. In the traffic-control model of Wang, Li, and Zhao, the environment is an infinite-horizon, complete-information, sequential game at an unsignalized intersection. The player set is the set of vehicles 5, the state 6 records the system at decision epoch 7, each vehicle has a finite action set such as 8, and discounted payoff is
9
with 0 for a stage cost combining delay, collision risk, and speed penalty. Level-0 is an instinctive rule that ignores interactions, while level-1 chooses a best response to the level-2 strategy in the current state: 3 The solution procedure recursively computes policies for 4 over reachable states and can terminate when 5. In simulations with up to 6 vehicles, level-1 reduces mean delay by approximately 7 and increases throughput by approximately 8 relative to level-0, while level-2 and level-3 differ by less than 9 from level-1; assigning emergency vehicles higher queuing priority cuts their mean delay by up to 00 without significant penalty to overall flow (Wang et al., 2020).
Extensive-form generalizations pursue a different issue: how a player should update beliefs about opponents’ sophistication when play itself reveals that some lower levels are impossible. Strong level-k thinking addresses precisely this question. In a two-step centipede example with uniform first-level beliefs, strong level-1 yields 01, but strong level-2 and all higher levels yield 02, because reaching the second player’s information set causes the first player to attribute a lower level consistent with that arrival. Schipper and Zhou use this apparatus to reanalyze Battle-of-the-Sexes-with-Outside-Option experiments. In Cooper et al. (1993), strong level-3 with uniform first-level beliefs predicts 03, and in the last 11 rounds 04 of row-player In choices and 05 of column-player “1” choices match that prediction. In other datasets, player 1 conforms more closely than player 2, indicating that the uniform level-0 premise may be role-dependent even when the extensive-form refinement is behaviorally relevant (Schipper et al., 2023).
Taken together, these developments show that “level” can index either bounded depth of strategic anticipation in dynamic control problems or bounded depth combined with path-dependent belief revision in trees. The two uses are compatible, but they solve different problems.
6. Level-k reasoning and LLMs
Recent work has used level-k both as a diagnostic language for LLM behavior and as an explicit prompting framework for strategic inference. In "K-Level Reasoning: Establishing Higher Order Beliefs in LLMs for Strategic Reasoning" (Zhang et al., 2024), the game unfolds over rounds 06, the environment updates as 07, and the level-08 decision is defined recursively by
09
The recursion is implemented by prompt orchestration rather than a new neural architecture: the model is repeatedly queried to simulate opponents’ lower-level moves and then to best-respond. Tested on a ten-round beauty contest and a ten-day survival auction, K-Level with 10 outperforms direct prompting, chain-of-thought, persona, reflection, refinement, and prediction-CoT baselines. In the beauty contest, its win-rate is 11 versus 12 for Direct and 13 for PCoT, and its adaptation index is 14 versus 15 for PCoT. In the survival auction, average survival round is 16 for K-Level versus 17 for Direct. Depth sensitivity is also consistent with classic level-k intuition: being one level ahead helps, while excessive depth can hurt against simpler opponents (Zhang et al., 2024).
A distinct line of work asks whether LLMs naturally behave like human level-k players. "LLM Agents as Static Level-k Players in Behavioural Games" (Teo, 26 Jun 2026) studies a 360-cell factorial over model scale, temperature, quantisation, alignment, and framing in a p-beauty contest and a public-goods game. The key result is not merely that LLM outputs can be mapped to level-k anchors, but that each fixed deployment setting produces a largely static reasoning depth. In the beauty contest, 18B models peak near 19 (level-0), 20B near 21 (level-1), 22B near 23 (level-2), and 24B–25B near zero; raising temperature broadens the spike slightly but does not create additional modes. Pooling across heterogeneous settings can recover human-like dispersion, but no single cell matches the full human distribution, with Kolmogorov–Smirnov distances around 26–27. In the public-goods game, humans begin around 28–29 and then contributions decay by approximately 30 tokens per round, whereas LLM contributions stay flat or rise slightly; when the final round is explicit, no last-round defection appears. The authors interpret this as evidence that current LLMs act as static, category-retrieved level-k players rather than as agents performing within-game belief updating or backward induction (Teo, 26 Jun 2026).
These computational uses reinforce a longstanding theoretical point. Level-k is not only a descriptive hierarchy over human strategic sophistication; it is also a modular template for constructing bounded recursive reasoners. The open question, highlighted in both the behavioral and LLM literatures, is how to augment that template with endogenous belief revision, richer level-0 anchors, and game-dependent heterogeneity without losing its tractability.