Papers
Topics
Authors
Recent
Search
2000 character limit reached

Performatively Stable Equilibrium

Updated 9 July 2026
  • Performatively stable equilibrium is a fixed-point solution concept where a model or strategy is optimal for the distribution it induces.
  • It formalizes stability via repeated retraining dynamics and fixed-point conditions in both single-agent and multiplayer settings.
  • Extensions to neural networks and mixed equilibria reveal computational thresholds and convergence challenges in decision-dependent environments.

to=arxiv_search 天天中彩票微信 򐂕json {"query":"\"performatively stable equilibrium\" OR \"performatively stable\" performative prediction", "max_results": 10} to=arxiv_search 񎔲json {"query":"(Cetin et al., 2023) OR (Mofakhami et al., 2023) OR (Farina et al., 27 Feb 2026) OR (Zhong et al., 1 Sep 2025) OR (Anagnostides et al., 28 Jan 2026) OR (Narang et al., 2022) OR (Chen et al., 6 Oct 2025)", "max_results": 10} Performatively stable equilibrium is a fixed-point solution concept for settings in which deployed decisions alter the data distribution or environment that later governs retraining. In its most basic form, a model, policy, or strategy profile is performatively stable when it is optimal for the distribution or environment induced by itself. In single-agent performative prediction, this is the fixed point of repeated risk minimization; in multiplayer settings, it is the Nash equilibrium of the frozen game induced by the candidate profile; and in recent online-learning formulations, the notion extends to mixed equilibria over model iterates. Across these formulations, performative stability formalizes the equilibrium reached by myopic retraining under endogenous distribution shift, rather than the solution of a fully strategic or globally performative objective (Mofakhami et al., 2023, Narang et al., 2022, Farina et al., 27 Feb 2026).

1. Formal definitions and equivalent fixed-point views

In the performative prediction framework, a deployed model θΘRd\theta \in \Theta \subseteq \mathbb{R}^d induces a distribution D(θ)D(\theta), and the performative risk is

PR(θ)=EzD(θ)[(z;θ)].PR(\theta) = \mathbb{E}_{z \sim D(\theta)}[\ell(z;\theta)].

A model is performatively stable if it is optimal for the distribution it induces. Equivalently, with the retraining operator

G(θ)=argminθΘEzD(θ)[(z;θ)],G(\theta)=\arg\min_{\theta' \in \Theta}\mathbb{E}_{z\sim D(\theta)}[\ell(z;\theta')],

performative stability is the fixed-point condition θ=G(θ)\theta = G(\theta). A standard first-order relaxation used in convex settings is

xx, EzD(x)[x(x;z)]ϵxX,\langle x' - x,\ \mathbb E_{z\sim \mathcal D(x)}[\nabla_x \ell(x; z)] \rangle \ge -\epsilon \quad \forall x' \in \mathcal X,

which coincides with exact performative stability when (;z)\ell(\cdot;z) is convex (Anagnostides et al., 28 Jan 2026).

A closely related formulation uses the decoupled performative risk

DPR(θ,D(θ))=EzD(θ)[(z;θ)],DPR(\theta, D(\theta')) = \mathbb{E}_{z \sim D(\theta')}[\ell(z;\theta)],

which makes explicit the distinction between evaluating a model against its own induced distribution and evaluating it against a frozen distribution induced by another model. This decoupling is central to both repeated retraining analyses and online-to-batch reductions (Farina et al., 27 Feb 2026).

In multiplayer performative prediction, player ii chooses xiXix_i\in X_i, the joint action is D(θ)D(\theta)0, and player D(θ)D(\theta)1’s loss is

D(θ)D(\theta)2

For each reference point D(θ)D(\theta)3, the static frozen game D(θ)D(\theta)4 is defined by

D(θ)D(\theta)5

A profile D(θ)D(\theta)6 is a performatively stable equilibrium if it is a Nash equilibrium of D(θ)D(\theta)7, equivalently

D(θ)D(\theta)8

In decision-dependent games this same idea is written as

D(θ)D(\theta)9

so that the profile is a Nash equilibrium of the game induced by its own distribution (Narang et al., 2022, Zhong et al., 1 Sep 2025).

Performative stability is distinct from both Nash equilibrium and performative optimality. In multiplayer games, Nash equilibrium requires the full derivative of each player’s loss with respect to the induced distribution response, whereas performative stability keeps the distribution frozen at the candidate profile. In performative reinforcement learning, the performatively stable policy is explicitly described as a surrogate target rather than the desired objective, because the desired notion is the performatively optimal policy that maximizes the true matched policy-environment value (Narang et al., 2022, Chen et al., 6 Oct 2025).

2. Retraining dynamics, contraction regimes, and stability criteria

The canonical dynamic associated with performative stability is repeated retraining. In the single-agent case,

PR(θ)=EzD(θ)[(z;θ)].PR(\theta) = \mathbb{E}_{z \sim D(\theta)}[\ell(z;\theta)].0

and in multiplayer settings,

PR(θ)=EzD(θ)[(z;θ)].PR(\theta) = \mathbb{E}_{z \sim D(\theta)}[\ell(z;\theta)].1

Performatively stable equilibria are precisely the fixed points of these updates (Mofakhami et al., 2023, Narang et al., 2022).

Classical positive results establish convergence by proving that the retraining map is a contraction. In the parameter-space formulation of performative prediction, if

PR(θ)=EzD(θ)[(z;θ)].PR(\theta) = \mathbb{E}_{z \sim D(\theta)}[\ell(z;\theta)].2

and the loss is PR(θ)=EzD(θ)[(z;θ)].PR(\theta) = \mathbb{E}_{z \sim D(\theta)}[\ell(z;\theta)].3-strongly convex and PR(θ)=EzD(θ)[(z;θ)].PR(\theta) = \mathbb{E}_{z \sim D(\theta)}[\ell(z;\theta)].4-smooth in PR(θ)=EzD(θ)[(z;θ)].PR(\theta) = \mathbb{E}_{z \sim D(\theta)}[\ell(z;\theta)].5, then the retraining operator satisfies

PR(θ)=EzD(θ)[(z;θ)].PR(\theta) = \mathbb{E}_{z \sim D(\theta)}[\ell(z;\theta)].6

so repeated risk minimization converges when PR(θ)=EzD(θ)[(z;θ)].PR(\theta) = \mathbb{E}_{z \sim D(\theta)}[\ell(z;\theta)].7 (Mofakhami et al., 2023).

The multiplayer analogue is governed by the quantity

PR(θ)=EzD(θ)[(z;θ)].PR(\theta) = \mathbb{E}_{z \sim D(\theta)}[\ell(z;\theta)].8

Under the standing assumptions that each frozen game is PR(θ)=EzD(θ)[(z;θ)].PR(\theta) = \mathbb{E}_{z \sim D(\theta)}[\ell(z;\theta)].9-strongly monotone, each G(θ)=argminθΘEzD(θ)[(z;θ)],G(\theta)=\arg\min_{\theta' \in \Theta}\mathbb{E}_{z\sim D(\theta)}[\ell(z;\theta')],0 is Lipschitz in Wasserstein distance, and the gradient in G(θ)=argminθΘEzD(θ)[(z;θ)],G(\theta)=\arg\min_{\theta' \in \Theta}\mathbb{E}_{z\sim D(\theta)}[\ell(z;\theta')],1 is Lipschitz with constant G(θ)=argminθΘEzD(θ)[(z;θ)],G(\theta)=\arg\min_{\theta' \in \Theta}\mathbb{E}_{z\sim D(\theta)}[\ell(z;\theta')],2, repeated retraining converges linearly whenever G(θ)=argminθΘEzD(θ)[(z;θ)],G(\theta)=\arg\min_{\theta' \in \Theta}\mathbb{E}_{z\sim D(\theta)}[\ell(z;\theta')],3: G(θ)=argminθΘEzD(θ)[(z;θ)],G(\theta)=\arg\min_{\theta' \in \Theta}\mathbb{E}_{z\sim D(\theta)}[\ell(z;\theta')],4 The same work also studies repeated gradient play,

G(θ)=argminθΘEzD(θ)[(z;θ)],G(\theta)=\arg\min_{\theta' \in \Theta}\mathbb{E}_{z\sim D(\theta)}[\ell(z;\theta')],5

and repeated stochastic gradient play, obtaining linear convergence under stronger thresholds and finite-variance assumptions (Narang et al., 2022).

A more localized notion of stability appears in analyses of policy iteration and best-response maps. There, the criterion is not global contraction but local attraction of the fixed point: an equilibrium is locally stable if all sufficiently small perturbations converge back under repeated updates. This local notion is the one used in the analysis of Kyle’s discrete-time equilibrium, where the Jacobian of the policy-iteration map determines whether perturbations decay or expand (Cetin et al., 2023).

3. Prediction-space formulations and neural-network performative stability

A major reformulation of performative stability replaces sensitivity in parameter space by sensitivity in prediction space. Instead of writing the shifted distribution as G(θ)=argminθΘEzD(θ)[(z;θ)],G(\theta)=\arg\min_{\theta' \in \Theta}\mathbb{E}_{z\sim D(\theta)}[\ell(z;\theta')],6, the distribution map is written as G(θ)=argminθΘEzD(θ)[(z;θ)],G(\theta)=\arg\min_{\theta' \in \Theta}\mathbb{E}_{z\sim D(\theta)}[\ell(z;\theta')],7, where G(θ)=argminθΘEzD(θ)[(z;θ)],G(\theta)=\arg\min_{\theta' \in \Theta}\mathbb{E}_{z\sim D(\theta)}[\ell(z;\theta')],8 is the predictive function. The performative risk becomes

G(θ)=argminθΘEzD(θ)[(z;θ)],G(\theta)=\arg\min_{\theta' \in \Theta}\mathbb{E}_{z\sim D(\theta)}[\ell(z;\theta')],9

and a performatively stable classifier is again the fixed point of repeated risk minimization, but now interpreted in function space (Mofakhami et al., 2023).

The central assumption is θ=G(θ)\theta = G(\theta)0-sensitivity with respect to Pearson θ=G(θ)\theta = G(\theta)1 divergence: θ=G(θ)\theta = G(\theta)2 where

θ=G(θ)\theta = G(\theta)3

Under this formulation, the loss need only be θ=G(θ)\theta = G(\theta)4-strongly convex in the predictions θ=G(θ)\theta = G(\theta)5, and the gradient in prediction space must be bounded: θ=G(θ)\theta = G(\theta)6 This removes the requirement of convexity in the model parameters, which is the restrictive part of earlier convergence guarantees for modern neural networks (Mofakhami et al., 2023).

For convex function classes θ=G(θ)\theta = G(\theta)7, the retraining operator contracts in function space: θ=G(θ)\theta = G(\theta)8 and therefore repeated risk minimization converges linearly whenever

θ=G(θ)\theta = G(\theta)9

For nonconvex function classes, including neural-network parameterizations, the analysis gives a near-convergence result under approximation error xx, EzD(x)[x(x;z)]ϵxX,\langle x' - x,\ \mathbb E_{z\sim \mathcal D(x)}[\nabla_x \ell(x; z)] \rangle \ge -\epsilon \quad \forall x' \in \mathcal X,0: xx, EzD(x)[x(x;z)]ϵxX,\langle x' - x,\ \mathbb E_{z\sim \mathcal D(x)}[\nabla_x \ell(x; z)] \rangle \ge -\epsilon \quad \forall x' \in \mathcal X,1 so the iterates converge to a neighborhood of a stable solution when the contraction factor is below one and xx, EzD(x)[x(x;z)]ϵxX,\langle x' - x,\ \mathbb E_{z\sim \mathcal D(x)}[\nabla_x \ell(x; z)] \rangle \ge -\epsilon \quad \forall x' \in \mathcal X,2 is small (Mofakhami et al., 2023).

The paper also introduces the Resample-if-Rejected (RIR) procedure as a concrete performative shift model. One samples xx, EzD(x)[x(x;z)]ϵxX,\langle x' - x,\ \mathbb E_{z\sim \mathcal D(x)}[\nabla_x \ell(x; z)] \rangle \ge -\epsilon \quad \forall x' \in \mathcal X,3 from the base distribution xx, EzD(x)[x(x;z)]ϵxX,\langle x' - x,\ \mathbb E_{z\sim \mathcal D(x)}[\nabla_x \ell(x; z)] \rangle \ge -\epsilon \quad \forall x' \in \mathcal X,4, rejects it with probability xx, EzD(x)[x(x;z)]ϵxX,\langle x' - x,\ \mathbb E_{z\sim \mathcal D(x)}[\nabla_x \ell(x; z)] \rangle \ge -\epsilon \quad \forall x' \in \mathcal X,5, and if rejected draws another sample. The shifted density is

xx, EzD(x)[x(x;z)]ϵxX,\langle x' - x,\ \mathbb E_{z\sim \mathcal D(x)}[\nabla_x \ell(x; z)] \rangle \ge -\epsilon \quad \forall x' \in \mathcal X,6

For

xx, EzD(x)[x(x;z)]ϵxX,\langle x' - x,\ \mathbb E_{z\sim \mathcal D(x)}[\nabla_x \ell(x; z)] \rangle \ge -\epsilon \quad \forall x' \in \mathcal X,7

the map is xx, EzD(x)[x(x;z)]ϵxX,\langle x' - x,\ \mathbb E_{z\sim \mathcal D(x)}[\nabla_x \ell(x; z)] \rangle \ge -\epsilon \quad \forall x' \in \mathcal X,8-sensitive in xx, EzD(x)[x(x;z)]ϵxX,\langle x' - x,\ \mathbb E_{z\sim \mathcal D(x)}[\nabla_x \ell(x; z)] \rangle \ge -\epsilon \quad \forall x' \in \mathcal X,9, and the bounded norm ratio property holds with

(;z)\ell(\cdot;z)0

The empirical illustration uses the Kaggle “Give Me Some Credit” dataset, a two-layer neural network with a hidden layer of size (;z)\ell(\cdot;z)1, LeakyReLU nonlinearity, and a scaled sigmoid output, and repeated retraining is run approximately by gradient descent until the risk change is below (;z)\ell(\cdot;z)2. The reported behavior is that performative risk decreases and stabilizes for large (;z)\ell(\cdot;z)3, while convergence is less reliable for smaller (;z)\ell(\cdot;z)4, consistent with stronger performative effects (Mofakhami et al., 2023).

4. Multiplayer performative prediction and decision-dependent games

In multiplayer settings, performatively stable equilibrium is the equilibrium concept associated with myopic retraining rather than full strategic optimization. The distinction is made explicit by the decomposition

(;z)\ell(\cdot;z)5

A Nash equilibrium of the full decision-dependent game satisfies (;z)\ell(\cdot;z)6 for all players, whereas a performatively stable equilibrium requires only the frozen-distribution stationarity (;z)\ell(\cdot;z)7. This makes performatively stable equilibria computationally easier than Nash equilibria because the distribution-response term (;z)\ell(\cdot;z)8 is ignored (Narang et al., 2022).

The decision-dependent game formulation makes the induced-distribution structure explicit. Player (;z)\ell(\cdot;z)9 solves

DPR(θ,D(θ))=EzD(θ)[(z;θ)],DPR(\theta, D(\theta')) = \mathbb{E}_{z \sim D(\theta')}[\ell(z;\theta)],0

and the induced game at a reference point DPR(θ,D(θ))=EzD(θ)[(z;θ)],DPR(\theta, D(\theta')) = \mathbb{E}_{z \sim D(\theta')}[\ell(z;\theta)],1 is

DPR(θ,D(θ))=EzD(θ)[(z;θ)],DPR(\theta, D(\theta')) = \mathbb{E}_{z \sim D(\theta')}[\ell(z;\theta)],2

The equilibrium condition is

DPR(θ,D(θ))=EzD(θ)[(z;θ)],DPR(\theta, D(\theta')) = \mathbb{E}_{z \sim D(\theta')}[\ell(z;\theta)],3

and repeated retraining is

DPR(θ,D(θ))=EzD(θ)[(z;θ)],DPR(\theta, D(\theta')) = \mathbb{E}_{z \sim D(\theta')}[\ell(z;\theta)],4

This framework emphasizes that the feedback loop operates through the entire joint action profile, not just an individual player’s action (Zhong et al., 1 Sep 2025).

To avoid reliance on an unknown or impractical DPR(θ,D(θ))=EzD(θ)[(z;θ)],DPR(\theta, D(\theta')) = \mathbb{E}_{z \sim D(\theta')}[\ell(z;\theta)],5-smoothness constant, the paper "Towards Performatively Stable Equilibria in Decision-Dependent Games for Arbitrary Data Distribution Maps" (Zhong et al., 1 Sep 2025) introduces the gradient-based sensitivity measure DPR(θ,D(θ))=EzD(θ)[(z;θ)],DPR(\theta, D(\theta')) = \mathbb{E}_{z \sim D(\theta')}[\ell(z;\theta)],6-sensitivity: DPR(θ,D(θ))=EzD(θ)[(z;θ)],DPR(\theta, D(\theta')) = \mathbb{E}_{z \sim D(\theta')}[\ell(z;\theta)],7 With the joint expected gradient map DPR(θ,D(θ))=EzD(θ)[(z;θ)],DPR(\theta, D(\theta')) = \mathbb{E}_{z \sim D(\theta')}[\ell(z;\theta)],8, this yields

DPR(θ,D(θ))=EzD(θ)[(z;θ)],DPR(\theta, D(\theta')) = \mathbb{E}_{z \sim D(\theta')}[\ell(z;\theta)],9

Under ii0-strong monotonicity,

ii1

the retraining dynamics are contractive: ii2 and repeated retraining converges to a performatively stable equilibrium whenever

ii3

The paper further proposes the Sensitivity-Informed Repeated Retraining algorithm, SIRii4, which selects

ii5

adds the quadratic regularizer

ii6

and updates empirical sensitivities during retraining. The reported experiments cover a prediction error minimization game, Cournot competition, and a ride-share revenue maximization game. In the prediction error minimization game, SIRii7 achieves the lowest total RMSE, around ii8, and converges within about ii9 iterations; in the Cournot and ride-share tasks it converges within about xiXix_i\in X_i0 iterations and outperforms the baselines listed as RR, RGD, SFB, AGM, and OPGD (Zhong et al., 1 Sep 2025).

5. Mixed equilibria, optimality gaps, and computational complexity

A major extension of performative stability replaces deterministic fixed points by mixed equilibria over model iterates. If xiXix_i\in X_i1 are produced by any online algorithm with sublinear regret on stochastic losses

xiXix_i\in X_i2

and xiXix_i\in X_i3 is the uniform distribution over xiXix_i\in X_i4, then

xiXix_i\in X_i5

This result requires no smoothness, strong convexity, or continuity assumptions on the distribution map xiXix_i\in X_i6. It also explains why stable single points are not the only relevant solution concept: a mixed performatively stable equilibrium can exist even when no deterministic stable model exists. The standard example is squared loss xiXix_i\in X_i7 with

xiXix_i\in X_i8

on xiXix_i\in X_i9, for which there is no performatively stable point and retraining oscillates as D(θ)D(\theta)00 (Farina et al., 27 Feb 2026).

Recent reinforcement-learning work makes the distinction between stability and optimality especially explicit. In performative reinforcement learning, a performatively stable policy is optimal for the environment induced by itself, but this is not the same as the performatively optimal policy that maximizes the true matched value D(θ)D(\theta)01. The paper "Achieve Performatively Optimal Policy for Performative Reinforcement Learning" (Chen et al., 6 Oct 2025) states that there is a provably positive constant gap between the performatively stable policy and the desired performatively optimal policy in prior work, and it emphasizes that PS is neither necessary nor sufficient for PO. Under the regularizer dominance condition, the value function satisfies a gradient dominance property; on the compact lower-bounded policy set

D(θ)D(\theta)02

a zeroth-order Frank-Wolfe method is used to return a D(θ)D(\theta)03-stationary policy with probability at least D(θ)D(\theta)04, and if D(θ)D(\theta)05, that policy is also an D(θ)D(\theta)06-PO policy. The stated complexities are

D(θ)D(\theta)07

This suggests that performative stability, although natural for repeated retraining, need not identify the desired objective in environments where the policy should be evaluated in its own induced dynamics (Chen et al., 6 Oct 2025).

The computational difficulty of finding performatively stable points is now understood to exhibit a sharp threshold. Under the Euclidean assumptions of D(θ)D(\theta)08-strong convexity, D(θ)D(\theta)09-smoothness, and distribution sensitivity

D(θ)D(\theta)10

the key parameter is

D(θ)D(\theta)11

When D(θ)D(\theta)12, repeated risk minimization is contractive and converges linearly. By contrast, computing an D(θ)D(\theta)13-performatively stable point is D(θ)D(\theta)14-hard even when

D(θ)D(\theta)15

and the problem is PPAD-complete. The hardness persists even for the quadratic loss

D(θ)D(\theta)16

and affine distribution shifts, extends to any well-bounded convex compact domain, and is accompanied by an exponential ERM-query lower bound. The same paper also proves that computing a strategic local optimum is D(θ)D(\theta)17-hard, and that performatively stable points in strategic classification with endogenous costs are D(θ)D(\theta)18-hard (Anagnostides et al., 28 Jan 2026).

6. Kyle’s equilibrium as a sharp local-stability threshold

A particularly sharp characterization of performative stability appears in the discrete-time insider-trading model of Kyle. There are D(θ)D(\theta)19 trading times, noise-trader orders satisfy D(θ)D(\theta)20, the terminal asset value satisfies D(θ)D(\theta)21, aggregate order flow is

D(θ)D(\theta)22

market makers set efficient prices

D(θ)D(\theta)23

and the insider maximizes

D(θ)D(\theta)24

The unique linear equilibrium is described by coefficients D(θ)D(\theta)25 such that

D(θ)D(\theta)26

The paper studies stability via the insider’s policy-iteration map

D(θ)D(\theta)27

and defines local stability by the existence of an D(θ)D(\theta)28 such that every initial D(θ)D(\theta)29 with

D(θ)D(\theta)30

satisfies D(θ)D(\theta)31 (Cetin et al., 2023).

The main theorem is exact: Kyle’s discrete-time equilibrium is performatively stable if and only if there are at most two trading dates. For D(θ)D(\theta)32, the map satisfies

D(θ)D(\theta)33

so the equilibrium is super-attractive. For D(θ)D(\theta)34, the Jacobian at equilibrium is

D(θ)D(\theta)35

and the equilibrium is locally stable. For every D(θ)D(\theta)36, and for all D(θ)D(\theta)37, the equilibrium is not locally stable. More precisely, the induced scalar map in the D(θ)D(\theta)38-th coordinate satisfies

D(θ)D(\theta)39

so perturbations expand in at least one direction. After normalization D(θ)D(\theta)40, the derivative becomes a universal constant for all D(θ)D(\theta)41, and for D(θ)D(\theta)42

D(θ)D(\theta)43

(Cetin et al., 2023).

This example is notable for two reasons. First, the stability classification is parameter-independent: it depends only on the number of trading dates, not on D(θ)D(\theta)44, D(θ)D(\theta)45, or D(θ)D(\theta)46. Second, the equilibrium is not unstable when D(θ)D(\theta)47: some perturbations, such as those in the last coordinate or the D(θ)D(\theta)48-th coordinate, can still converge back to equilibrium, even though generic perturbations in earlier coordinates do not. The result therefore isolates a dimension threshold in performative stability: one and two-dimensional policy-iteration maps are locally contracting, while three or more trading periods generate an expanding direction and destroy local attraction (Cetin et al., 2023).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Performatively Stable Equilibrium.