---
title: Private Markovian Equilibrium (PME)
url: https://www.emergentmind.com/topics/private-markovian-equilibrium-pme
type: topic
---

# Private Markovian Equilibrium (PME)

Searching arXiv for recent papers on "Private Markovian Equilibrium" and related mean-field / Markov game terminology.
Private Markovian Equilibrium (PME) denotes an equilibrium concept for dynamic games in which each agent’s strategy is restricted to be Markovian and to depend on information that is private to that agent, together with whatever public state variables the model treats as commonly observed. In the nonatomic-game formulation of Adlakha, Johari, Weintraub, and Goldsmith, a PME is a private Markov policy \(\phi\) that is optimal against the deterministic environment path generated when all players use \(\phi\) [1510.06813]. In the non-stationary mean-field-game formulation of Vasal, PME is a symmetric Markov perfect equilibrium in which the prescription \(\sigma_t(a\mid \theta,z)\) depends on a player’s private type \(\theta\) and the current population state \(z\), with equilibrium characterized by a Bellman-type fixed point and a mean-field consistency condition [1905.04154]. In the Stackelberg Markov-game formulation for smart-grid demand response, PME refers to a lower-level equilibrium of users whose private states are storage levels and whose strategies are private Markovian strategies (PMS) [2509.12225]. These formulations share the same structural idea—Markovian decision rules conditioned on private state information—but differ materially in what counts as the relevant public environment and in how equilibrium consistency is enforced.

## 1. Conceptual scope and terminology

The defining feature of PME is the restriction to **private Markovian** strategies. In the nonatomic-game model, a private (mixed) Markov policy is a measurable kernel
\[
\phi:S\longrightarrow\mathcal P(A),
\]
or equivalently a measurable function \(x:S\times G\to A\), where the player uses only her own current state \(s_t\) and an idiosyncratic pre-action shock \(g_t\) to randomize over actions [1510.06813]. The resulting equilibrium is “private” because the policy depends only on the player’s own state and private randomization, not on the realized states or actions of others.

In Vasal’s mean-field model, the terminology is slightly different. The game contains a large population of homogeneous players, each with a private type \(x_t^i\in X\), and a commonly observed aggregate state \(z_t\in Z=\Delta(X)\) equal to the empirical distribution of private types. A Markovian strategy profile is a sequence of prescription functions \(\sigma_t(\cdot\mid \theta,z)=\gamma_t(a\mid \theta)\), so the equilibrium depends on both a player’s private type and the current dynamic population state [1905.04154]. Here, privacy does not mean that the strategy ignores the aggregate environment; rather, it means that the individual’s private type is a primitive argument of the prescription.

In the smart-grid Stackelberg Markov game, the lower-level users have private states \(\boldsymbol s_i^t=(e_N^t,b_i^t)\), where \(b_i^t\) is user \(i\)’s storage level, and a PMS is a profile \(\boldsymbol\pi_i=(\pi_i^1,\dots,\pi_i^T)\) with
\[
\pi_i^t:\mathcal S_i^t\longrightarrow\Delta(\mathcal A_i^t).
\]
A PME is then a joint profile \(\boldsymbol\pi^*\) such that no user can improve by deviating to another PMS [2509.12225]. This suggests that PME is best understood as a family of equilibrium concepts rather than a single canonical definition: what remains invariant is the Markovian restriction together with conditioning on agent-specific private state information.

## 2. Formal definitions in the main model classes

In nonatomic games, the environment at date \(t\) is a joint distribution \(\mu_t\in\mathcal P(S\times A)\) of other players’ states and actions, the one-period payoff is
\[
u:S\times A\times\mathcal P(S\times A)\longrightarrow\mathbb R,
\]
and state evolution is governed by
\[
s_{t+1}\sim P(\cdot\mid s_t,a_t,\mu_t,\xi_t).
\]
When all players use the same \(\phi\), the law of \((s_t,a_t)\) is deterministic and evolves as
\[
\mu_{t+1}=\mathcal T(\mu_t,\phi).
\]
Given an initial state distribution \(\mu_0\), the pair \(\bigl(\phi,\{\mu_t\}\bigr)\) is a PME if
\[
J(\phi;\mu_0)\ge J(\psi;\mu_0)\quad \forall \psi,
\]
where \(J\) is the discounted expected payoff of a single agent facing the deterministic environment induced by \(\phi\) [1510.06813].

In the non-stationary mean-field-game model, a PME is a collection \(\{\sigma_t,V_t,z_t\}_{t=1,\dots,T}\) satisfying three coupled conditions. First, the value functions \(V_t:Z\times X\to\mathbb R\) satisfy the Bellman optimality equation
\[
V_t(z,\theta)=\max_{\gamma(\cdot\mid \theta)}
\sum_{a\in A}\gamma(a\mid \theta)\Bigl[R(\theta,a,z)+\delta\sum_{\theta'}Q_x(\theta'\mid \theta,a,z)V_{t+1}(\phi(z,\gamma),\theta')\Bigr].
\]
Second, once the maximizer is selected, one sets
\[
\sigma_t(a\mid \theta,z):=\gamma_t(a\mid \theta).
\]
Third, the mean field evolves consistently according to
\[
z_{t+1}=\phi(z_t,\gamma_t).
\]
For finite horizon the boundary condition is \(V_{T+1}\equiv 0\); for infinite horizon, the definition becomes a time-stationary fixed point [1905.04154].

In the Stackelberg Markov-game setting, the lower-level user value under PMS is
\[
V_{t,i}^{\boldsymbol\pi}(\boldsymbol s)=
\E\Bigl[\sum_{t'=t}^T r_i(\boldsymbol s^{t'},\boldsymbol a^{t'}) \Bigm|\boldsymbol s^t=\boldsymbol s\Bigr],
\]
and \(\boldsymbol\pi^*\) is a PME if, for every \(i\), every \(\boldsymbol s^1\in\mathcal S^1\), and every alternative PMS \(\pi_i\),
\[
V_{1,i}^{\pi_i^*,\pi_{-i}^*}(\boldsymbol s^1)\ge
V_{1,i}^{\pi_i,\pi_{-i}^*}(\boldsymbol s^1).
\]
The paper also gives an equivalent Bellman-type best-response characterization against fixed \(\pi^*_{-i}\) [2509.12225].

## 3. State variables, information structures, and equilibrium consistency

A central distinction across PME formulations is the object that summarizes the strategic environment. In Vasal’s model, the public environment is the mean-field state \(z_t\in\Delta(X)\), the empirical distribution of all players’ private types. If all players use the same prescription \(\gamma_t:X\to\Delta(A)\), then the mean field updates by
\[
z_{t+1}(y)=\sum_{\theta\in X}\sum_{a\in A}z_t(\theta)\gamma_t(a\mid \theta)Q_x(y\mid \theta,a,z_t),
\]
or equivalently \(z_{t+1}=\phi(z_t,\gamma_t)\) [1905.04154]. Consistency is therefore endogenous: the same prescription that is optimal at the individual level also determines the evolution of the aggregate state.

In the nonatomic-game model, the analogous object is \(\mu_t\), the deterministic law of states and actions under the common policy \(\phi\). Individual agents do not condition on the realized profile of others’ states or actions; instead, they optimize against the deterministic sequence \(\{\mu_t\}\) implied by mass behavior. The paper emphasizes that such equilibria “pay no attention to players’ external environments” and are therefore adoptable in settings where awareness of others’ states can be anywhere between full and non-existent [1510.06813].

In the smart-grid model, the public signal is renewable generation \(e_N^t\), while the private signal is the user’s storage level \(b_i^t\). The leader chooses real-time-pricing parameters \((\alpha,\beta,\gamma_1,\gamma_2)\in A^0\subset\mathbb R_+^4\), which define the price
\[
P(d^t,e_N^t)=\frac{\alpha}{n\,e_N^t+\gamma_1\,d^t}+\frac{\beta}{e_N^t+\gamma_2},
\qquad d^t=\sum_i d_i^t.
\]
Users choose \((d_i^t,c_i^t)\) subject to
\[
b_i^t+d_i^t-b_i^{\max}\le c_i^t\le b_i^t+d_i^t,
\]
and their stage reward is
\[
r_i(\boldsymbol s^t,\boldsymbol a^t)=\theta_i c_i^t-P(d^t,e_N^t)d_i^t.
\]
Here the consistency requirement is not a mean-field fixed point in the sense of [1905.04154]; instead, it is a lower-level Markov-game equilibrium among users under the leader’s announced pricing rule [2509.12225].

A common misconception is to treat “private” as implying complete ignorance of aggregate variables. The cited models show otherwise. In [1510.06813], privacy means conditioning only on own state and private randomization. In [1905.04154] and [2509.12225], private strategies may still depend on public state variables such as \(z_t\) or \(e_N^t\). What is excluded is conditioning on the full profile of other players’ private states.

## 4. Existence and equilibrium structure

Existence results for PME are established under different structural assumptions in each model class. In the nonatomic-game setting, assumptions (A1)–(A4) require compact metric spaces \(S\) and \(A\), continuity of \(u(s,a,\mu)\), joint continuity of the transition kernel \(P(\cdot\mid s,a,\mu,\xi)\), and fixed atomless shock laws \(\gamma\) and \(\lambda\). Under these conditions and \(\beta\in(0,1)\), there exists at least one private Markov policy \(\phi^*\) and environment path \(\{\mu_t^*\}\) such that \(\phi^*\) is a PME [1510.06813].

In Vasal’s non-stationary mean-field game, Assumption A1 requires \(R(\theta,a,z)\) and \(Q_x(\theta' \mid \theta,a,z)\) to be continuous in \(z\), with \(R\) bounded. Under A1, the paper states that there exists at least one solution \((\sigma_t,V_t,z_t)\) to the backward-recursive system (2.1–2.3) for finite and infinite horizon [1905.04154]. The proof sketch proceeds by showing that the stagewise maximization yields a nonempty, upper-hemicontinuous best-response correspondence, that the mean-field update \(\phi\) is continuous in \((z,\gamma)\), and that Kakutani’s fixed-point theorem yields existence of \(\gamma_t\), after which backward induction constructs \(V_t\). The paper further states that under further monotonicity assumptions on \(R\) in \(z\) one can also prove uniqueness of the PME [1905.04154].

In the smart-grid Stackelberg Markov game, Theorem 1 states that in the \(T\)-stage Markov game there exists at least one pure PME \(\boldsymbol\pi^*\) of the user subgame [2509.12225]. The proof sketch uses backward induction. At stage \(T\), the single-stage game parameterized by \(e_N^T,b_i^T\) is an exact potential game and admits a pure equilibrium in demands \(d_i^{*T}(e_N^T)\). By eliminating dominated \(c_i^T\) actions, one shows \(c_i^T=d_i^T+b_i^T\). Folding the continuation value into stage \(T-1\) preserves the potential structure, and recursion to \(t=1\) yields a sequence of pure policies \(\mathbbm1_{a_i^{*t}(\cdot)}\) that forms a pure PME [2509.12225].

These results indicate that PME is not tied to a single proof technique. Dynamic programming, contraction arguments, fixed-point theorems, and potential-game structure all appear, depending on whether the model is nonatomic, mean-field, or finite-horizon Markovian with private states.

## 5. Computation and algorithms

The computational treatment of PME differs sharply across the cited papers.

Vasal develops a **backward-recursive algorithm** for finite horizon \(T\), with \(V_{T+1}(z,\theta)\leftarrow 0\) and, for each \(t=T,\dots,1\) and each mean field \(z\in Z\), an inner fixed-point problem for the prescription \(\gamma_t:X\to\Delta(A)\). The prescription must satisfy, for each \(\theta\in X\),
\[
\gamma_t(\cdot\mid \theta)\in \arg\max_{\gamma(\cdot\mid \theta)}
\sum_a \gamma(a\mid \theta)\Bigl[R(\theta,a,z)+\delta\sum_{\theta'}Q_x(\theta'\mid \theta,a,z)V_{t+1}(\phi(z,\gamma_t),\theta')\Bigr],
\]
where \(\gamma_t\) appears on both sides through \(\phi(z,\gamma_t)\). The algorithm then sets \(\sigma_t(a\mid \theta,z)\leftarrow \gamma_t(a\mid \theta)\), updates the mean-field map \(\phi_t(z)=\phi(z,\gamma_t)\), and computes
\[
V_t(z,\theta)\leftarrow
\sum_a \gamma_t(a\mid \theta)\Bigl[R(\theta,a,z)+\delta\sum_{\theta'}Q_x(\theta'\mid \theta,a,z)V_{t+1}(\phi_t(z),\theta')\Bigr].
\]
For the infinite-horizon case, the paper states that one solves the stationary equation by value iteration. It also states that, at each \(t\) and each \(z\), one must solve a fixed point over the compact set \(\Delta(A)^{|X|}\), typically by iterating best responses or by Kakutani’s theorem; if \(\delta<1\) and \(R,Q_x\) are continuous in \(z\), the mapping is a contraction in the infinite-horizon case, so value iteration converges; and the overall complexity is \(O(T\cdot |Z|\cdot |X|\cdot |A|\cdot \mathrm{iter}_0)\) [1905.04154].

In the nonatomic-game treatment, computation in stationary settings proceeds by initializing a guess \(\mu^{(0)}\), solving the single-agent Bellman equation
\[
V^{(k)}(s)=\max_{a\in A}\Bigl\{u(s,a,\mu^{(k)})+\beta\int V^{(k)}(s')\,P(ds'\mid s,a,\mu^{(k)})\Bigr\},
\]
extracting the optimal policy \(\phi^{(k)}\), updating \(\mu^{(k+1)}=\mathcal T(\mu^{(k)},\phi^{(k)})\), and repeating until \(\|\mu^{(k+1)}-\mu^{(k)}\|\) is small. Under compactness and continuity, the iteration converges to \((\phi^*,\mu^*)\) [1510.06813].

The smart-grid paper provides both a **centralized** and a **decentralized** computational route. The key theoretical step is the construction of an auxiliary Markov potential game \(\mathcal G_2\), where each user’s state is only \(e_N^t\), action is \(d_i^t\), and stage payoff is
\[
g_i(e_N^t,\boldsymbol d^t)=\theta_i d_i^t-P\Bigl(\sum_j d_j^t,e_N^t\Bigr)d_i^t.
\]
The game has explicit potential
\[
\phi(e_N^t,\boldsymbol d^t)=\sum_i\Bigl[\theta_i-\frac{\beta}{e_N^t+\gamma_2}\Bigr]d_i^t
-\frac{\alpha}{n e_N^t+\gamma_1}\sum_i(d_i^t)^2
-\frac{\alpha}{n e_N^t+\gamma_1}\sum_{i<j} d_i^t d_j^t.
\]
Any pure ME of \(\mathcal G_2\) can be found by maximizing this potential via an FIP best-response process. The paper states that the centralized FIP converges in polynomial\((n,|\mathcal E_N^t|,|\mathcal D_i|)\) steps and reconstructs the PME in the original game by
\[
c_i^t(\boldsymbol s_i^t)=d_i^t(e_N^t)+b_i^t.
\]
It also presents a decentralized FP + MDP procedure in which each user maintains an estimate \(\hat\pi_{-i}\) of opponents’ aggregate-demand law, solves a personal MDP, observes \((t,e_N^t,d_{-i}^t)\), and updates by fictitious-play averaging; under mild conditions this process empirically converges to PME [2509.12225].

## 6. Applications, finite-player interpretation, and empirical illustrations

The applications in the cited papers make clear that PME is designed for settings with dynamic private states and strategic externalities.

Vasal studies a cyber-physical security problem in which a node’s private state is binary, \(X=\{0\text{ (“healthy”)},1\text{ (“infected”)}\}\), actions are \(A=\{0\text{ (“do nothing”)},1\text{ (“repair/harden”)}\}\), and infected nodes impose a negative externality through the infected fraction \(z(1)\). The state transition is
\[
x_{t+1}=
\begin{cases}
x_t+(1-x_t)\cdot w_t & \text{if } a_t=0,\\
0 & \text{if } a_t=1,
\end{cases}
\qquad w_t\sim\mathrm{Bernoulli}(q),
\]
and the one-stage payoff is
\[
R(\theta,a,z)=-(k+z(1))\cdot \theta-\lambda\cdot a.
\]
For \(q=0.9\), \(k=0.2\), \(\lambda=0.5\), and \(\delta=0.9\), the paper numerically computes the infinite-horizon equilibrium and reports three qualitative findings: healthy nodes adopt a threshold policy in which \(\gamma(1\mid 0)\) is a non-decreasing function of the infected fraction \(z(1)\); infected nodes almost always choose repair, \(\gamma(1\mid 1)\approx 1\), independent of \(z\); and the equilibrium value \(V(z,\theta)\) decreases in \(z\) for both \(\theta=0,1\) [1905.04154].

The smart-grid application models private battery levels in demand response. The numerical setup uses \(T=7\), \(n=50\), \(\mathcal D_i=\{0,\dots,4\}\), real solar data for \(\{\hat e_N^t\}\), discretized \(\omega_N^t\in\{-20,0,20\}\), transition \(\hat q\) fitted from data, \(\theta_i\sim U[0.9,1.5]\), and price parameters \(\alpha=19,\beta=20,\gamma_1=\gamma_2=1\). The paper states that centralized FIP finds PME in under 100 iterations, that users’ equilibrium demands increase in \(\theta_i\) and respond sensibly to \(e_N^t\), and that the aggregator can search \((\alpha,\beta)\in\{19,20,21\}^2\) while fixing \(\gamma_1,\gamma_2\) to maximize \(U(\alpha,\beta)\) [2509.12225].

The nonatomic-game paper supplies a different form of application-oriented significance: it proves that equilibria derived for nonatomic games can be used by large finite counterparts to achieve near-equilibrium performances. In the finite \(N\)-player version, if every player uses the oblivious PME \(\phi^*\), then there exists a sequence \(\varepsilon_N\downarrow 0\), explicitly \(\varepsilon_N=O(N^{-1/2})\) under Lipschitz or bounded-difference conditions, such that the profile “all use \(\phi^*\)” is an \(\varepsilon_N\)-Nash equilibrium of the \(N\)-player game [1510.06813]. This connects PME to approximation theory for large games: the nonatomic equilibrium provides a decentralized policy rule that remains nearly optimal in large but finite populations.

Taken together, these results show that PME serves at least three roles in current research. It is an equilibrium notion for nonatomic dynamic competition [1510.06813], a solution concept for non-stationary mean-field games with private types and dynamic population states [1905.04154], and a tractable equilibrium restriction for lower-level strategic interaction with private storage states in Stackelberg smart-grid control [2509.12225]. A plausible implication is that PME is most useful when the modeler seeks a compromise between strategic richness and computational or informational tractability: private state dependence is retained, while conditioning on the full joint state of all players is deliberately excluded or aggregated.

Source: https://www.emergentmind.com/topics/private-markovian-equilibrium-pme