---
title: Norm-Augmented Markov Games
url: https://www.emergentmind.com/topics/norm-augmented-markov-games
type: topic
---

# Norm-Augmented Markov Games

A Norm-Augmented Markov Game (NMG) is a formal framework extending classical Markov Games by explicitly representing, learning, and enforcing normative systems—sets of prohibitive and obligative rules—alongside the conventional dynamics and reward structures. This unifies algorithmic approaches to multi-agent strategic interaction with mechanisms for the emergence, inference, and collective enforcement of social and institutional norms. In NMGs, agents form beliefs about which norms govern joint behavior, update these beliefs through Bayesian inference from observed actions, and plan using norm-penalized rewards, enabling the modeling and engineering of artificial societies with shared normative structures [2402.13399].

## 1. Formal Definition and Structure

An NMG is formally specified by
\[
\mathcal{M} = \bigl\langle \mathcal{I},\,\mathcal{S},\,\{\mathcal{A}_i\}_{i \in \mathcal{I}},\,T,\,\{R_i\}_{i \in \mathcal{I}},\,\gamma,\,\mathcal{N},\,\{\mathcal{P}^0_i\}_{i \in \mathcal{I}},\,\{C_i\}_{i \in \mathcal{I}}\bigr\rangle,
\]
where:
- $\mathcal{I}$ denotes the finite agent set.
- $\mathcal{S}$ is the finite state space.
- $\mathcal{A}_i$ are finite action sets; $\mathcal{A} = \prod_i \mathcal{A}_i$ is the joint action space.
- $T: \mathcal{S} \times \mathcal{A} \to \Delta(\mathcal{S})$ is the transition kernel.
- $R_i: \mathcal{S} \times \mathcal{A}_i \times \mathcal{S} \to \mathbb{R}$ denotes material rewards.
- $\gamma \in [0,1)$ is the discount factor.
- $\mathcal{N}$ is the finite space of norms $\nu$ (each $\nu: \mathcal{H} \to \{0,1\}$ mapping histories to violation flags).
- $\mathcal{P}^0_i$ is agent $i$'s prior over sets of norms in force.
- $C_i: \mathcal{N} \to \mathbb{R}_{\geq 0}$ assigns cost to norm violations.

A normative system $N \subseteq \mathcal{N}$ is stable if, given common belief in $N$, rational agents find it optimal to comply with all $\nu \in N$ with high frequency [2402.13399].

## 2. Norm Types and Formalization

Norms in NMGs adopt first-order, rule-based templates and fall into two primary categories:

- **Prohibitive norms (maintenance):** $\nu_P = \langle \mathsf{Prohib}(x), \mathsf{Post}(x) \rangle$, where $\mathsf{Prohib}(x)$ denotes a set of forbidden actions and $\mathsf{Post}(x)$ is the predicate over successor states. Violation: $\nu_P(h,s,a,s') = 1$ if $a \in s[\mathsf{Prohib}]$ and $s'[\mathsf{Post}] = 1$.
- **Obligative norms (achievement):** $\nu_O = \langle \mathsf{Pre}, \mathsf{Post},\tau \rangle$ with trigger $\mathsf{Pre}$, discharge condition $\mathsf{Post}$, and deadline $\tau$. Violation occurs if, for a $k \le \tau$, $\mathsf{Pre}$ is satisfied at time $t-k$ and $\mathsf{Post}$ is not achieved by $t$.

Typical examples include resource conservation restrictions (“do not pick apples if local density is low”) and role-conditioned obligations (“if a cleaner observes river dirtiness $>0.3$, must clean within 20 steps”) [2402.13399].

## 3. Learning Norms via Bayesian Rule Induction

Agents in NMGs posit a latent set of governing norms $N \subseteq \mathcal{N}$ and maintain a posterior
\[
P_i^t(N) \propto P_i^0(N) P_i(h_t \mid N),
\]
where $P_i^0(N)$ is a factored prior, and $P_i(h_t \mid N)$ quantifies the likelihood of observed joint behavior given universal compliance with $N$. Due to the exponential scale in $|\mathcal{N}|$, a mean-field approximation is adopted: Each agent $i$ maintains marginals $\tilde{P}_i^t(\nu)$ for candidate norms, updated as
\[
\tilde{P}_i^{t+1}(\nu) \propto \pi_{-i}(a_{-i,t} \mid h_t, s_t, \nu)\, \tilde{P}_i^t(\nu),
\]
where $\pi_{-i}(a_{-i} \mid h,s,\nu)$ is the policy-likelihood of observed non-$i$ actions assuming $\nu$ is enforced. In practice, this policy-likelihood uses the $Q^\nu$ action-value from model-based planning with norm-augmented reward [2402.13399].

## 4. Norm Integration into Planning and Control

At each timestep, agent $i$ forms beliefs $\tilde{P}_i^t(\nu)$ and plans with a norm-augmented instantaneous reward,
\[
R'_i(h',s,a,s') = R_i(s,a,s') - \sum_{\nu \in \mathcal{N}} \tilde{P}_i^t(\nu) C_i(\nu) \nu(h',s,a,s').
\]
For prohibitions, standard Real-Time Dynamic Programming (RTDP) is used. Obligations, when triggered, cause a switch to obligation-oriented planning with a modified reward penalizing lapsed obligations and rewarding successful discharge before deadline [2402.13399]. At execution, norm compliance can be implemented by deterministic thresholding on belief, or by sampled norm-sets (Thompson sampling) for implicit coordination.

## 5. Empirical Results and Evaluative Metrics

Experimental realization of NMGs was conducted in gridworld domains combining resource management, prosocial labor (cleaning), and role-based payments [2402.13399]. Agents, with roles (e.g., Farmer, Cleaner), operated in environments with 68 enumerated candidate norms.

Key results include:
- **Passive norm learning:** Single agents can infer ground-truth active norms within 150–300 timesteps, achieving $\sim$80% recall/precision when using a 0.95 posterior threshold.
- **Norm-enabled welfare:** Activation of shared norm-sets leads to increased social welfare and sustainability (orchard health), while absence leads to welfare stagnation and >90% environmental degradation.
- **Intergenerational transmission:** Beliefs over prohibitive norms are reliably transmitted for agent lifespans $L \ge 50$; obligative norms require longer, $L \ge 300$.
- **Norm emergence/convergence:** Multiple naive agents converge beliefs on salient norms (posterior $\ge$95%) within 300 steps, typically achieving unanimous consensus.

Metrics included posterior belief accuracy (precision/recall), mean team reward, sustainability ratios (e.g., fraction of desiccated orchards), intergenerational stability, and consensus convergence rates [2402.13399].

## 6. Connections to Norm-Augmented Optimization in Markov Games

An orthogonal approach termed “norm-augmentation” appears in the equilibrium computation literature for general-sum MGs. Here, “norm” refers not to rules governing behavior, but to explicit $L_p$ constraints on Bellman-like residuals [1606.08718].

For a $\gamma$-discounted $N$-player general-sum MG, an $\epsilon$-Nash equilibrium is defined by bounding the $L_p$ distance between player values under the joint strategy and their best response values:
\[
Err(\pi) = \left( \sum_{i=1}^N \rho(i) \left( \sum_{s \in S} \mu(s) |v^{*i}_{\pi^{-i}}(s) - v^i_\pi(s)|^p \right)^{1/p} \right)^{1/p} \leq \epsilon.
\]
The optimization minimizes the $L_p$ norm of two Bellman residuals per player, resulting in theoretical guarantees on equilibrium quality. Practically, this is implemented in the NashNetwork, which jointly trains value and policy networks to minimize Lp-aggregated residuals over batch data [1606.08718].

## 7. Limitations and Future Directions

NMGs, as formulated, currently assume:
- **Batch setting** and access to a fixed transition kernel, limiting stochastic applicability unless remedied with model-estimation in RKHS or double-sampling.
- **Non-convex objectives:** Neural network optimization is not globally guaranteed; only stationary points guarantee weak-$\epsilon$ Nash bounds in the case of norm-augmented Bellman residuals.
- **Single equilibrium discovery:** The NashNetwork approach finds one equilibrium, despite multiplicity [1606.08718].

Prospective directions include scaling to richer state representations, exploring optimization techniques tailored to Bellman-residual objectives, incorporating explicit kernel-based model estimation to handle stochastic transitions, and extending normative inference to more complex relational and simultaneous-move environments [2402.13399, 1606.08718]. 

The NMG formalism provides a principled substrate for the emergence, inference, and enforcement of social norms in artificial societies, while norm-augmented optimization offers robust methods for learning equilibria in complex multi-agent settings.

Source: https://www.emergentmind.com/topics/norm-augmented-markov-games