---
title: Coherent Utility Measure Games
url: https://www.emergentmind.com/topics/coherent-utility-measure-games
type: topic
---

# Coherent Utility Measure Games

Searching arXiv for recent and directly relevant papers on coherent utility measure games, coherent utility/risk measures in games, and coherence-based gamble frameworks.
Coherent utility measure games are strategic models in which each player evaluates random payoffs through a coherent utility functional rather than through ordinary expectation. In the formulation developed for data-driven distributionally robust games, a player’s payoff from a mixed profile is the coherent utility of the induced random payoff, which yields an equivalent worst-case expectation representation over an ambiguity set of probability measures [2605.19302]. Related work places these games within a broader coherence-based lineage: coherence between ex-ante and ex-post evaluations of psychological gambles characterizes subjective expected utility, minimax expected utility, and Choquet expected utility [2307.10328], while function-coherent gambles generalize de Finetti–Walley desirability to non-linear utility via utility-space convexity and a continuous linear representation [2503.01855].

## 1. Coherent utility as a payoff functional

In the reward convention used for coherent utility measure games, a functional $\rho:\mathcal{X}\to\overline{\mathbb{R}}$ on measurable real random variables is a coherent utility if it satisfies four axioms: concavity, monotonicity, translation equivariance, and positive homogeneity [2605.19302]. Concavity is
\[
\rho(\alpha X + (1-\alpha)Y) \ge \alpha \rho(X)+(1-\alpha)\rho(Y),
\]
monotonicity requires
\[
Y(\omega)\ge X(\omega)\ \forall\omega \Rightarrow \rho(Y)\ge \rho(X),
\]
translation equivariance requires
\[
\rho(X+a)=\rho(X)+a,
\]
and positive homogeneity requires
\[
\rho(tX)=t\rho(X).
\]

These axioms are the reward version of the Artzner–Delbaen–Eber–Heath axioms for coherent risk measures. For a loss variable $L$, a coherent risk measure $\varphi(L)$ corresponds to a coherent utility on reward $X=-L$ via $\rho(X)=-\varphi(-X)$ [2605.19302]. Under standard regularity, coherent utilities admit the dual representation
\[
\rho(X)=\inf_{\mathbb{Q}\in U}\mathbb{E}^{\mathbb{Q}}[X],
\]
for a nonempty set $U$ of probability measures. This representation is the formal bridge between risk-theoretic modeling and distributional robustness: utility is evaluated as a lower envelope of linear expectations.

The importance of this formulation is structural rather than merely notational. It makes risk sensitivity a primitive feature of player preferences while retaining a representation by linear expectations over a suitably chosen ambiguity set. In consequence, coherent utility measure games inherit both the geometry of coherent risk theory and the strategic interpretation of distributionally robust games [2605.19302].

## 2. Coherence-based foundations in gambles and acts

A deeper conceptual foundation comes from the theory of psychological gambles, where acts are embedded in a space of finitely supported nonnegative gambles on acts, and ex-ante utility $V$ is compared with an ex-post, statewise order induced by an outcome utility $u$ [2307.10328]. In that framework, for a gamble $\theta$ and state $\omega$,
\[
u(\theta|\omega)=\sum_{f\in\mathcal{F}}\theta(f)\,u(f(\omega)),
\]
and the statewise order is
\[
\theta\ge_u\eta \iff u(\theta|\omega)\ge u(\eta|\omega)\ \forall\omega\in\Omega.
\]
Full coherence requires
\[
\theta\ge_u\eta \Rightarrow V(\theta)\ge V(\eta),\quad \forall \theta,\eta\in\Theta.
\]

Under Assumption $\ref{structure}$, coherence is equivalent to a generalized expected-utility representation
\[
V(f)=\Phi(u(f))+\int_\Omega u(f)\,dm,
\]
where $m$ is a finitely additive probability and $\Phi$ is a positive linear functional vanishing on bounded functions; under a continuity assumption, $\Phi=0$ and the representation collapses to subjective expected utility
\[
V(f)=\int_\Omega u(f)\,dm
\]
[2307.10328]. Weaker coherence notions generate alternative models: $\Theta_0$-coherence yields maxmin expected utility
\[
V(f)=\inf_{\lambda\in\Lambda}\int_\Omega u(f)\,d\lambda,
\]
while $\Theta_1$-coherence yields Choquet expected utility
\[
V(f)=\int u(f)\,d\nu_V
\]
with $\nu_V$ a convex capacity [2307.10328].

The same paper explicitly connects coherence to absence of arbitrage and to coherent utility and risk measures. Full coherence yields linear expectation under a single prior; weaker coherence yields lower envelopes of expectations or Choquet integrals, both closely related to coherent utility measures [2307.10328]. In a game-theoretic environment, each player’s payoff vector over states can be treated as an act, and the player’s ex-ante evaluation can take SEU, MMEU, or CEU form. This places coherent utility measure games within a broader de Finetti–no-arbitrage interpretation of strategic evaluation.

## 3. From gamble coherence to non-linear utility coherence

Function-coherent gambles extend the desirable gambles framework to non-linear utility while preserving a Dutch-book–style coherence structure [2503.01855]. Let $X$ be a real vector space of gambles and let $u:X\to V$ be a strictly increasing, continuous map into a locally convex Hausdorff topological vector space, normalized by $u(0)=0$. Acceptance is defined by
\[
\mathbb{D}=\{f\in X:u(f)\ge 0\}.
\]
A set $\mathbb{D}\subseteq X$ is function-coherent with respect to $u$ if it satisfies F1 (Avoid partial losses), F2 (Monotonicity), and F3 ($u$-convexity). The $u$-convexity condition requires that for $f,g\in\mathbb{D}$ and $\lambda,\mu\ge 0$,
\[
h=u^{-1}\bigl(\lambda u(f)+\mu u(g)\bigr)
\]
is again acceptable, assuming it is well defined [2503.01855].

Under F1–F3, the utility-transformed acceptance set
\[
U(\mathbb{D})=\{u(f):f\in\mathbb{D}\}\subseteq V
\]
is a convex cone. With regularity conditions $R1$ and $R2$, there exists a continuous linear functional $\ell:V\to\mathbb{R}$, unique up to positive scaling, such that
\[
f\in\mathbb{D}\iff \ell(u(f))\ge 0.
\]
Equivalently, the composed functional
\[
\rho(f)=\ell(u(f))
\]
represents acceptance [2503.01855].

This generalization is directly relevant to coherent utility measure games because it separates preferences from beliefs. The map $u$ encodes non-linear evaluation of consequences, while $\ell$ aggregates utility values linearly across states [2503.01855]. This suggests a broad design space for strategic models in which players are risk-averse, discounted, or otherwise non-linear at the consequence level, yet still coherent in the sense of avoiding sure loss and respecting dominance.

## 4. Formal model of coherent utility measure games

A coherent utility measure game is a distributionally robust game in which each player’s payoff evaluation is given directly by a coherent utility functional of the player’s random payoff [2605.19302]. The model has $m$ players, finite pure action sets $\mathbf{A}_i$, mixed strategy simplices $\mathbf{X}_i$, a random state $\xi\in\Xi\subseteq\mathbb{R}^k$, and random payoffs
\[
u_i(\mathbf{x}_i,\mathbf{x}_{-i}\mid\xi)
=
\sum_{\mathbf{a}\in\mathbf{A}}u_i(\mathbf{a}\mid\xi)\prod_{s=1}^m x_s(a_{j_s}).
\]
Given a coherent utility $\rho_i$, player $i$’s payoff from $(\mathbf{x}_i,\mathbf{x}_{-i})$ is
\[
\rho_i(\mathbf{x}_i,\mathbf{x}_{-i})
:=
\rho_i\bigl(u_i(\mathbf{x}_i,\mathbf{x}_{-i}\mid\xi)\bigr).
\]
By duality, there is an associated ambiguity set $U_i$ such that
\[
\rho_i(\mathbf{x}_i,\mathbf{x}_{-i})
=
\inf_{\mathbb{Q}\in U_i}\mathbb{E}^{\mathbb{Q}}\bigl[u_i(\mathbf{x}_i,\mathbf{x}_{-i}\mid\xi)\bigr].
\]

The paper is explicitly data-driven. There are $K$ i.i.d. payoff samples $\xi_1,\dots,\xi_K$, the nominal distribution $\mathbb{P}$ is empirical with $\mathbb{P}(\xi_k)=1/K$, and empirical expected payoffs are
\[
\mu_i^\mathbb{P}(\mathbf{a})
=
\frac{1}{K}\sum_{k=1}^K u_i(\mathbf{a}\mid\xi_k).
\]
Ambiguity sets considered in the existence analysis include $f$-divergence balls, Wasserstein balls, and coherent utility–induced dual sets [2605.19302]. In the coherent utility specializations, the analysis often proceeds directly with closed-form risk-adjusted utilities on the empirical distribution.

The principal coherent utilities instantiated in this framework are summarized below.

| Utility | Formula | Parameter range |
|---|---|---|
| Mean–semideviation | $\rho_{\mathrm{MSD}}(X)=\mathbb{E}^\mathbb{P}[X]-\gamma_s\,\mathbb{E}^\mathbb{P}[\max(0,\mathbb{E}^\mathbb{P}[X]-X)]$ | $\gamma_s\in[0,1]$ |
| Mean–deviation | $\rho_{\mathrm{MD}}(X)=\mathbb{E}^\mathbb{P}[X]-\gamma_d\,\mathbb{E}^\mathbb{P}[|X-\mathbb{E}^\mathbb{P}[X]|]$ | $\gamma_d\in[0,1/2]$ |
| CVaR-based coherent utility | $\rho_{\mathrm{CVaR}}(X)=(1-\gamma_c)\mathbb{E}^\mathbb{P}[X]+\gamma_c\,\mathrm{CVaR}_\alpha[X]$ | $\gamma_c\in[0,1]$, $\alpha\in(0,1)$ |

For reward $X$, the reward-CVaR is
\[
\mathrm{CVaR}_\alpha[X]
=
\sup_{z\in\mathbb{R}}
\left\{
z+\frac{1}{\alpha}\mathbb{E}^\mathbb{P}[\min(0,X-z)]
\right\},
\]
that is, the expectation of the worst $100\alpha\%$ of outcomes [2605.19302].

## 5. Equilibrium structure and the continuous-game character

The equilibrium notion is the distributionally robust equilibrium, or risk-aware Nash equilibrium:
\[
\mathbf{x}_i^* \in \arg\max_{\mathbf{x}_i\in\mathbf{X}_i}\rho_i(\mathbf{x}_i,\mathbf{x}_{-i}^*),\quad \forall i
\]
[2605.19302]. Existence follows from a Kakutani fixed-point argument when the first moment of utilities is uniformly bounded on the ambiguity set:
\[
|\mu_i^\mathbb{Q}(\mathbf{a})|
=
|\mathbb{E}^\mathbb{Q}[u_i(\mathbf{a}\mid\xi)]|
\le M<\infty
\quad \forall i,\mathbf{a},\mathbb{Q}\in U.
\]
In the data-driven setting described in the paper, this boundedness condition is satisfied for several ambiguity classes, including $f$-divergence balls, Wasserstein balls under Lipschitz assumptions, and coherent utility dual sets [2605.19302].

A central structural result is that distributionally robust games, and therefore coherent utility measure games, are inherently continuous, rather than finite matrix games [2605.19302]. Although pure action sets are finite, mixed-strategy payoffs cannot in general be reduced to convex combinations of componentwise worst-case pure-action payoffs. Formally,
\[
\rho_i(\mathbf{x})
=
\inf_{\mathbb{Q}\in U}\sum_{\mathbf{a}}\mu_i^\mathbb{Q}(\mathbf{a})\prod_s x_s(a_{j_s})
\ge
\sum_{\mathbf{a}}\inf_{\mathbb{Q}\in U}\mu_i^\mathbb{Q}(\mathbf{a})\prod_s x_s(a_{j_s}),
\]
with strict inequality possible [2605.19302]. A common misconception is therefore that one can replace a coherent utility measure game by a finite normal-form game whose pure payoffs are already worst-case-adjusted; the paper shows that this reduction fails.

This continuous lifted structure also alters correlated equilibrium. Standard finite-game correlated equilibrium definitions rely on linear extension from pure to mixed actions, but that linear extension does not reproduce the actual $\rho_i(\mathbf{x})$ in coherent utility measure games [2605.19302]. The appropriate continuous-game definition uses measurable deviation maps $\phi_i:\mathbf{X}_i\to\mathbf{X}_i$ and requires
\[
\mathbb{E}^{\mathbb{C}}\big[\rho_i(\phi_i(\mathbf{x}_i),\mathbf{x}_{-i})-\rho_i(\mathbf{x}_i,\mathbf{x}_{-i})\big]\le 0.
\]
This point is not a technicality; it marks a change in equilibrium structure that precludes direct extensions of standard correlated equilibrium notions.

## 6. Complexity and complementarity formulations

Approximate equilibrium is defined by
\[
\rho_i(\mathbf{x}_i^*,\mathbf{x}_{-i}^*)
\ge
\max_{\mathbf{x}_i\in\mathbf{X}_i}\rho_i(\mathbf{x}_i,\mathbf{x}_{-i}^*)-\epsilon,
\quad \forall i.
\]
Under bounded payoffs and polynomial-time computable, Lipschitz utilities, computing an approximate distributionally robust equilibrium in data-driven distributionally robust games is PPAD-complete [2605.19302]. Hardness follows because any finite matrix game is a distributionally robust game with singleton ambiguity set, while membership follows from the fact that these are concave games with compact convex strategy sets and suitable separation oracles.

For mean–semideviation, mean–deviation, and CVaR coherent utility games with $\gamma>0$, approximate equilibrium computation belongs to PPAD [2605.19302]. The proofs establish polynomial-time computability of the utility functionals and Lipschitz continuity in mixed strategies. For MSD and MD, the penalties are computed by summations of max or absolute-deviation terms over finitely many samples. For CVaR, computation reduces to a small linear program in $K+1$ variables, and Danskin’s theorem is used to control Lipschitz constants [2605.19302].

The same paper derives multilinear complementarity program formulations for several coherent utility measure games. In the mean–semideviation case, auxiliary variables encode the negative of semideviation terms, and the player best-response problem becomes a linear program whose KKT conditions yield a mixed complementarity system [2605.19302]. In the CVaR case, auxiliary variables $\nu_{i,k}$ and a threshold $z_i$ encode the $\min(0,\cdot)$ terms, again producing a mixed complementarity system. A recurrent equilibrium implication is that all actions in the support of a mixed strategy have equal risk-adjusted value.

These complementarity systems can be normalized into multilinear complementarity problems without simplex equalities, providing a bridge to PATH and other complementarity solvers [2605.19302]. Even for two-player games, however, the resulting formulations do not reduce to a linear complementarity problem, so Lemke–Howson does not apply [2605.19302]. This sharply distinguishes coherent utility measure games from ordinary bimatrix games.

## 7. Strategic behavior, dynamic extensions, and open directions

The available examples show that coherent utility measure games can materially change equilibrium selection and out-of-sample behavior [2605.19302]. In a coordination game with mean–semideviation utility, increasing downside risk aversion enlarges the parameter region in which only the conservative mixed equilibrium remains. In a small-$K$ prisoner’s-dilemma variant, more conservative MSD equilibria improve average out-of-sample expected payoff in the reported experiments, although the relationship with the risk parameter is not monotone over the entire range because equilibrium structure changes [2605.19302]. In a CVaR game, the equilibrium exhibits a regime change at $\alpha=0.5$, and increasing $\gamma_c$ reduces both expected payoff and variance for the risk-sensitive player while preserving an explicit tail guarantee through the VaR threshold $z_i^*$ [2605.19302].

The performance analysis in the same framework quantifies the loss in expected utility from risk aversion. Under bounded payoffs and concentration assumptions, if the coherent utility has the form
\[
\rho_i(\mathbf{x})
=
\mathbb{E}^\mathbb{P}[u_i(\mathbf{x}\mid\xi)]
-\gamma\,R(\mathbb{P},u_i(\mathbf{x}\mid\xi)),
\]
then with probability at least $1-2\delta$,
\[
\mathbb{E}^\mathbb{T}[u_i(\mathbf{s}^*\mid\xi)]
\le
\mathbb{E}^\mathbb{T}\big[u_i(b_\rho(\mathbf{s}_{-i}^*),\mathbf{s}_{-i}^*\mid\xi)\big]
+
2\sqrt{\frac{\ln(4|\mathbf{A}|/\delta)}{2K}+\gamma B.
}
\]
The corresponding expressions for $B$ are given explicitly for MSD, MD, and CVaR utilities [2605.19302]. The bound suggests that one should scale $\gamma\sim 1/K$ to maintain vanishing regret as more data arrives.

Broader coherence-based work indicates two extension paths. First, the psychological-gamble framework shows that full coherence, $\Theta_1$-coherence, and $\Theta_0$-coherence correspond respectively to SEU, Choquet expected utility, and maxmin expected utility, which suggests a taxonomy of strategic models based on how restrictive the underlying coherence requirement is [2307.10328]. Second, function-coherent gambles show that exponential, hyperbolic, quasi-hyperbolic, generalized hyperbolic, scale-dependent, state-dependent, and hybrid discounting can all be embedded into a coherent acceptance framework of the form $\ell\circ u$ [2503.01855]. This suggests that dynamic or repeated coherent utility measure games can be formulated with non-linear and time-dependent utility while preserving frame-by-frame coherence.

Several open directions are identified explicitly. These include a full theory and computation of continuous correlated equilibria in coherent utility measure games, efficient algorithms beyond generic MLCP solvers, dynamic, repeated, and incomplete-information coherent utility measure games, and characterization of classes with monotone comparative statics in risk-aversion parameters [2605.19302]. In that sense, coherent utility measure games are both a concrete equilibrium model and a synthesis point for distributional robustness, coherent risk theory, non-additive utility, and coherence-based gamble foundations.

Source: https://www.emergentmind.com/topics/coherent-utility-measure-games