---
title: Ergodic-Risk Constraints in Control
url: https://www.emergentmind.com/topics/ergodic-risk-constraints
type: topic
---

# Ergodic-Risk Constraints in Control

Ergodic-risk constraints are long-run control constraints in which admissibility is specified by an asymptotic risk quantity rather than by an ordinary average-cost bound. In the classical ergodic risk-sensitive control literature, the relevant quantity is a logarithmic exponential growth rate,
$$
J_x(c,\zeta) = \limsup_{T\to\infty}\frac{1}{\gamma T} \log \mathbb{E}_x^\zeta\!\left[\exp\!\left(\int_0^T \gamma\, c(X_t,\zeta_t)\,dt\right)\right],
$$
or its discrete-time analogue, so that both the objective and the constraint can be posed as long-run log-moment functionals. More recent work also uses the term for constraints on the asymptotic fluctuation level of a risk functional, typically through a functional central limit theorem and an asymptotic conditional variance. In both cases, the central theme is that the constraint regulates long-run variability, tail behavior, or cumulative uncertainty, rather than only mean performance [2301.00224] [2409.10767] [2503.05878].

## 1. Classical ergodic risk-sensitive formulation

Ergodic risk-sensitive control is an infinite-horizon stochastic control framework in which performance is measured not by an ordinary average cost, but by an exponential-of-integral criterion. In the ergodic setting, the horizon \(T\to\infty\), and the value is normalized by \(T\), leading to a long-run growth rate. For discrete-time controlled Markov chains, the corresponding criterion is
$$
J_x(c,\zeta) = \limsup_{T\to\infty}\frac{1}{\gamma T} \log \mathbb{E}_x^\zeta\!\left[\exp\!\left(\sum_{t=0}^{T-1}\gamma\, c(X_t,\zeta_t)\right)\right],
$$
while for controlled diffusions the same form appears with the time integral. The sign of \(\gamma\) encodes risk attitude: \(\gamma>0\) is risk-averse, \(\gamma=0\) is the risk-neutral limit recovering classical ergodic control, and \(\gamma<0\) is risk-seeking [2301.00224].

A major point emphasized in the literature is that the exponential criterion accounts for fluctuations around the mean, unlike a plain average-cost criterion. Two motivations recur. First, a linear average cost ignores fluctuations, whereas exponential criteria weight higher moments and therefore penalize variability and tail behavior. Second, the exponential criterion is multiplicative and therefore supports a multiplicative dynamic programming principle, whereas quadratic mean-variance type objectives often fail to satisfy dynamic programming because they are not time-consistent. This is the structural background from which ergodic-risk constraints emerge [2301.00224].

The standard additive ergodic criterion,
$$
\limsup_{T\to\infty}\frac{1}{T}\mathbb{E}\!\left[\int_0^T c(X_t,\zeta_t)\,dt\right],
$$
is therefore replaced by
$$
\limsup_{T\to\infty}\frac{1}{\gamma T} \log \mathbb{E}\!\left[e^{\gamma \int_0^T c\,dt}\right].
$$
This changes the optimization from additive to multiplicative, converts the Bellman equation into a nonlinear eigenvalue problem, and naturally penalizes variability and rare bad excursions. In the limit \(\gamma\to 0\), the risk-sensitive criterion reduces to the classical ergodic cost, at least formally [2301.00224].

## 2. Exponential-growth-rate constraints

The survey literature formulates an explicit constrained problem in which the control minimizes one ergodic risk-sensitive criterion while satisfying another:
$$
\text{minimize } J_x(c,\zeta) \quad\text{subject to}\quad \limsup_{T\to\infty}\frac{1}{T} \log \mathbb{E}\!\left[ e^{\sum_{t=0}^{T-1} k(X_t,\zeta_t)} \right] \le C.
$$
Here the second cost \(k\) is constrained through its exponential growth rate. This is an ergodic risk constraint: the constraint is not on the mean of \(k\), but on its exponential growth rate. The survey states that this is stronger and more tail-sensitive than an average-cost constraint [2301.00224].

The constrained problem is converted into an unconstrained Lagrangian form,
$$
\min_v \left\{ \max_{q}\widehat\Phi(q,v,c) + \Gamma\Big(\max_{q'}\widehat\Phi(q',v,k)-C\Big) \right\}, \qquad \Gamma\ge 0,
$$
where \(\Gamma\) is the Lagrange multiplier and \(\widehat\Phi\) is the ergodic risk-sensitive game value induced by the Kullback–Leibler variational form. The significance of this transformation is explicit: it turns an ergodic-risk constraint into a standard unconstrained risk-sensitive game or eigenvalue problem. The same treatment yields a corresponding primal linear program with variables \(\beta_i,\beta_i',V_i,V_i'\), and a dual program over occupation measures, embedding the constrained problem into a convex optimization framework [2301.00224].

This constrained formulation is closely tied to the variational representation of exponential moments. In the discrete-time variational formulas, terms of the form
$$
c(i,u)-D_{\mathrm{KL}(q(\cdot|i)\|P(\cdot|i,u))
$$
appear, where the Kullback–Leibler divergence measures the cost of deviating from the nominal transition kernel. The same large-deviation structure underlies the general formula
$$
\log\lambda(\mathbb A) = \sup_{(\pi,\tilde P)} \sum_i \pi(i)\left[\kappa_i - D_{\mathrm{KL}(\tilde p(\cdot|i)\|p(\cdot|i))\right].
$$
Accordingly, ergodic-risk constraints in this sense are inseparable from entropy-penalized robustness: the constraint is imposed on a long-run cumulant generating function, and the resulting optimization admits a robust game interpretation [2301.00224].

## 3. Fluctuation-based ergodic-risk constraints

A second, newer construction defines ergodic risk through the long-run cumulative fluctuation of a measurable risk functional \(g(X_t,U_t)\). The one-step uncertain part relative to past information is
$$
C_t \coloneqq g(X_t,U_t)-[g(X_t,U_t)\mid \mathcal F_{t-1}], \qquad t\ge 0,
$$
or, equivalently,
$$
C_t = g(X_t,U_t)-[g(X_t,U_t)\mid F_{t-1}].
$$
The corresponding ergodic-risk quantity is specified through the normalized cumulative uncertainty
$$
\frac{1}{\sqrt{t} \sum_{s=1}^t C_s \xrightarrow{d} C_\infty,
$$
together with the asymptotic conditional variance
$$
\frac{1}{t}\sum_{s=1}^t [C_s^2\mid \mathcal F_{s-1}] \xrightarrow{a.s.} \gamma_N^2.
$$
This criterion is not about the mean cost, but about the asymptotic distribution of accumulated fluctuations. It is therefore a long-horizon risk measure aimed at long-term cumulative uncertainty and extreme deviations beyond mean performance [2409.10767] [2503.05878].

In the linear quadratic case, the constrained optimization problem takes the form
$$
\min_{K\in \mathcal S} \; J(K) \quad \text{s.t.}\quad \gamma_N^2(K)\le \bar\beta,
$$
where \(\mathcal S=\{K : A_K \text{ is Schur stable}\}\) and
$$
J(K)=\limsup_{T\to\infty}\mathbb E\!\left[\frac{1}{T}\sum_{t=0}^T X_t^\top QX_t + U_t^\top RU_t\right].
$$
For stabilizing linear feedback \(U_t=KX_t\), the state covariance satisfies
$$
\Sigma_K = A_K \Sigma_K A_K^\top + H\Sigma_W H^\top.
$$
With quadratic risk functionals
$$
g(x,u)=x^\top Q^c x + u^\top R^c u,
$$
the asymptotic conditional variance becomes the risk budget variable. The constraint therefore does not alter the average-cost objective directly; it imposes a robustness specification on the long-run fluctuation level of the risk process [2503.05878].

The explicit quadratic formula is central:
$$
\gamma_N^2(K) = 4\,Q_K^c H\Sigma_W H^\top Q_K^c(\Sigma_K - H\Sigma_W H^\top) + m_4[Q_K^c],
$$
with
$$
m_4[Q_K^c] \coloneqq \big[Q_K^c H (W_1W_1^\top-\Sigma_W)H^\top\big]^2.
$$
The corresponding asymptotic normality statement is
$$
\frac{1}{\sqrt{t} \sum_{s=1}^t C_s \xrightarrow{d} C_\infty \sim \mathcal N(0,\gamma_M^2(K)),
$$
when \(\gamma_M^2(K)>0\), together with
$$
\frac{1}{t}\sum_{s=1}^t C_s \xrightarrow{a.s.} 0.
$$
The data explicitly emphasize that this framework handles heavy-tailed process noise as long as the noise has a finite fourth moment, and does not require exponential integrability. In simulations with Student-\(t\) noise with \(\nu=5\), the risk-constrained controller reduces long-run fluctuation sensitivity by about \(20\%\) compared with LQR, while the average cost increases by only about \(0.25\%\) [2503.05878].

## 4. Dynamic programming, eigenvalue structure, and duality

In classical ergodic risk-sensitive control, the optimality equations are nonlinear eigenvalue problems rather than additive Bellman equations. For a controlled discrete-time chain with transition kernel \(P(\cdot|x,u)\), the optimality equation is
$$
e^{\gamma\lambda}\Psi(x) = \min_{u\in U(x)} \left[ e^{\gamma c(x,u)} \sum_y \Psi(y)P(y|x,u) \right]
$$
for \(\gamma>0\). For a controlled diffusion with generator
$$
\mathcal L_u f(x)=\operatorname{trace}(a(x)D^2f(x))+b(x,u)\cdot\nabla f(x),
$$
the ergodic risk-sensitive HJB eigen-equation is
$$
\min_{u\in U} \left\{ \mathcal L_u\Psi(x)+\gamma c(x,u)\Psi(x) \right\} = \gamma\lambda^*\,\Psi(x).
$$
For continuous-time Markov chains,
$$
\gamma\lambda\,\Psi(i) = \min_{u\in U(i)} \left[ \sum_j q(j|i,u)\Psi(j) + \gamma c(i,u)\Psi(i) \right].
$$
These are the baseline structures from which constrained formulations inherit their analytical machinery [2301.00224].

The fluctuation-based constrained LQR problem is treated through a Lagrangian,
$$
L(K,\lambda) = J(K)+\lambda\big(\gamma_N^2(K)-\bar\beta\big), \qquad \lambda\ge 0,
$$
or, in the general affine formulation,
$$
L(K,\ell,\lambda)=Q_K(\Sigma_K+\bar x_K\bar x_K^\top)+\lambda(\gamma_N^2-\bar\beta).
$$
Under Slater’s condition,
$$
\exists K\in\mathcal S \text{ such that } \gamma_N^2(K)<\bar\beta,
$$
the constrained problem admits strong duality and can be solved through the saddle-point problem
$$
\sup_{\lambda\ge 0}\inf_{K\in\mathcal S} L(K,\lambda),
$$
with complementary slackness
$$
\lambda^*\big(\gamma_N^2(K^*(\lambda^*))-\bar\beta\big)=0.
$$
In the tractable case \(R^c=0\), the minimizing controller has the Riccati-like form
$$
K^*(\lambda) = -\big(R+B^\top P_{(K^*(\lambda),\lambda)}B\big)^{-1}B^\top P_{(K^*(\lambda),\lambda)}A,
$$
where
$$
P_{(K,\lambda)} = A_K^\top P_{(K,\lambda)}A_K + Q_K + 4\lambda Q^cH\Sigma_WH^\top Q^c.
$$
This makes the constraint enter the synthesis as an additional quadratic penalty in the Riccati equation [2503.05878].

Algorithmically, the same source gives a primal-dual policy optimization method. The inner step uses a Riemannian quasi-Newton or Hewer-style update,
$$
G \gets -(R+B^\top P_{K,\lambda}B)^{-1}\nabla L_\lambda(K)\Sigma_K^{-1}, \qquad
K \gets K+\frac12 G,
$$
while the outer step updates the multiplier by projected ascent,
$$
\lambda_{m+1} = \max\left[0,\lambda_m+\eta_m\big(\gamma_N^2(K)-\bar\beta\big)\right].
$$
The stated complexity for obtaining an \(\epsilon\)-accurate solution is
$$
\mathcal O\!\left(\frac{\ln(\ln(\epsilon))}{\epsilon^2}\right).
$$
This preserves the average-performance objective while enforcing the ergodic-risk constraint through the dual variable [2503.05878].

## 5. Stability, existence, and admissibility conditions

The classical literature repeatedly ties ergodic-risk constraints to existence and stability conditions. For discrete-time chains, the survey lists finite state spaces, countable state spaces with Doeblin or near-monotone conditions, and general Borel spaces via discounted approximation. For controlled diffusions it lists local Lipschitz and nondegeneracy assumptions, coercive or near-monotone running cost, blanket stability or Lyapunov conditions, and uniform ellipticity in some results. For continuous-time Markov chains it lists irreducibility under stationary policies, stability or simultaneous Doeblin conditions, near-monotone costs, and finite jump rates or Lyapunov drift conditions. In many of these settings, the risk-sensitive optimality equation has a positive solution \(\Psi\), and any minimizing selector is optimal [2301.00224].

Near-monotonicity plays a special role because it encourages stabilizing controls. The survey summarizes it as
$$
\liminf_{|x|\to\infty}\min_u c(x,u)>\rho,
$$
for an appropriate threshold \(\rho\), typically the optimal ergodic value. This condition helps ensure existence of an optimal stationary policy even without strong a priori stability assumptions. Later diffusion work sharpens this by combining a two-region structural hypothesis with a Foster–Lyapunov drift condition on one subset and near-monotonicity with inf-compact running cost on the complement, yielding a unique positive solution to the multiplicative HJB equation and a complete characterization of optimal stationary Markov controls [2301.00224] [2511.01100].

The fluctuation-based ergodic-risk constraint literature imposes a different but related admissibility regime. The closed-loop policy is restricted to affine stationary Markov policies
$$
U_t = KX_t+\ell,
$$
with \(K\) stabilizing and \((A,B)\) stabilizable. To obtain irreducibility and a unique invariant measure, the papers further assume that \((A_K,H)\) is controllable, and that the chain is positive Harris recurrent and \(V\)-uniformly ergodic. For quadratic ergodic-risk, the finite fourth moment condition
$$
\mathbb{E}\|W_t\|^4 < \infty
$$
is explicitly required. The papers emphasize that this is what makes the framework applicable to heavy-tailed noise, as long as the fourth moment exists [2409.10767] [2503.05878].

## 6. Related directions, adjacent formulations, and common confusions

The direct constrained formulations above should be distinguished from nearby uses of “ergodic” and “risk.” In multi-agent ergodic exploration under smoke-based visibility, smoke density defines a visibility coefficient
$$
m(s, t) = \begin{cases} 1-\frac{\boldsymbol{\rho}(s,t)}{c} & \boldsymbol{\rho}(s,t) \leq \mathrm{c} \\ 0 & \boldsymbol{\rho}(s,t) > \mathrm{c} \end{cases}
$$
that modifies the expected information distribution, for example through
$$
\Phi_{1} = \mathcal{V} \odot M.
$$
The paper states explicitly that smoke is not formulated as a hard constraint on the optimization problem and is not a paper about explicit “ergodic-risk constraints” or formal risk-sensitive optimization. It is therefore conceptually adjacent, not mathematically equivalent, to ergodic-risk constrained planning [2503.04998].

A second nearby line is finite-horizon multistage risk-constrained control with nested conditional risk mappings. There the proposed constraint is
$$
\bar{r}_t[G_{j,t}] = r_{\mid 0}\Big[ r_{\mid 1}\big[ \cdots r_{\mid t}[G_{j,t}] \big] \Big] \le 0,
$$
which accounts for propagation of uncertainty in time on a scenario tree. The same source states that it does not use “ergodic” in the classical infinite-horizon stationary-average sense. This clarifies a common terminological confusion: nested time-consistent risk constraints and ergodic-risk constraints both regulate cumulative uncertainty over time, but they are not the same construction [1903.06749].

Finance supplies another related but distinct interpretation. Forward entropic risk measures built from exponential forward performance processes are governed by an ergodic BSDE,
$$
dY_t=\big(-F(V_t,Z_t)+\lambda\big)\,dt+Z_t^{tr}\,dW_t,
$$
with forward utility
$$
U(x,t)=-e^{-\gamma x+Y_t-\lambda t}.
$$
The risk measure is represented by
$$
\rho_t(\xi_T)=Y_t^{-\xi_T},
$$
and for long maturities it converges exponentially fast to a constant independent of the initial factor state. This is a long-run risk statement with an ergodic constant \(\lambda\), but it is not an explicit control constraint of the Lagrangian type above. A plausible implication is that the expression “ergodic-risk constraint” now covers a family of long-run, stationary, and tail-sensitive restrictions whose exact mathematical form depends on whether the model is built from exponential growth rates, asymptotic fluctuation limits, or ergodic BSDEs [1607.02289].

Across these strands, the most stable technical picture is the following. Ergodic-risk constraints replace mean-only admissibility by a long-run risk specification; the specification is either a log-moment growth-rate bound or a bound on asymptotic fluctuation variance; and the analysis proceeds through nonlinear eigenvalue problems, ergodic occupation measures, entropy or quadratic penalties, strong duality, and stability conditions that guarantee the existence of optimal stationary controls [2301.00224] [2409.10767] [2503.05878].

Source: https://www.emergentmind.com/topics/ergodic-risk-constraints