Ergodic Risk Measures
- Ergodic risk measures are long-run risk functionals defined via asymptotic limits rather than fixed terminal payoffs, providing a framework for understanding stability in stochastic systems.
- They encompass diverse approaches including forward entropic risk via ergodic BSDEs, exponential risk-sensitive control, and asymptotic fluctuation analysis in reinforcement learning and Markov processes.
- These measures offer practical insights for robust control and risk management by quantifying long-run uncertainty and adapting to dynamic, continual decision-making settings.
Ergodic risk measures are long-run risk functionals for stochastic systems whose defining object is asymptotic rather than single-horizon behavior. In the literature, this label covers several distinct constructions: maturity-independent forward entropic risk measures generated by ergodic backward stochastic differential equations (BSDEs), exponential long-run growth criteria in ergodic risk-sensitive control, innovation-based asymptotic fluctuation criteria for controlled Markov chains, and dynamic risk-measure sequences with asymptotic forgetting in continual reinforcement learning. What unifies these strands is that risk is evaluated through stationary, ergodic, or large-time limits rather than through a fixed terminal payoff alone (Biswas et al., 2022, Chong et al., 2016, Rojas et al., 3 Oct 2025).
1. Multiple notions in the literature
The literature does not use a single canonical definition of an ergodic risk measure. Instead, several mathematically distinct objects are used to quantify long-run risk.
| Strand | Defining object | Emphasis |
|---|---|---|
| Forward entropic risk | maturity-independence | |
| Ergodic risk-sensitive control | exponential long-run growth | |
| Ergodic-risk criterion | cumulative uncertainty | |
| Continual-RL ergodic risk measure | satisfying asymptotic plasticity and local time consistency | forgetting of remote past |
In the survey literature on ergodic risk-sensitive control, the core object is the long-run growth rate of an exponential moment of accumulated cost,
or its discrete-time analogue. This is presented as a control-theoretic framework for ergodic risk measures because it evaluates an asymptotic log-moment generating rate rather than only an expected average cost (Biswas et al., 2022).
A different recent strand defines ergodic-risk criteria from the unpredictable component of a stagewise risk functional. If
then the normalized cumulative uncertainty
is studied through functional central limit theorems, together with the asymptotic conditional variance
Here the emphasis is not an exponential utility, but long-run stochastic fluctuations around predictable evolution (Talebi et al., 2024, Talebi et al., 10 Feb 2025).
In continual reinforcement learning, the term is formalized differently again. A sequence of conditional risk measures
is called an ergodic risk measure if it satisfies asymptotic plasticity and local time consistency. This construction is designed for continual settings in which risk evaluation must adapt and asymptotically forget remote history (Rojas et al., 3 Oct 2025).
2. Forward entropic risk measures and ergodic BSDEs
A mathematically explicit example of an ergodic risk measure arises in the theory of forward entropic risk. The forward utility is
where 0 solves the ergodic BSDE
1
with driver
2
Under Assumptions 1–2, this equation admits a unique Markovian solution
3
with 4, 5 of at most linear growth, and 6 bounded. The resulting forward entropic risk measure is maturity-independent because it is built from a forward performance process defined for all times, rather than from a single fixed terminal date (Chong et al., 2016).
For a bounded risk position 7, the 8-normalized forward entropic risk measure 9 is defined implicitly by
0
and for general 1 with maturity 2 one sets
3
Its main representation is another BSDE: 4 where
5
and the key identity is
6
Thus the forward risk measure is the solution of a backward equation whose driver depends on the ergodic BSDE volatility 7 defining the forward utility.
Because 8 is convex, 9 is convex in 0, with convex dual
1
The corresponding dual representation is
2
This yields anti-positivity, convexity, and cash translativity,
3
The large-maturity regime is especially characteristic. For claims of the form 4 with 5 bounded and Lipschitz, there exists a constant 6, independent of the initial factor 7, such that
8
with exponential rate
9
The associated hedging strategy decays in the sense that for each fixed finite 0,
1
The paper attributes this stability to dissipativity of the factor drift,
2
which forces exponential contraction of trajectories. The same work also proves the parity identity
3
expressing the forward entropic risk measure as a difference of two classical entropic risk measures (Chong et al., 2016).
3. Exponential long-run risk in stochastic control
The oldest and most developed control-theoretic notion of ergodic risk is ergodic risk-sensitive control. Its basic object is the infinite-horizon exponential criterion
4
or, with explicit risk parameter,
5
This criterion is emphasized as fluctuation-sensitive, connected with principal eigenvalues, multiplicative dynamic programming, large deviations, and 6 control. It is also explicitly distinguished from simply minimizing average cost and then adding a variance penalty (Biswas et al., 2022).
Across model classes, the mathematical structure is an eigenvalue or multiplicative HJB equation. For controlled diffusions, the survey gives
7
and the logarithmic transform 8 converts this into a nonlinear additive HJB/HJI-type equation. For controlled Markov processes on countable state spaces, the discrete-time equation is
9
while the continuous-time analogue is
0
Under blanket stability assumptions and, separately, under near-monotonicity conditions, existence, uniqueness up to normalization, verification for optimal stationary Markov controls, and policy improvement algorithms are established (Biswas et al., 2021).
Recent diffusion results extend this program under a mixed structural hypothesis. One considers a partition of state space with inf-compactness of the running cost on one set and a Foster–Lyapunov-type drift condition on its complement. Under these conditions there exists a unique positive 1 solution to
2
optimal stationary Markov controls are exactly the minimizers of the HJB, and the admissible and stationary optimal values coincide,
3
The proof uses the Boué–Dupuis variational formula for exponential Brownian functionals and an extended diffusion with an auxiliary control carrying quadratic penalty 4 (Anugu et al., 2 Nov 2025).
A queueing-theoretic version appears in multiclass many-server systems with abandonment in the Halfin--Whitt regime. There the ergodic risk-sensitive criterion is the long-run exponential average cost
5
and the limiting diffusion satisfies the eigenvalue equation
6
The main result is the asymptotic optimality statement
7
Because the exponential criterion is not directly an occupation-measure functional, the analysis relies on Brownian and Poisson variational representations, auxiliary controls, and tightness of mean empirical measures for extended processes (Anugu et al., 2024).
4. Innovation-based ergodic-risk criteria and constrained synthesis
A separate line of work defines ergodic risk through the asymptotic fluctuations of the unpredictable part of a stagewise risk functional. For the controlled linear system
8
one starts from a measurable risk functional
9
and defines the one-step ergodic-risk increment
0
The long-run ergodic-risk criterion is the normalized sum
1
with asymptotic conditional variance
2
The key technical point is that the summands are correlated through the dynamics, so ordinary CLTs do not apply directly; the theory instead uses 3-uniform ergodicity, Foster–Lyapunov drift conditions, and functional central limit theorems for additive functionals of general-state Markov chains (Talebi et al., 2024).
This framework is designed to handle heavy-tailed disturbances. For quadratic risk functionals on stochastic linear systems, finite fourth moments of the process noise are sufficient to obtain well-defined asymptotic variances and CLT-type limits. In particular, if
4
then under stabilizing stationary affine policies the normalized cumulative uncertainty converges to a Gaussian limit when the asymptotic variance is positive, and the asymptotic conditional variance converges almost surely. The framework is explicitly contrasted with standard average-cost LQR and with exponential risk-sensitive control, and it is presented as remaining meaningful for non-Gaussian, even heavy-tailed, disturbances provided relevant moments exist (Talebi et al., 2024).
A specialized formulation for linear stationary Markov policies 5 produces an ergodic-risk constrained LQR. With
6
and invariant covariance 7 solving
8
the average objective is
9
and the risk constraint is
0
The theorem gives the explicit formula
1
Under 2, full row rank of 3, and Slater’s condition, the constrained problem admits strong duality and is solved by a primal-dual algorithm with an inner Riemannian Newton/Hewer-type step and an outer subgradient ascent update for the multiplier. The reported convergence complexity is
4
for an 5-accurate solution (Talebi et al., 10 Feb 2025).
5. Continual reinforcement learning and asymptotic plasticity
In continual reinforcement learning, ergodic risk measures are introduced to address a failure of classical risk-measure theory in indefinite-horizon adaptive settings. The motivating problem is an agent operating across a stream of changing environments or tasks 6, indexed by 7, where risk evaluation should evolve rather than remain fixed forever. The paper argues that the same stability–plasticity tension governing continual learning should govern risk assessment (Rojas et al., 3 Oct 2025).
Two incompatibility results structure the argument. First, static risk measures
8
do not satisfy fixed or asymptotic plasticity. Second, nested dynamic risk measures of the form
9
also fail both plasticity properties because the recursion propagates dependence through the entire trajectory. The replacement is a sequence 0 satisfying local time consistency
1
for all 2, and asymptotic plasticity, meaning that after some finite time step 3 the influence of sufficiently old history vanishes: 4 with 5 depending only on post-6 information. An ergodic risk measure is then defined precisely as a sequence satisfying asymptotic plasticity and local time consistency (Rojas et al., 3 Oct 2025).
The associated long-run objective is
7
Under the unichain assumption for prediction or the communicating assumption for control, the objective becomes independent of initial conditions in the long run. Theorem 1 states that, given such an ergodicity-like assumption and a stationary policy 8, this objective corresponds to an ergodic risk measure. The proof splits the long-run average at a large finite time and uses Birkhoff’s Ergodic Theorem to show that early transient terms vanish, which yields asymptotic plasticity.
The case study uses CVaR. With
9
and, when the distribution is continuous at the quantile,
0
the continual objective becomes
1
The paper studies two continual versions of the red-pill/blue-pill environment and implements a tabular RED CVaR Q-learning algorithm using the transformed reward
2
The reported experiments are intended as a case study rather than a general algorithmic theory (Rojas et al., 3 Oct 2025).
6. Nonlinear expectations, invariance, and stability issues
Under model uncertainty, the relation between long-run risk and invariant behavior becomes subtler. For G-diffusions driven by G-Brownian motion, the long-time limit
3
defines a unique invariant expectation, while the time-average limit
4
defines an ergodic expectation. Both are sublinear expectations represented by weakly compact families of probability measures, but the paper shows that they need not coincide, unlike in the classical linear case. The ergodic quantity is characterized through the fully nonlinear elliptic PDE
5
This directly contradicts the common classical intuition that invariant and ergodic objects are interchangeable in the long run (Hu et al., 2014).
A different caution comes from ergodic invariant measures with infinite entropy. For generic continuous maps and, in dimension at least two, generic homeomorphisms on compact manifolds, one can construct ergodic invariant measures 6 with
7
that converge in the weak8 topology to a periodic-orbit measure 9 with
00
This is not a paper on risk measures in finance or control, but it shows that entropy-based ergodic quantities can be extremely unstable under weak01 perturbations. A plausible implication is that any ergodic risk functional increasing with entropy, complexity, or asymptotic unpredictability needs to account for this non-upper-semicontinuous behavior (Catsigeras et al., 2019).
Taken together, these results indicate that ergodic risk is not exhausted by a single invariant-law statistic. In some settings it is a principal eigenvalue; in others it is a BSDE value process, an asymptotic fluctuation variance, or a plastic dynamic risk sequence. This suggests that the modern theory of ergodic risk measures is best understood as a family of long-run risk formalisms, each tied to a specific asymptotic regime, structural assumption, and notion of admissible adaptation.