Iterated Risk Measures in Decision Processes
- Iterated Risk Measures (IRMs) are dynamic risk measures defined via backward recursion of one-step conditional evaluations that enable time-consistent decision-making.
- IRMs improve upon discounted expected utility by capturing risk preferences in Markov decision processes using conditional mappings like CVaR and entropic risk.
- IRMs support dynamic programming across diverse applications, offering recursive evaluation and convergence guarantees for risk-averse control and modern planning.
Searching arXiv for the cited IRM papers and related dynamic risk-measure work. Iterated Risk Measures (IRMs) are dynamic risk measures built by backward recursion of one-step conditional risk evaluations. In the discounted-cost Markov decision process (MDP) setting, they were introduced to address a limitation of discounted expected utility for representing risk preferences under intertemporal discounting, and to recover decision rules that are consistent over time (Osogami, 2012). In a broader filtered-space formulation, time-consistent dynamic risk measures admit an iterated-composition representation through one-step maps, and under adapted-law invariance and Fatou regularity these one-step maps are conditional lifts of static law-invariant risk measures (Beiglböck et al., 5 Jul 2026). Across discrete-time control, continuous-time finance, and partially observable planning, IRMs provide a recursive alternative to one-shot evaluation of cumulative discounted cost.
1. Formal definition and recursive construction
A static risk measure is a mapping from bounded random variables to that quantifies “riskiness.” Common examples are expectation , the entropic risk measure , and Conditional Tail Expectation (Osogami, 2012). In a finite-horizon MDP with state–action–transition , per-stage cost , and discount factor , one uses a conditional version that maps -measurable costs to 0-measurable quantities. Typical examples are
1
for risk-neutral evaluation,
2
for entropic risk, and
3
The iterated risk measure is defined by backward recursion: 4 for 5. Unfolding the recursion yields the nested form
6
An equivalent discrete-time dynamic-risk formulation uses a filtered probability space 7 with finite filtration
8
and a family of conditional monetary risk mappings
9
In the time-consistent case, the one-step maps
0
give the backward-composition representation
1
(Beiglböck et al., 5 Jul 2026). This places the MDP recursion and the general dynamic-risk recursion in the same iterative framework.
2. Axioms, coherence, and time consistency
The key properties of IRMs are stated through the one-step maps. Monotonicity requires that if 2 almost surely, then 3 almost surely. Translation-invariance requires that for any 4-measurable constant 5,
6
Positive homogeneity requires that for any nonnegative 7-measurable scalar 8,
9
If each 0 is convex in its argument, the IRM is called coherent; strong monotonicity 1 translation 2 homogeneity 3 subadditivity 4 coherence (Osogami, 2012).
In the filtered-space formulation, a dynamic risk measure 5 satisfies monotonicity, cash additivity, normalization 6, and terminal condition 7 (Beiglböck et al., 5 Jul 2026). Time-consistency is defined by
8
for every 9. Standard equivalent forms are the semigroup property, the strong form
0
and the equality form
1
(Beiglböck et al., 5 Jul 2026).
In the discounted-cost MDP setting, time consistency means that the policy that is optimal under 2 today remains optimal under 3 tomorrow. In particular, if for two future-cost profiles 4 one has 5, then also
6
(Osogami, 2012). This property is the structural condition that allows recursive control.
3. Dynamic programming and Bellman recursions
For a policy 7, the value function in a finite-horizon MDP is
8
If each one-step 9 is strongly monotonic, translation-invariant, and, when 0, positively homogeneous, then the optimal value function
1
satisfies a Bellman-like recursion and can be computed by backward induction. An optimal Markov policy 2 is obtained by choosing at each state the 3 in the preceding display (Osogami, 2012).
The algorithmic implications are direct. Value iteration initializes 4, then computes
5
for 6. Policy iteration fixes a policy, evaluates its IRM-value by nested one-step risk recursions, and improves the policy via the risk-sensitive Bellman update. Complexity is comparable to risk-neutral dynamic programming provided each one-step 7 admits an 8 or 9 implementation, as is true for expectation, 0, and 1 (Osogami, 2012).
A canonical partially observable specialization is Iterated Conditional Value-at-Risk (ICVaR). With discounted return 2, the recursion is
3
4
The optimal Bellman equations are
5
6
(Pariente et al., 28 Jan 2026). In that formulation, 7 recovers standard expectation-based planning and 8 induces increasing risk aversion.
4. Limitation of discounted expected utility
A central motivation for IRMs is a limitation of discounted expected utility (DEU) in discounted-cost problems. The example compares two payment methods over 20 days. Method 9 pays 0 on each Day 1 only with probability 2 independently of the day, so the total 3-outlay is 4, with mean 5. A risk-averse agent prefers 6 to 7, because 8 has no chance of 9. Yet for any increasing disutility function 0 and any discount factor 1,
2
so DEU always rates 3 at least as good as 4. Hence no DEU can represent 5 (Osogami, 2012).
A natural modification is to apply expected utility to the cumulative discounted cost, 6, rather than 7. With exponential 8, this yields the entropic risk measure of the sum. Under discounting, however, this loses time consistency. A two-period example with a one-year delay versus a two-year delay shows that a choice that maximizes 9 today can flip one year later, thus violating time consistency (Osogami, 2012).
This is the basic separation between one-shot evaluation of cumulative discounted cost and recursive evaluation through one-step conditional risk maps. A plausible implication is that the distinctive contribution of IRMs lies not only in modeling risk aversion, but in preserving a dynamically stable ordering of future cost streams.
5. Representation theorems, adapted-law invariance, and entropic rigidity
In the discounted-cost setting, the main representation result states that any preference over discounted-cost streams that is time-consistent and respects monotonicity can be represented as an IRM built from suitable one-step conditional risk measures. Preferences such as 0, which no DEU can capture, can be represented by choosing the one-step map to be Conditional Tail Expectation (Osogami, 2012).
The illustrative example uses 1, the IRM whose one-step maps are 2-3 at each stage. For method 4, 5 is constant 6, so by translation-invariance and positive-homogeneity,
7
For method 8, with probability 9 there is no payment and with probability 00 there is payment of 01 each day. By backward recursion, for 02,
03
Hence there is a nonempty region of 04 for which
05
capturing 06 while preserving time consistency (Osogami, 2012).
A more general characterization is given by adapted-law invariance. For a filtered model, the adapted law is encoded by the vector of iterated conditional laws
07
A dynamic risk measure is adapted-law invariant if equality of adapted laws implies equality of 08 for all 09. On a rich atomless filtered space, a relevant, time-consistent, Fatou-regular dynamic risk measure is adapted-law invariant if and only if its one-step maps are conditional lifts of static law-invariant risk measures: 10 Consequently,
11
Moreover, 12 is convex if and only if each 13 is convex, and coherent if and only if each 14 is coherent (Beiglböck et al., 5 Jul 2026).
In the coherent case, each 15 admits an adapted Kusuoka representation through sets 16 of probability measures on 17: 18 The one-step maps become
19
and the full dynamic risk measure is a nested adapted Kusuoka operator (Beiglböck et al., 5 Jul 2026).
The principal contrast is with terminal-law invariance, which requires only that 20 depend on the terminal distribution 21. The finite-horizon Kupper–Schachermayer theorem states that if 22 is relevant, time-consistent, and terminal-law invariant, then there exists 23 such that
24
with expectation and worst-case as limiting cases. Thus entropic risk measures, and their expectation or worst-case limits, are the only time-consistent, relevant maps that ignore the information-structure of 25 (Beiglböck et al., 5 Jul 2026). This clarifies a common misconception: equality of terminal laws does not determine dynamic risk unless one is willing to accept entropic rigidity.
6. Extensions to continuous time and modern planning
IRMs extend beyond finite discounted-cost MDPs. For infinite-horizon discounted MDPs, the theory extends under suitable contraction properties of 26. More general one-step risk measures, including mixtures 27, remain strongly monotonic and allow dynamic programming. Relaxing homogeneity or translation invariance leads to state-augmentation for dynamic programming, increasing complexity. Behavioral modeling—how well IRMs, as compared with DEU or EUD, fit actual human choices under risk plus time trade-off—remains an open question (Osogami, 2012).
A continuous-time extension is given by dynamic spectral risk measures (DSRs). In the discrete-time precursor, one iterates conditional Choquet-integrals with a given probability-distortion 28 on a time grid 29, obtaining an 30-step composition
31
Because each step is a coherent conditional risk measure, the resulting map is strongly time-consistent in the discrete sense (Madan et al., 2013).
In continuous time, on a filtered probability space supporting a Poisson random measure, the dynamic coherent risk measure 32 is defined via the solution of a backward stochastic differential equation with spectral driver 33. The dynamic spectral risk measure is 34, where 35 solves the BSDE. A functional limit theorem shows that any DSR arises in the limit of a sequence of iterated spectral risk-measures driven by lattice-random walks, under suitable scaling and vanishing time- and spatial-mesh sizes. This yields a strongly time-consistent continuous-time extension of iterated spectral risk-measures (Madan et al., 2013).
In partially observable control, ICVaR has been used to define online risk-averse planning algorithms for POMDPs. Sparse Sampling, Particle Filter Trees with Double Progressive Widening, and Partially Observable Monte Carlo Planning with Observation Widening are extended to optimize the ICVaR value function rather than the expectation of the return. Theorem 1 gives finite-time policy-evaluation guarantees, Theorem 2 extends the bounds to sparse sampling, and in both cases the error shrinks like 36 as 37 (Pariente et al., 28 Jan 2026). Empirically, in LaserTag and LightDark, ICVaR-based planners with 38 report lower 39 of cumulative cost than their risk-neutral counterparts; for example, 40-POMCPOW reports 41 versus 42 in LaserTag and 43 versus 44 in LightDark (Pariente et al., 28 Jan 2026).
Taken together, these developments place IRMs at the intersection of risk-averse dynamic programming, coherent and law-invariant risk theory, continuous-time BSDE methods, and modern online planning. The unifying structure is the same: a multi-period risk measure is obtained by iterating one-step conditional risk evaluations backward through time.