Papers
Topics
Authors
Recent
Search
2000 character limit reached

Iterated Risk Measures in Decision Processes

Updated 19 July 2026
  • Iterated Risk Measures (IRMs) are dynamic risk measures defined via backward recursion of one-step conditional evaluations that enable time-consistent decision-making.
  • IRMs improve upon discounted expected utility by capturing risk preferences in Markov decision processes using conditional mappings like CVaR and entropic risk.
  • IRMs support dynamic programming across diverse applications, offering recursive evaluation and convergence guarantees for risk-averse control and modern planning.

Searching arXiv for the cited IRM papers and related dynamic risk-measure work. Iterated Risk Measures (IRMs) are dynamic risk measures built by backward recursion of one-step conditional risk evaluations. In the discounted-cost Markov decision process (MDP) setting, they were introduced to address a limitation of discounted expected utility for representing risk preferences under intertemporal discounting, and to recover decision rules that are consistent over time (Osogami, 2012). In a broader filtered-space formulation, time-consistent dynamic risk measures admit an iterated-composition representation through one-step maps, and under adapted-law invariance and Fatou regularity these one-step maps are conditional lifts of static law-invariant risk measures (Beiglböck et al., 5 Jul 2026). Across discrete-time control, continuous-time finance, and partially observable planning, IRMs provide a recursive alternative to one-shot evaluation of cumulative discounted cost.

1. Formal definition and recursive construction

A static risk measure ρ\rho is a mapping from bounded random variables to R\mathbb{R} that quantifies “riskiness.” Common examples are expectation E[]E[\cdot], the entropic risk measure ERMγ[X]=(1/γ)lnE[eγX]ERM_\gamma[X]=(1/\gamma)\cdot \ln E[e^{\gamma X}], and Conditional Tail Expectation CTE(α)[X]CTE_{(\alpha)}[X] (Osogami, 2012). In a finite-horizon MDP with state–action–transition (X,A,P)(\mathcal{X},\mathcal{A},P), per-stage cost ct(x,a,x)c_t(x,a,x'), and discount factor 0<γ10<\gamma\le 1, one uses a conditional version ρt\rho_t that maps Ft+1\mathcal{F}_{t+1}-measurable costs to R\mathbb{R}0-measurable quantities. Typical examples are

R\mathbb{R}1

for risk-neutral evaluation,

R\mathbb{R}2

for entropic risk, and

R\mathbb{R}3

The iterated risk measure is defined by backward recursion: R\mathbb{R}4 for R\mathbb{R}5. Unfolding the recursion yields the nested form

R\mathbb{R}6

(Osogami, 2012).

An equivalent discrete-time dynamic-risk formulation uses a filtered probability space R\mathbb{R}7 with finite filtration

R\mathbb{R}8

and a family of conditional monetary risk mappings

R\mathbb{R}9

In the time-consistent case, the one-step maps

E[]E[\cdot]0

give the backward-composition representation

E[]E[\cdot]1

(Beiglböck et al., 5 Jul 2026). This places the MDP recursion and the general dynamic-risk recursion in the same iterative framework.

2. Axioms, coherence, and time consistency

The key properties of IRMs are stated through the one-step maps. Monotonicity requires that if E[]E[\cdot]2 almost surely, then E[]E[\cdot]3 almost surely. Translation-invariance requires that for any E[]E[\cdot]4-measurable constant E[]E[\cdot]5,

E[]E[\cdot]6

Positive homogeneity requires that for any nonnegative E[]E[\cdot]7-measurable scalar E[]E[\cdot]8,

E[]E[\cdot]9

If each ERMγ[X]=(1/γ)lnE[eγX]ERM_\gamma[X]=(1/\gamma)\cdot \ln E[e^{\gamma X}]0 is convex in its argument, the IRM is called coherent; strong monotonicity ERMγ[X]=(1/γ)lnE[eγX]ERM_\gamma[X]=(1/\gamma)\cdot \ln E[e^{\gamma X}]1 translation ERMγ[X]=(1/γ)lnE[eγX]ERM_\gamma[X]=(1/\gamma)\cdot \ln E[e^{\gamma X}]2 homogeneity ERMγ[X]=(1/γ)lnE[eγX]ERM_\gamma[X]=(1/\gamma)\cdot \ln E[e^{\gamma X}]3 subadditivity ERMγ[X]=(1/γ)lnE[eγX]ERM_\gamma[X]=(1/\gamma)\cdot \ln E[e^{\gamma X}]4 coherence (Osogami, 2012).

In the filtered-space formulation, a dynamic risk measure ERMγ[X]=(1/γ)lnE[eγX]ERM_\gamma[X]=(1/\gamma)\cdot \ln E[e^{\gamma X}]5 satisfies monotonicity, cash additivity, normalization ERMγ[X]=(1/γ)lnE[eγX]ERM_\gamma[X]=(1/\gamma)\cdot \ln E[e^{\gamma X}]6, and terminal condition ERMγ[X]=(1/γ)lnE[eγX]ERM_\gamma[X]=(1/\gamma)\cdot \ln E[e^{\gamma X}]7 (Beiglböck et al., 5 Jul 2026). Time-consistency is defined by

ERMγ[X]=(1/γ)lnE[eγX]ERM_\gamma[X]=(1/\gamma)\cdot \ln E[e^{\gamma X}]8

for every ERMγ[X]=(1/γ)lnE[eγX]ERM_\gamma[X]=(1/\gamma)\cdot \ln E[e^{\gamma X}]9. Standard equivalent forms are the semigroup property, the strong form

CTE(α)[X]CTE_{(\alpha)}[X]0

and the equality form

CTE(α)[X]CTE_{(\alpha)}[X]1

(Beiglböck et al., 5 Jul 2026).

In the discounted-cost MDP setting, time consistency means that the policy that is optimal under CTE(α)[X]CTE_{(\alpha)}[X]2 today remains optimal under CTE(α)[X]CTE_{(\alpha)}[X]3 tomorrow. In particular, if for two future-cost profiles CTE(α)[X]CTE_{(\alpha)}[X]4 one has CTE(α)[X]CTE_{(\alpha)}[X]5, then also

CTE(α)[X]CTE_{(\alpha)}[X]6

(Osogami, 2012). This property is the structural condition that allows recursive control.

3. Dynamic programming and Bellman recursions

For a policy CTE(α)[X]CTE_{(\alpha)}[X]7, the value function in a finite-horizon MDP is

CTE(α)[X]CTE_{(\alpha)}[X]8

If each one-step CTE(α)[X]CTE_{(\alpha)}[X]9 is strongly monotonic, translation-invariant, and, when (X,A,P)(\mathcal{X},\mathcal{A},P)0, positively homogeneous, then the optimal value function

(X,A,P)(\mathcal{X},\mathcal{A},P)1

satisfies a Bellman-like recursion and can be computed by backward induction. An optimal Markov policy (X,A,P)(\mathcal{X},\mathcal{A},P)2 is obtained by choosing at each state the (X,A,P)(\mathcal{X},\mathcal{A},P)3 in the preceding display (Osogami, 2012).

The algorithmic implications are direct. Value iteration initializes (X,A,P)(\mathcal{X},\mathcal{A},P)4, then computes

(X,A,P)(\mathcal{X},\mathcal{A},P)5

for (X,A,P)(\mathcal{X},\mathcal{A},P)6. Policy iteration fixes a policy, evaluates its IRM-value by nested one-step risk recursions, and improves the policy via the risk-sensitive Bellman update. Complexity is comparable to risk-neutral dynamic programming provided each one-step (X,A,P)(\mathcal{X},\mathcal{A},P)7 admits an (X,A,P)(\mathcal{X},\mathcal{A},P)8 or (X,A,P)(\mathcal{X},\mathcal{A},P)9 implementation, as is true for expectation, ct(x,a,x)c_t(x,a,x')0, and ct(x,a,x)c_t(x,a,x')1 (Osogami, 2012).

A canonical partially observable specialization is Iterated Conditional Value-at-Risk (ICVaR). With discounted return ct(x,a,x)c_t(x,a,x')2, the recursion is

ct(x,a,x)c_t(x,a,x')3

ct(x,a,x)c_t(x,a,x')4

The optimal Bellman equations are

ct(x,a,x)c_t(x,a,x')5

ct(x,a,x)c_t(x,a,x')6

(Pariente et al., 28 Jan 2026). In that formulation, ct(x,a,x)c_t(x,a,x')7 recovers standard expectation-based planning and ct(x,a,x)c_t(x,a,x')8 induces increasing risk aversion.

4. Limitation of discounted expected utility

A central motivation for IRMs is a limitation of discounted expected utility (DEU) in discounted-cost problems. The example compares two payment methods over 20 days. Method ct(x,a,x)c_t(x,a,x')9 pays 0<γ10<\gamma\le 10 on each Day 0<γ10<\gamma\le 11 only with probability 0<γ10<\gamma\le 12 independently of the day, so the total 0<γ10<\gamma\le 13-outlay is 0<γ10<\gamma\le 14, with mean 0<γ10<\gamma\le 15. A risk-averse agent prefers 0<γ10<\gamma\le 16 to 0<γ10<\gamma\le 17, because 0<γ10<\gamma\le 18 has no chance of 0<γ10<\gamma\le 19. Yet for any increasing disutility function ρt\rho_t0 and any discount factor ρt\rho_t1,

ρt\rho_t2

so DEU always rates ρt\rho_t3 at least as good as ρt\rho_t4. Hence no DEU can represent ρt\rho_t5 (Osogami, 2012).

A natural modification is to apply expected utility to the cumulative discounted cost, ρt\rho_t6, rather than ρt\rho_t7. With exponential ρt\rho_t8, this yields the entropic risk measure of the sum. Under discounting, however, this loses time consistency. A two-period example with a one-year delay versus a two-year delay shows that a choice that maximizes ρt\rho_t9 today can flip one year later, thus violating time consistency (Osogami, 2012).

This is the basic separation between one-shot evaluation of cumulative discounted cost and recursive evaluation through one-step conditional risk maps. A plausible implication is that the distinctive contribution of IRMs lies not only in modeling risk aversion, but in preserving a dynamically stable ordering of future cost streams.

5. Representation theorems, adapted-law invariance, and entropic rigidity

In the discounted-cost setting, the main representation result states that any preference over discounted-cost streams that is time-consistent and respects monotonicity can be represented as an IRM built from suitable one-step conditional risk measures. Preferences such as Ft+1\mathcal{F}_{t+1}0, which no DEU can capture, can be represented by choosing the one-step map to be Conditional Tail Expectation (Osogami, 2012).

The illustrative example uses Ft+1\mathcal{F}_{t+1}1, the IRM whose one-step maps are Ft+1\mathcal{F}_{t+1}2-Ft+1\mathcal{F}_{t+1}3 at each stage. For method Ft+1\mathcal{F}_{t+1}4, Ft+1\mathcal{F}_{t+1}5 is constant Ft+1\mathcal{F}_{t+1}6, so by translation-invariance and positive-homogeneity,

Ft+1\mathcal{F}_{t+1}7

For method Ft+1\mathcal{F}_{t+1}8, with probability Ft+1\mathcal{F}_{t+1}9 there is no payment and with probability R\mathbb{R}00 there is payment of R\mathbb{R}01 each day. By backward recursion, for R\mathbb{R}02,

R\mathbb{R}03

Hence there is a nonempty region of R\mathbb{R}04 for which

R\mathbb{R}05

capturing R\mathbb{R}06 while preserving time consistency (Osogami, 2012).

A more general characterization is given by adapted-law invariance. For a filtered model, the adapted law is encoded by the vector of iterated conditional laws

R\mathbb{R}07

A dynamic risk measure is adapted-law invariant if equality of adapted laws implies equality of R\mathbb{R}08 for all R\mathbb{R}09. On a rich atomless filtered space, a relevant, time-consistent, Fatou-regular dynamic risk measure is adapted-law invariant if and only if its one-step maps are conditional lifts of static law-invariant risk measures: R\mathbb{R}10 Consequently,

R\mathbb{R}11

Moreover, R\mathbb{R}12 is convex if and only if each R\mathbb{R}13 is convex, and coherent if and only if each R\mathbb{R}14 is coherent (Beiglböck et al., 5 Jul 2026).

In the coherent case, each R\mathbb{R}15 admits an adapted Kusuoka representation through sets R\mathbb{R}16 of probability measures on R\mathbb{R}17: R\mathbb{R}18 The one-step maps become

R\mathbb{R}19

and the full dynamic risk measure is a nested adapted Kusuoka operator (Beiglböck et al., 5 Jul 2026).

The principal contrast is with terminal-law invariance, which requires only that R\mathbb{R}20 depend on the terminal distribution R\mathbb{R}21. The finite-horizon Kupper–Schachermayer theorem states that if R\mathbb{R}22 is relevant, time-consistent, and terminal-law invariant, then there exists R\mathbb{R}23 such that

R\mathbb{R}24

with expectation and worst-case as limiting cases. Thus entropic risk measures, and their expectation or worst-case limits, are the only time-consistent, relevant maps that ignore the information-structure of R\mathbb{R}25 (Beiglböck et al., 5 Jul 2026). This clarifies a common misconception: equality of terminal laws does not determine dynamic risk unless one is willing to accept entropic rigidity.

6. Extensions to continuous time and modern planning

IRMs extend beyond finite discounted-cost MDPs. For infinite-horizon discounted MDPs, the theory extends under suitable contraction properties of R\mathbb{R}26. More general one-step risk measures, including mixtures R\mathbb{R}27, remain strongly monotonic and allow dynamic programming. Relaxing homogeneity or translation invariance leads to state-augmentation for dynamic programming, increasing complexity. Behavioral modeling—how well IRMs, as compared with DEU or EUD, fit actual human choices under risk plus time trade-off—remains an open question (Osogami, 2012).

A continuous-time extension is given by dynamic spectral risk measures (DSRs). In the discrete-time precursor, one iterates conditional Choquet-integrals with a given probability-distortion R\mathbb{R}28 on a time grid R\mathbb{R}29, obtaining an R\mathbb{R}30-step composition

R\mathbb{R}31

Because each step is a coherent conditional risk measure, the resulting map is strongly time-consistent in the discrete sense (Madan et al., 2013).

In continuous time, on a filtered probability space supporting a Poisson random measure, the dynamic coherent risk measure R\mathbb{R}32 is defined via the solution of a backward stochastic differential equation with spectral driver R\mathbb{R}33. The dynamic spectral risk measure is R\mathbb{R}34, where R\mathbb{R}35 solves the BSDE. A functional limit theorem shows that any DSR arises in the limit of a sequence of iterated spectral risk-measures driven by lattice-random walks, under suitable scaling and vanishing time- and spatial-mesh sizes. This yields a strongly time-consistent continuous-time extension of iterated spectral risk-measures (Madan et al., 2013).

In partially observable control, ICVaR has been used to define online risk-averse planning algorithms for POMDPs. Sparse Sampling, Particle Filter Trees with Double Progressive Widening, and Partially Observable Monte Carlo Planning with Observation Widening are extended to optimize the ICVaR value function rather than the expectation of the return. Theorem 1 gives finite-time policy-evaluation guarantees, Theorem 2 extends the bounds to sparse sampling, and in both cases the error shrinks like R\mathbb{R}36 as R\mathbb{R}37 (Pariente et al., 28 Jan 2026). Empirically, in LaserTag and LightDark, ICVaR-based planners with R\mathbb{R}38 report lower R\mathbb{R}39 of cumulative cost than their risk-neutral counterparts; for example, R\mathbb{R}40-POMCPOW reports R\mathbb{R}41 versus R\mathbb{R}42 in LaserTag and R\mathbb{R}43 versus R\mathbb{R}44 in LightDark (Pariente et al., 28 Jan 2026).

Taken together, these developments place IRMs at the intersection of risk-averse dynamic programming, coherent and law-invariant risk theory, continuous-time BSDE methods, and modern online planning. The unifying structure is the same: a multi-period risk measure is obtained by iterating one-step conditional risk evaluations backward through time.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Iterated Risk Measures (IRMs).