---
title: Optimized Certainty Equivalents (OCE)
url: https://www.emergentmind.com/topics/optimized-certainty-equivalents-oce
type: topic
---

# Optimized Certainty Equivalents (OCE)

Optimized certainty equivalents (OCEs) are certainty-equivalent and risk-measure constructions that evaluate a random payoff or loss by jointly optimizing over a deterministic shift variable and a nonlinear utility or disutility transform. In the literature covered here, two sign conventions are standard: a utility-based reward convention of the form \(\sup_{\eta\in\mathbb R}\{\eta+\mathbb E[u(Z-\eta)]\}\), and a loss-based convention of the form \(\inf_{\xi\in\mathbb R}\{\xi+\mathbb E[\phi(X-\xi)]\}\) or \(\inf_{\eta\in\mathbb R}\{E[l(\eta-X)]-\eta\}\). These formulations generate broad families of law-invariant, cash-additive, convex or concave functionals that include expected loss, entropic risk, CVaR/AVaR, mean-variance, and monotone mean-variance, and they now appear in statistical estimation, robust finance, stochastic control, reinforcement learning, conformal risk control, and individualized decision-making [1212.6732][1908.10742][2405.20933].

## 1. Definition, conventions, and mathematical structure

In the utility convention used for rewards, the classical OCE is
\[
\mathcal O_u(\mathcal Z)\triangleq \sup_{\eta\in\mathbb R}\left\{\eta+\mathbb E[u(\mathcal Z-\eta)]\right\},
\]
where \(u\) is upper semicontinuous, satisfies \(u(0)=0\), and \(1\in\partial u(0)\). In this formulation, \(\eta\) is interpreted as present consumption, \(\mathcal Z-\eta\) as uncertain future consumption, and \(-\mathcal O_u(\mathcal Z)\) is a convex risk measure [1908.10742]. In the loss convention used in risk measurement and estimation, one writes
\[
\oce(X):=\inf_{\xi}\left\{\xi+\mathbb E[\phi(X-\xi)]\right\},
\]
with \(\phi:\mathbb R\to\mathbb R^+\cup\{0\}\) nondecreasing, closed, convex, \(\phi(0)=0\), and \(\phi'(0)=1\); under these assumptions, \(\oce(X)\) is a convex risk measure with translation invariance, consistency on constants, and monotonicity [2405.20933]. A closely related loss-function form is
\[
\rho(X):=\inf_{\eta\in\mathbb R}\{E[l(\eta-X)]-\eta\},
\]
where \(l\) is increasing, convex, satisfies \(l(0)=0\) and \(l(x)\ge x\), and often also \(l(x)>x\) for all sufficiently large \(x\) [1212.6732].

A central structural fact is that OCE reduces an infinite-dimensional risk evaluation to a scalar optimization. In the differentiable loss convention of [1212.6732], the optimal allocation \(\eta^*\) satisfies
\[
E[l'(\eta^*-X)]=1,
\]
while in the differentiable disutility convention of [2405.20933], the optimizer \(e^*\) satisfies
\[
\mathbb E[\phi'(X-e^*)]=1.
\]
In the nondifferentiable case, subgradient inequalities replace equality. This scalar first-order condition is the core mechanism behind both numerical algorithms and statistical estimators [1212.6732][2405.20933].

A recurring source of confusion is sign convention rather than substance. The reward-side supremum and the loss-side infimum are both standard in this literature, and several papers explicitly switch between them according to whether the primitive random variable is interpreted as reward, payoff, income, or loss [1908.10742][2405.20933][2512.02386].

## 2. Canonical special cases

Many familiar risk criteria arise by specifying the utility or disutility generator. Several papers emphasize that CVaR is only one member of a much larger OCE family, not the defining example [1212.6732][2405.20933].

| Generator | Resulting OCE | Notes |
|---|---|---|
| \(\phi(t)=t\) | \(\oce(X)=\mathbb E[X]\) | Expected loss [2405.20933] |
| \(l(x)=\frac{e^{\gamma x}-1}{\gamma}\) or \(\phi_\gamma(t)=\frac{1}{\gamma}(e^{\gamma t}-1)\) | \(\frac{1}{\gamma}\ln(E[e^{\gamma X}])\) in loss convention | Entropic risk [1212.6732][2405.20933] |
| Piecewise linear \(l\) or \(\phi_\alpha(t)=\frac{1}{1-\alpha}[t]_+\) | CVaR / AVaR | Quantile-based OCE [1212.6732][2405.20933] |
| \(\phi_c(t)=t+ct^2\) | \(\mathbb E[X]+c\,\mathrm{Var}(X)\) | Mean-variance [2405.20933] |
| \(l(x)=\frac{([1+x]^+)^\gamma-1}{\gamma}\), \(\gamma=2\) | Monotone mean-variance | Polynomial OCE family [1212.6732] |

For CVaR in the loss convention, [1212.6732] gives
\[
CV@R_\lambda(X)=-\frac{1}{\lambda}\int_0^\lambda q_X^+(s)\,ds
=\frac{1}{\lambda}\int_0^\lambda V@R_s(X)\,ds,
\]
and shows that the OCE optimizer is the quantile \(\eta^*=q_X^+(\lambda)\). For mean-variance in [2405.20933], the OCE optimizer is \(e^*=\mathbb E[X]\), giving the exact decomposition \(\oce(X)=\mathbb E[X]+c\,\mathrm{Var}(X)\). The polynomial family in [1212.6732] includes the \(\gamma=2\) monotone mean-variance case and higher-order loss functions for which no closed form is available.

Recursive reinforcement-learning work extends this catalog. In discounted MDPs, the paper on sample complexity lists entropic risk, CVaR, and mean-variance as learnable recursive OCE objectives, and also identifies an essential-infimum-type utility with non-full domain as a contrasting non-PAC-learnable case [2605.21763]. Continuous-time work likewise lists linear utility, exponential utility, power utility, logarithmic utility, CVaR, mean-variance, and monotone mean-variance as OCE examples in a unified control framework [2512.02386].

## 3. Duality, robustness, and generalizations

Classical OCEs admit robust dual representations. In the loss-function setting of [1212.6732], if \(l^*\) is the convex conjugate, then
\[
\rho(X)=\max_{Q\in\mathcal M_{1,l^*}(P)}
\left\{E_Q[-X]-E_P\!\left[l^*\!\left(\frac{dQ}{dP}\right)\right]\right\}.
\]
This identifies OCE as a penalized worst-case expectation over absolutely continuous measures and underlies later robust, conditional, and dynamic extensions [1212.6732].

Robust OCE under model uncertainty can often be reduced from infinite-dimensional optimization over distributions to finite-dimensional optimization over a scalar transport-penalty parameter. For transport-penalized ambiguity,
\[
\mathcal{R}(f):=\sup_{\mu\in\mathcal M_1(\mathbb R)}
\left(\int f\,d\mu-\varphi(d_c(\mu_0,\mu))\right),
\]
the robust OCE satisfies
\[
\mathcal{OCE}(l)=\inf_{\lambda\ge 0}\left(\mathrm{OCE}(l^{\lambda c})+\varphi^*(\lambda)\right),
\]
and in several AVaR cases this yields explicit uncertainty premiums, such as \(\delta/\alpha\) under a Wasserstein-ball specification with \(c(x,y)=|x-y|\) [1706.10186].

The conditional counterpart replaces deterministic cash shifts by \(\mathcal G\)-measurable random shifts. On \(L^\infty(\mathcal F\mid\mathcal G)\), the conditional OCE studied by Principi and Maccheroni is
\[
\sup_{a\in L^0(\mathcal G)}
\left\{a-E_\mu[\varphi^*(a-x)\mid\mathcal G]\right\},
\]
and its conditional variational formula is
\[
\inf_{\nu\in\mathcal M(\mathcal G)}
\left\{E_\nu[x\mid\mathcal G]+D_{\varphi,\mathcal G}(\nu\|\mu)\right\}
=
\sup_{a\in L^0(\mathcal G)}
\left\{a-E_\mu[\varphi^*(a-x)\mid\mathcal G]\right\}.
\]
The random-modular approach is essential here because both coefficients and dual variables live in \(L^0(\mathcal G)\), not in \(\mathbb R\) [2211.04592].

Multivariate OCE replaces a scalar cash shift by a vector allocation \(w\in\mathbb R^d\):
\[
R(X)=\inf_{w\in\mathbb R^d}\left\{\sum_{i=1}^d w_i+E[l(-X-w)]\right\}.
\]
This yields a scalar systemic risk measure together with an optimal allocation \(m^*\), characterized in the differentiable case by
\[
1=E[\nabla l(-X-m^*)].
\]
The resulting risk measure is convex, monotone, cash invariant, continuous, and admits the dual representation
\[
R(X)=\max_{Q\in\mathcal D^{l^*}}\{E_Q[-X]-E[l^*(dQ/dP)]\},
\]
with the dependence structure entering through a genuinely multivariate loss function rather than through prior scalar aggregation [2210.13825].

Several papers also enlarge the OCE template itself. In precision medicine, the optimized covariate-dependent equivalent replaces the scalar allocation by a measurable function \(\alpha(X)\):
\[
\mathcal O^{\,d}_{(u,\mathcal F)}(\mathcal Z)
=
\sup_{\alpha\in\mathcal F}
\left\{\mathbb E[\alpha(X)]+\mathbb E^d[u(\mathcal Z-\alpha(X))]\right\},
\]
thereby turning classical OCE into a criterion for individualized decision rules [1908.10742]. A different variant, the modified OCE,
\[
M_u(\xi):=\sup_{x\in\mathbb R}\{u(x)+\mathbb E_P[u(\xi-x)]\},
\]
puts present and future terms in the same utility units, and the preference-robust version replaces \(u\) by a worst-case utility over an ambiguity set. The paper presenting this variant proves law invariance, monotonicity, concavity, and positive subhomogeneity, while also stressing that the modified formulation is not presented as translation invariant in the classical cash-additive sense [2203.10762].

## 4. Computation and statistical estimation

A major computational theme is that OCE often reduces to scalar optimization plus transform or sampling machinery. For univariate OCE with known moment generating function \(M_X\), [1212.6732] develops a Fourier method in which computation consists of: first, solving a one-dimensional root-finding problem for the optimal allocation \(\eta^*\); and second, evaluating one or two Fourier integrals. In the differentiable case, \(\eta^*\) is the unique root of
\[
f(\eta)=
\frac{1}{2\pi}\int_{\mathbb R}
e^{(R'-iu)\eta}M_X(iu-R')\widehat{l'}(u+iR')\,du-1,
\]
and then
\[
\rho(X)=
\frac{1}{2\pi}\int_{\mathbb R}
e^{(R-iu)\eta^*}M_X(iu-R)\widehat l(u+iR)\,du-\eta^*.
\]
The same paper derives a specialized Fourier formula for CVaR and argues that this makes CVaR computation comparable in time to VaR, because the method replaces repeated quantile evaluations by a single root solve and a single Fourier integral [1212.6732].

Statistical estimation from i.i.d. data is treated in [2405.20933]. The sample average approximation is
\[
\oce_n^\phi=
\inf_\xi\left\{\xi+\frac1n\sum_{i=1}^n\phi(X_i-\xi)\right\},
\]
with empirical optimizer \(\hat e_n\) satisfying
\[
\frac1n\sum_{i=1}^n \phi'(X_i-\hat e_n)=1.
\]
Under strong convexity, smoothness, and a differentiation-under-expectation condition, the paper proves \(O(1/n)\) mean-squared error for \(\hat e_n\), sub-Gaussian concentration
\[
P\!\left(|\hat e_n-e^*|\ge \epsilon\right)
\le 2\exp\!\left(-\frac{n\mu^2\epsilon^2}{8L^2\sigma^2}\right),
\]
and sub-exponential concentration for \(\oce_n^\phi\). It also studies the streaming recursion
\[
t_j=t_{j-1}-\gamma_j(1-\phi'(X_j-t_{j-1})),
\]
with Polyak–Ruppert averaging, obtaining \(O(1/m)\) mean-squared error for the averaged optimizer and \(O(m^{-1/2})\) expected absolute error for the induced OCE estimate [2405.20933].

For multivariate OCE, [2210.13825] proposes projected Robbins–Monro updates
\[
m_{n+1}=\Pi_K[m_n+\gamma_n H_1(X_{n+1},m_n)],
\qquad
H_1(X,m)=\nabla l(-X-m)-1,
\]
proves almost sure convergence \(m_n\to m^*\), and derives a central limit theorem for averaged iterates. The same paper develops a companion recursion for the risk value itself and emphasizes that stochastic approximation provides asymptotic error quantification and confidence intervals, in contrast to direct Monte Carlo or deterministic minimization [2210.13825].

On the statistical learning side, [2006.08138] studies empirical OCE minimization over a hypothesis class and proves Rademacher-complexity generalization bounds. Its basic OCE is
\[
\oce^\phi(f;P)=\inf_{\lambda\in\mathbb R}
\left\{\lambda+\mathbb E_P[\phi(f(Z)-\lambda)]\right\},
\]
and the paper shows uniform convergence of \(\oce_n\) to \(\oce\), excess-OCE bounds for the empirical OCE minimizer, and a variance-based characterization
\[
C_\phi \sigma^2(f)\le \oce(f)-R(f)\le \frac{\Lip(\phi)}{2}\sigma(f),
\]
which yields expected-loss guarantees with weaker dependence on \(\Lip(\phi)\) than a direct Lipschitz analysis [2006.08138].

## 5. Dynamic control and reinforcement learning

A central distinction in dynamic settings is between static OCE, which is often time-inconsistent, and recursive OCE, which is designed to preserve dynamic programming structure. In controlled diffusions with terminal loss \(f(Y_T)\), [2001.10108] treats the static problem
\[
\inf_{a\in\mathcal A} p(f(Y_T^{0,y,a})),
\qquad
p(X)=\inf_{r\in\mathbb R}(E[l(X-r)]+r),
\]
and shows that generic OCE criteria are time-inconsistent. The main remedy is state augmentation via the dual density variable \(z\): the enlarged value
\[
V(t,y,z):=\inf_{r\in\mathbb R,\ a\in\mathcal A}
\left(E[l(f(Y_T^{t,y,a})-r)]+rz\right)
\]
satisfies a dynamic programming principle and is characterized as the viscosity solution of a singular Hamilton–Jacobi–Bellman–Isaacs equation. Under additional assumptions, the solution is unique. The paper identifies entropic risk as essentially the only time-consistent example under its assumptions, and treats AVaR/CVaR as a primary application of the general enlargement method [2001.10108].

By contrast, episodic tabular RL with recursive OCE starts from Bellman recursions that already insert OCE locally:
\[
Q_h^\pi(s,a)=r_h(s,a)+OCE^u_{s'\sim P_h(\cdot\mid s,a)}(V_{h+1}^\pi(s')),\qquad
V_h^\pi(s)=Q_h^\pi(s,\pi_h(s)).
\]
This recursive OCE formulation yields dynamically consistent value functions and supports optimistic value iteration. The UCB-style algorithm of [2301.12601] proves regret upper bounds and a minimax lower bound, with OCE-specific dependence on \(|u(-H+h)|\) and derivatives of \(u\). The paper emphasizes that the framework unifies recursive entropic risk, iterated CVaR, and recursive mean-variance in a single episodic RL formulation [2301.12601].

Sample-complexity analysis in discounted MDPs sharpens the structural picture. The paper on discounted recursive OCE proves that, aside from degenerate utilities that reduce to expectation, the PAC-learnable OCE objectives are exactly those with full-domain utility functions \(u\), namely \(\mathrm{dom}(u)=\mathbb R\). When \(\mathrm{dom}(u)\neq\mathbb R\), value and policy learning are not PAC-learnable in general. For \(\mathrm{CVaR}_\tau\), the paper derives a lower bound showing the correct dependence on the tail level is \(1/\tau^2\) [2605.21763]. This suggests that learnability in recursive OCE RL is governed as much by the domain geometry of the utility as by standard MDP parameters.

Recent RL extensions use OCE beyond ordinary Bellman recursions. In constrained discounted RL, [2510.20199] applies reward-side OCEs to the occupancy measure,
\[
\rho(Z)=\sup_t\{t+\mathbb E[g(Z-t)]\},
\]
and shows that, for fixed \(t\), the inner control problem is an ordinary expected-reward problem with transformed per-stage reward \(t+g(r-t)\). Under Slater-type conditions, the paper establishes parameterized strong Lagrangian duality and proposes an outer stochastic gradient descent-ascent scheme that can wrap standard solvers such as PPO [2510.20199]. In continuous time, [2512.02386] shows that when the objective functional is an OCE, the optimal policy is Markovian on an augmented state \((X,B_0,B_1)\), derives an HJB characterization, and proposes the martingale-based CT-RS-q algorithm. Its outer OCE step is the one-dimensional optimization
\[
J_0^*(t,x)=\max_{b\in\mathbb R}\{b+J^*(t,x,-b,1)\},
\]
which mirrors the scalar certainty-equivalent structure of the static theory [2512.02386].

## 6. Applications, variants, and recurring limitations

OCE now serves as a modeling language across several application domains. In precision medicine, the optimized covariate-dependent equivalent replaces a global scalar shift by a patient-specific function \(\alpha(X)\), and under decomposability the criterion becomes an expected conditional OCE,
\[
\mathcal O^{\,d}_{(u,\mathcal F)}(R)=
\mathbb E[\mathcal O_u(R\mid X,A=d(X))].
\]
This reframes individualized treatment learning as conditional OCE maximization and includes expected-value rules, CVaR-sensitive rules, and conditional mean-variance rules as special cases [1908.10742].

Conformal prediction and calibration work extends OCE into finite-sample uncertainty quantification. Conformal risk training defines an OCE risk
\[
R[X]=\inf_{t\in\mathbb R}\{t+\mathbb E[\phi(X-t)]\},
\]
and observes that for fixed \(t\), conformal calibration can act on the transformed loss \(t+\phi(L-t)\). This yields distribution-free finite-sample control of OCE risks, including CVaR, and enables end-to-end differentiation through the conformal calibration layer during model training [2510.08748]. A related prediction-set paper defines
\[
R_{\text{OCE}(\lambda)}=\inf_{t\in\mathbb R}
\left\{t+\mathbb E\!\left[\phi(\ell(y,\Gamma_\lambda(x))-t)\right]\right\},
\]
uses upper confidence bounds on the fixed-\(t\) surrogate, and proves
\[
\Pr\!\big[R_{\text{OCE}(\hat\lambda)}\le \alpha\big]\ge 1-\delta
\]
for bounded monotone losses such as miscoverage and false negative rate. In its experiments, OCE-RCPS consistently meets target satisfaction rates across CVaR and entropic-risk configurations, whereas OCE-CRC controls risk only on average over calibration datasets [2602.13660].

The literature also contains genuine variants rather than straightforward extensions. Modified OCE and robust modified OCE replace the present cash term by its utility,
\[
M_u(\xi)=\sup_{x\in\mathbb R}\{u(x)+\mathbb E_P[u(\xi-x)]\},
\qquad
R(\xi)=\sup_{x\in\mathbb R}\inf_{u\in\mathcal U}\{u(x)+\mathbb E_P[u(\xi-x)]\},
\]
motivated by the claim that the classical OCE objective mixes cash and utility units. The same paper proves law invariance, monotonicity, risk aversion when \(u(t)\le t\), concavity, and positive subhomogeneity, while noting that the modified criterion departs from classical cash-additive OCE structure [2203.10762].

Several limitations recur across this body of work. Static OCE control problems are often time-inconsistent unless one uses recursive formulations or state augmentation [2001.10108]. Sharp statistical guarantees for estimation commonly require strong convexity, smoothness, and sub-Gaussian sampling, and they do not directly cover nonsmooth special cases such as standard CVaR [2405.20933]. In recursive RL, utilities without full domain are structurally problematic for PAC learning [2605.21763]. Conformal OCE methods need bounded monotone losses and often require a separate optimization set for tuning the auxiliary parameter \(t\), which reduces effective calibration sample size and can make the resulting procedures conservative [2510.08748][2602.13660].

Taken together, these results suggest that OCE is best understood not as a single risk measure but as a flexible operational template: a scalar certainty-equivalent optimization that can be specialized, dualized, conditioned, robustified, vectorized, recursively embedded, or statistically estimated, with the precise behavior determined by the generator \(u\), \(\phi\), or \(l\) and by the dynamic or informational structure imposed around it.

Source: https://www.emergentmind.com/topics/optimized-certainty-equivalents-oce