Papers
Topics
Authors
Recent
Search
2000 character limit reached

Recursive Perturbed Utility (RPU)

Updated 18 March 2026
  • Recursive Perturbed Utility (RPU) is a stochastic differential utility framework that integrates entropy rewards into dynamic portfolio choice to model randomized investment strategies.
  • It generalizes the classical Merton problem by deriving optimal Gaussian portfolio policies with explicit mean and variance expressions that capture both myopic and hedging components.
  • RPU embeds entropy into the discount rate, dynamically reducing the marginal benefit of randomization to ensure the well-posedness of the optimization and mirror observed noisy investment behavior.

Recursive Perturbed Utility (RPU) is a stochastic differential utility framework that models investor preferences for randomization in dynamic portfolio choice, incorporating entropy-based rewards into a recursive aggregator. Developed to address the ill-posedness of additive perturbed utility models in dynamic, continuous-time settings, RPU offers a mathematically tractable and economically interpretable approach to balancing utility from portfolio randomization with traditional objectives such as bequest. RPU generalizes the classical Merton problem by allowing optimal portfolio policies to be inherently stochastic—Gaussian with explicitly characterized mean and variance—thereby unifying rational dynamic investment with empirically observed randomized behavior (Dai et al., 14 Feb 2026).

1. Market Model and Randomized Controls

The RPU framework operates in a standard continuous-time Markovian incomplete market comprising a risk-free asset (rate rr) and a single risky asset. The risky asset price StS_t evolves as

dStSt=μ(t,Xt) dt+σ(t,Xt) dBt\frac{dS_t}{S_t} = \mu(t,X_t)\,dt + \sigma(t,X_t)\,dB_t

where XtX_t is a Markovian state factor following

dXt=m(t,Xt) dt+ν(t,Xt)[ρ dBt+1−ρ2 dB~t]dX_t = m(t,X_t)\,dt + \nu(t,X_t)\bigl[\rho\,dB_t + \sqrt{1-\rho^2}\,d\widetilde{B}_t\bigr]

with independent Brownian motions BtB_t, B~t\widetilde{B}_t and correlation ρ\rho.

At each time tt, the investor's portfolio allocation is not deterministic but is specified by a probability density πt(⋅)\pi_t(\cdot) over StS_t0. The mean and variance of this allocation are denoted

StS_t1

The resulting wealth process under this "exploratory" control evolves as

StS_t2

where StS_t3 is independent of StS_t4.

2. Recursive Aggregator and Entropy Reward

Crucial to RPU is the entropy functional of the portfolio density,

StS_t5

For a temperature function StS_t6, RPU defines a recursive utility process StS_t7—a backward stochastic differential equation (BSDE)—given by

StS_t8

where StS_t9 is CRRA utility.

The recursion can be explicitly expressed as

dStSt=μ(t,Xt) dt+σ(t,Xt) dBt\frac{dS_t}{S_t} = \mu(t,X_t)\,dt + \sigma(t,X_t)\,dB_t0

The entropy reward is thereby discounted endogenously: the marginal utility of additional randomization decreases as cumulative past entropy increases (for dStSt=μ(t,Xt) dt+σ(t,Xt) dBt\frac{dS_t}{S_t} = \mu(t,X_t)\,dt + \sigma(t,X_t)\,dB_t1), preventing runaway entropy.

3. Dynamic Programming and Characterization via HJB Equation

The optimal value function under feedback policy dStSt=μ(t,Xt) dt+σ(t,Xt) dBt\frac{dS_t}{S_t} = \mu(t,X_t)\,dt + \sigma(t,X_t)\,dB_t2 is

dStSt=μ(t,Xt) dt+σ(t,Xt) dBt\frac{dS_t}{S_t} = \mu(t,X_t)\,dt + \sigma(t,X_t)\,dB_t3

The associated dynamic programming principle yields a Hamilton–Jacobi–Bellman (HJB) PDE: dStSt=μ(t,Xt) dt+σ(t,Xt) dBt\frac{dS_t}{S_t} = \mu(t,X_t)\,dt + \sigma(t,X_t)\,dB_t4 with dStSt=μ(t,Xt) dt+σ(t,Xt) dBt\frac{dS_t}{S_t} = \mu(t,X_t)\,dt + \sigma(t,X_t)\,dB_t5.

4. Structure of the RPU-Optimal Portfolio Policy

Maximizing entropy over densities with given mean and variance yields a Gaussian optimizer, i.e.,

dStSt=μ(t,Xt) dt+σ(t,Xt) dBt\frac{dS_t}{S_t} = \mu(t,X_t)\,dt + \sigma(t,X_t)\,dB_t6

Assuming the value function ansatz dStSt=μ(t,Xt) dt+σ(t,Xt) dBt\frac{dS_t}{S_t} = \mu(t,X_t)\,dt + \sigma(t,X_t)\,dB_t7, the optimal control is

dStSt=μ(t,Xt) dt+σ(t,Xt) dBt\frac{dS_t}{S_t} = \mu(t,X_t)\,dt + \sigma(t,X_t)\,dB_t8

with

dStSt=μ(t,Xt) dt+σ(t,Xt) dBt\frac{dS_t}{S_t} = \mu(t,X_t)\,dt + \sigma(t,X_t)\,dB_t9

The mean comprises a myopic Merton term and an intertemporal hedging correction, the latter captured by XtX_t0, which solves a nonlinear PDE: XtX_t1 with XtX_t2.

5. Asymptotic Expansion and Deviation from Classical Merton Policy

For small entropy weight XtX_t3, the solution admits an asymptotic expansion: XtX_t4 where XtX_t5 solves the classical Merton PDE. The mean policy expands as

XtX_t6

The forgone expected utility relative to the non-randomized (XtX_t7) policy is of order XtX_t8. The equivalent relative wealth loss,

XtX_t9

quantifies the small but nonzero financial cost of enjoying randomization.

6. Economic Interpretation and Well-Posedness

Additive perturbed utility formulations (e.g., dXt=m(t,Xt) dt+ν(t,Xt)[ρ dBt+1−ρ2 dB~t]dX_t = m(t,X_t)\,dt + \nu(t,X_t)\bigl[\rho\,dB_t + \sqrt{1-\rho^2}\,d\widetilde{B}_t\bigr]0) may be ill-posed for dXt=m(t,Xt) dt+ν(t,Xt)[ρ dBt+1−ρ2 dB~t]dX_t = m(t,X_t)\,dt + \nu(t,X_t)\bigl[\rho\,dB_t + \sqrt{1-\rho^2}\,d\widetilde{B}_t\bigr]1 or bounded-below utilities, allowing entropy to diverge. RPU overcomes this by embedding entropy into the discount rate, producing endogenous depreciation of the entropy reward as cumulative randomization accrues. This construction prevents explosion of entropy and ensures the well-posedness of the optimization.

The RPU model parallels Uzawa-style habit formation: the flow "reward" of entropy, tied to local randomization, is dynamically interdependent with the terminal wealth objective. The optimal portfolio’s variance dXt=m(t,Xt) dt+ν(t,Xt)[ρ dBt+1−ρ2 dB~t]dX_t = m(t,X_t)\,dt + \nu(t,X_t)\bigl[\rho\,dB_t + \sqrt{1-\rho^2}\,d\widetilde{B}_t\bigr]2 decreases with risk aversion dXt=m(t,Xt) dt+ν(t,Xt)[ρ dBt+1−ρ2 dB~t]dX_t = m(t,X_t)\,dt + \nu(t,X_t)\bigl[\rho\,dB_t + \sqrt{1-\rho^2}\,d\widetilde{B}_t\bigr]3 and stock volatility dXt=m(t,Xt) dt+ν(t,Xt)[ρ dBt+1−ρ2 dB~t]dX_t = m(t,X_t)\,dt + \nu(t,X_t)\bigl[\rho\,dB_t + \sqrt{1-\rho^2}\,d\widetilde{B}_t\bigr]4, while the mean is distorted from the Merton ratio by an order-dXt=m(t,Xt) dt+ν(t,Xt)[ρ dBt+1−ρ2 dB~t]dX_t = m(t,X_t)\,dt + \nu(t,X_t)\bigl[\rho\,dB_t + \sqrt{1-\rho^2}\,d\widetilde{B}_t\bigr]5 entropic hedging adjustment.

A plausible implication is that RPU offers a micro-founded and analytically tractable justification for stochasticity in policy selection, with explicit quantification of the cost of randomization, supporting empirical observations of noisy investment behavior.

7. Connections to Broader Literature and Applications

Recursive Perturbed Utility extends the additive perturbed utility theory of Fudenberg et al. (2015) for static decisions by establishing a well-posed dynamic counterpart. RPU admits explicit portfolio policy characterization in general Markovian incomplete markets with CRRA preferences and allows for closed-form expressions for both mean and variance of optimal policies. It provides a theoretical basis for observed stochastic choice in dynamic portfolio allocation, integrating entropy-regularized decision-making with foundational stochastic control.

Potential applications encompass asset management, behavioral finance, and stochastic control where preference for diversification or exploration is inherent. The RPU formalism is compatible with extensions involving state-dependent randomness preferences and more general stochastic environments (Dai et al., 14 Feb 2026).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Recursive Perturbed Utility (RPU).