Papers
Topics
Authors
Recent
Search
2000 character limit reached

Mean-Field Hamilton-Jacobi-Bellman Equation

Updated 27 November 2025
  • Mean-field HJB equations are partial differential equations characterizing the value function for stochastic control problems where dynamics depend on the state distribution.
  • They employ a dynamic programming principle on the Wasserstein space using Lions differentiability to tackle infinite-dimensional, non-local features.
  • Explicit linear-quadratic solutions demonstrate practical applications in finance and systemic risk, reducing the problem to coupled Riccati equations for optimal feedback controls.

A mean-field Hamilton-Jacobi-Bellman (HJB) equation is a partial differential equation characterizing the value function for stochastic optimal control problems in which the controlled system evolves according to McKean–Vlasov (mean-field) dynamics, i.e., where the drift, diffusion, or cost may depend on the distribution (law) of the state, and possibly also the control. This construct extends classical HJB theory to infinite-dimensional state spaces, typically the Wasserstein space of probability measures endowed with Lions differentiability. Mean-field HJB equations form the core of dynamic programming approaches for mean-field control and optimization, underpinning contemporary research in mean-field stochastic control, mean-field games, and large-population Markov decision processes.

1. Formulation of the Mean-Field Stochastic Control Problem

In the McKean–Vlasov framework, a controlled process (Xt)t[0,T](X_t)_{t\in[0,T]} follows the stochastic differential equation

dXt=b(t,Xt,PXt,αt)dt+σ(t,Xt,PXt,αt)dWt,dX_t = b(t, X_t, \mathbb{P}_{X_t}, \alpha_t)\,dt + \sigma(t, X_t, \mathbb{P}_{X_t}, \alpha_t)\,dW_t,

where bb and σ\sigma are Lipschitz functions in (x,μ,α)(x, \mu, \alpha), α\alpha is an admissible (progressively measurable, square-integrable) control with values in a compact set ARm\mathcal{A}\subset\mathbb{R}^m, and PXt\mathbb{P}_{X_t} is the law of XtX_t. The cost functional is

J(α)=E[0Tf(t,Xt,αt,P(Xt,αt))dt+g(XT,PXT)],J(\alpha) = \mathbb{E}\bigg[ \int_0^T f(t, X_t, \alpha_t, \mathbb{P}_{(X_t, \alpha_t)})\,dt + g(X_T, \mathbb{P}_{X_T}) \bigg],

with suitable growth and regularity constraints on dXt=b(t,Xt,PXt,αt)dt+σ(t,Xt,PXt,αt)dWt,dX_t = b(t, X_t, \mathbb{P}_{X_t}, \alpha_t)\,dt + \sigma(t, X_t, \mathbb{P}_{X_t}, \alpha_t)\,dW_t,0 and dXt=b(t,Xt,PXt,αt)dt+σ(t,Xt,PXt,αt)dWt,dX_t = b(t, X_t, \mathbb{P}_{X_t}, \alpha_t)\,dt + \sigma(t, X_t, \mathbb{P}_{X_t}, \alpha_t)\,dW_t,1. The goal is to minimize dXt=b(t,Xt,PXt,αt)dt+σ(t,Xt,PXt,αt)dWt,dX_t = b(t, X_t, \mathbb{P}_{X_t}, \alpha_t)\,dt + \sigma(t, X_t, \mathbb{P}_{X_t}, \alpha_t)\,dW_t,2 over admissible controls dXt=b(t,Xt,PXt,αt)dt+σ(t,Xt,PXt,αt)dWt,dX_t = b(t, X_t, \mathbb{P}_{X_t}, \alpha_t)\,dt + \sigma(t, X_t, \mathbb{P}_{X_t}, \alpha_t)\,dW_t,3.

For feedback controls of the form dXt=b(t,Xt,PXt,αt)dt+σ(t,Xt,PXt,αt)dWt,dX_t = b(t, X_t, \mathbb{P}_{X_t}, \alpha_t)\,dt + \sigma(t, X_t, \mathbb{P}_{X_t}, \alpha_t)\,dW_t,4 with dXt=b(t,Xt,PXt,αt)dt+σ(t,Xt,PXt,αt)dWt,dX_t = b(t, X_t, \mathbb{P}_{X_t}, \alpha_t)\,dt + \sigma(t, X_t, \mathbb{P}_{X_t}, \alpha_t)\,dW_t,5 Lipschitz, the flow of marginals dXt=b(t,Xt,PXt,αt)dt+σ(t,Xt,PXt,αt)dWt,dX_t = b(t, X_t, \mathbb{P}_{X_t}, \alpha_t)\,dt + \sigma(t, X_t, \mathbb{P}_{X_t}, \alpha_t)\,dW_t,6 evolves deterministically in dXt=b(t,Xt,PXt,αt)dt+σ(t,Xt,PXt,αt)dWt,dX_t = b(t, X_t, \mathbb{P}_{X_t}, \alpha_t)\,dt + \sigma(t, X_t, \mathbb{P}_{X_t}, \alpha_t)\,dW_t,7, the space of probability measures with finite second moment. The value function thus becomes dXt=b(t,Xt,PXt,αt)dt+σ(t,Xt,PXt,αt)dWt,dX_t = b(t, X_t, \mathbb{P}_{X_t}, \alpha_t)\,dt + \sigma(t, X_t, \mathbb{P}_{X_t}, \alpha_t)\,dW_t,8, the minimum cost starting from time dXt=b(t,Xt,PXt,αt)dt+σ(t,Xt,PXt,αt)dWt,dX_t = b(t, X_t, \mathbb{P}_{X_t}, \alpha_t)\,dt + \sigma(t, X_t, \mathbb{P}_{X_t}, \alpha_t)\,dW_t,9 and marginal law bb0 (Pham et al., 2015).

2. Dynamic Programming Principle and Bellman Equation on Wasserstein Space

Under this reformulation, the dynamic programming principle (DPP) holds in the space of probability measures: bb1 where bb2 and bb3 are the mean-field extensions of the running and terminal cost: bb4 The DPP takes a recursive form on bb5.

Key to the analytic machinery is the notion of differentiability with respect to probability measures, as formalized by Lions. For a function bb6, the "lift" to bb7 is defined by bb8. If bb9 is Fréchet-differentiable, the Lions derivative σ\sigma0 exists and forms the infinitesimal generator for Itô's calculus on the Wasserstein space (Pham et al., 2015).

Applying the extended Itô formula yields: σ\sigma1 Plugging this chain rule, together with the DPP, leads to the mean-field HJB equation.

3. Mean-Field Hamilton-Jacobi-Bellman Equation: Structure and Interpretation

The resulting Bellman PDE on σ\sigma2 takes the form: σ\sigma3

σ\sigma4

The Hamiltonian is given explicitly in terms of drift, diffusion, and cost, with infimum over Markov controls σ\sigma5. The appearance of Lions derivatives reflects the infinite-dimensional geometry of σ\sigma6, making the equation genuinely non-local and nonlinear in distributional argument (Pham et al., 2015).

4. Solution Concepts: Classical, Verification, and Viscosity

Existence and uniqueness of solutions depend on regularity:

  • Classical solution and Verification: If σ\sigma7 solves the Bellman equation and the infimum is achieved by a Lipschitz feedback σ\sigma8, then σ\sigma9 and the associated closed-loop control is optimal. The proof proceeds by applying Itô's formula to (x,μ,α)(x, \mu, \alpha)0 and leveraging the PDE to dominate the cost functional, with equality achieved on the feedback minimizer (Pham et al., 2015).
  • Viscosity solutions: If smoothness fails, the equation is lifted to (x,μ,α)(x, \mu, \alpha)1, and a notion of viscosity solution is constructed via test functions that themselves are lifts from (x,μ,α)(x, \mu, \alpha)2. The value function (x,μ,α)(x, \mu, \alpha)3 is shown to satisfy the viscosity solution conditions, and a comparison principle holds for sub/supersolutions with suitable growth controls, implying uniqueness.

These results ensure the well-posedness of the mean-field Bellman equation in wide generality, even in the presence of measure dependence and degenerate diffusion.

5. Linear-Quadratic Explicit Solutions and Applications

In the linear-quadratic (LQ) case, (x,μ,α)(x, \mu, \alpha)4 and (x,μ,α)(x, \mu, \alpha)5 are affine in (x,μ,α)(x, \mu, \alpha)6 and (x,μ,α)(x, \mu, \alpha)7, and the costs are quadratic: (x,μ,α)(x, \mu, \alpha)8 where (x,μ,α)(x, \mu, \alpha)9, α\alpha0. The value function admits a closed form as a quadratic function of α\alpha1: α\alpha2 with α\alpha3 evolving according to coupled Riccati and linear ODEs.

These explicit solutions undergird applications including:

  • Mean-variance portfolio selection, recovering classical optimal investment formulas.
  • Systemic risk in inter-bank models, where the optimal borrowing-lending rates are obtained as affine functions of deviation from the population mean (Pham et al., 2015).

6. Open-Loop vs. Closed-Loop Controls and Equivalence

Under open-loop controls, where policies need not be feedback in α\alpha4, the DPP remains valid on α\alpha5 with a modified Hamiltonian. However, under mild integrability conditions, one can show that the infima coincide for open- and closed-loop formulations, so that the HJB equation and optimal values coincide in both settings. This equivalence is particularly robust in the LQ case (Pham et al., 2015).

7. Relation to Mean-Field Game Theory and Extensions

The mean-field HJB equation forms the optimality condition in mean-field type control, mean-field games, and certain large-system Markov decision frameworks. It connects to the mean-field game master equation, which describes the limit of Nash equilibria for large populations and dualizes with Fokker–Planck equations for the state law, as in the foundational theory developed by Lions and subsequent works (Pham et al., 2015, Bensoussan et al., 2014, Gast et al., 2010).

Extensions include:

  • Master equations coupling the value function with the law and individual state (mean-field “social optimization” and necessary conditions for α\alpha6-person-by-person optimality) (Huang et al., 19 Aug 2025).
  • Infinite-dimensional PDEs arising in storage models or delayed systems, formulated in Hilbert or Banach spaces and linked analytically to the mean-field HJB structure (Bertucci et al., 2022, Fouque et al., 2018).
  • Weak/viscosity solutions in spaces of measures (e.g., via Fourier mode truncation) accommodating highly singular or non-convex data (Cecchin et al., 2022).

The rigorous study of mean-field HJB equations continues to drive both mathematical theory and applications in stochastic control, finance, systemic risk, and large-scale engineered or physical systems.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Mean-Field Hamilton-Jacobi-Bellman Equation.