---
title: Path-Space ELBO in Diffusion Modeling
url: https://www.emergentmind.com/topics/path-space-evidence-lower-bound-elbo
type: topic
---

# Path-Space ELBO in Diffusion Modeling

The path-space Evidence Lower Bound (ELBO) is a variational lower bound on the marginal likelihood formulated not in terms of finite-dimensional random variables, but as a functional on measures over entire continuous-time trajectories—i.e., paths of a stochastic process. Path-space ELBOs arise naturally in the context of diffusion-based generative modeling, stochastic optimal control, and, more recently, in posterior inference for models such as deep Gaussian processes that utilize stochastic differential equations (SDEs) to define priors and generative mechanisms [2211.01364] [2605.23434].

## 1. Forward and Reverse SDEs: Path-Space Laws

Diffusion-based generative models characteristically define a forward SDE (the inference SDE) on a time interval $[0, T]$ in $\mathbb{R}^d$ via
$$
dY_t = f(Y_t, t)\,dt + \sigma(t)\,dB_t, \qquad Y_0 \sim \mathcal{D},
$$
where $f$ is the data-driven drift, $\sigma$ is the volatility coefficient (typically independent of state), and $B_t$ is standard Brownian motion. The law of the path $(Y_t)_{0 \leq t \leq T}$, denoted $P_Y$, lives on the function space $C([0,T], \mathbb{R}^d)$.

The generative process is defined by reversing time in the SDE. The resulting process, $X_t \equiv Y_{T-t}$, satisfies
$$
dX_t = \mu(X_t, t)\,dt + \sigma(t)\,d\widetilde{B}_t, \qquad X_0 \sim Y_T,
$$
with reversed drift
$$
\mu(x, t) = \sigma(t)\sigma(t)^\top \nabla \log p_Y(x, T-t) - f(x, T-t),
$$
governed by the time-reversal of diffusion (Nelson–Haussmann–Föllmer theory). As the exact marginal score $\nabla \log p_Y$ is generally intractable, practical implementations substitute a (learned) approximation [2211.01364].

## 2. The Hamilton–Jacobi–Bellman Equation and Stochastic Control

The log-density of the time-reversed (generative) process admits a nonlinear PDE characterizing its evolution. Setting $q(x, t) = p_Y(x, T-t)$, one obtains $V(x,t) = -\log q(x,t)$, which evolves by the Hamilton–Jacobi–Bellman (HJB) equation:
$$
\partial_t V = -\mathrm{Tr}(D(t)\nabla^2 V) + \mu \cdot \nabla V - \mathrm{div}\,\mu + \frac{1}{2}\| \sigma(t)^\top \nabla V \|^2,
$$
where $D(t) = \frac{1}{2}\sigma(t)\sigma(t)^\top$ and the terminal condition is $V(x,T) = -\log p_{X_0}(x)$, with $p_{X_0}$ the prior density [2211.01364]. The HJB equation directly links stochastic optimal control and diffusion-based generative modeling via the value functional over trajectories.

## 3. Verification Theorem: Derivation of the Path-Space ELBO

Given a family of admissible controls $u \in \mathcal{U} \subset C^1(\mathbb{R}^d \times [0,T], \mathbb{R}^d)$, define the controlled forward SDE:
$$
dY_t^u = [\sigma(t) u(Y_t^u, t) - \mu(Y_t^u, t)]dt + \sigma(t) dB_t, \qquad Y_0^u \sim \mathcal{D}.
$$
Define the running cost as
$$
\mathcal{R}^u_\mu(Y^u) = \int_0^T \left[ \mathrm{div}\,\mu(Y^u_s,s) + \frac{1}{2}\|u(Y^u_s,s)\|^2 \right] ds,
$$
and the terminal cost
$$
C(Y_T^u) = -\log p_{X_0}(Y_T^u).
$$
The verification theorem shows that the value function at the process start is minimized over controls by the expected sum of these costs, with the minimizer given by a feedback control involving the log-density gradient. For any control $u$,
$$
\log p_{X_T}(Y^u_0) = -V(Y^u_0, 0) \geq \mathbb{E}\left[\log p_{X_0}(Y^u_T) - \mathcal{R}^u_\mu(Y^u) \mid Y_0^u\right].
$$
The right-hand side is the path-space ELBO, representing a variational lower bound on the log marginal density at the endpoint [2211.01364].

## 4. Path-Space Kullback–Leibler Divergence and the ELBO Decomposition

Let $P_{Y^u}$ and $P_{Y^{u^*}}$ denote the path-space laws under a general and optimal control, respectively. Girsanov’s theorem gives
$$
\log \frac{dP_{Y^u}}{dP_{Y^{u^*}}}(Y^u) = \mathcal{R}^u_\mu(Y^u) + \int_0^T u_s \cdot dB_s + \log \frac{p_{X_T}(Y^u_0)}{p_{X_0}(Y^u_T)},
$$
and taking expectations yields
$$
D_{\mathrm{KL}}(P_{Y^u} \Vert P_{Y^{u^*}}) = \mathbb{E}[\mathcal{R}^u_\mu(Y^u)] - \mathbb{E}[\log p_{X_0}(Y^u_T)] + \mathbb{E}[\log p_{X_T}(Y^u_0)].
$$
Upon rearrangement, the path-space ELBO is
$$
\mathrm{ELBO}(u) := \mathbb{E}\left[\log p_{X_0}(Y^u_T) - \mathcal{R}^u_\mu(Y^u)\right] = \mathbb{E}\left[\log p_{X_T}(Y^u_0)\right] - D_{\mathrm{KL}}(P_{Y^u} \Vert P_{Y^{u^*}}).
$$
Maximizing the path-space ELBO is thus equivalent to minimizing the path-space KL divergence to the time-reversed (optimal) diffusion law [2211.01364].

## 5. Path-Space ELBOs in Diffusion-Based Generative Modeling and Deep Gaussian Processes

In diffusion probabilistic models, the path-space ELBO describes a fully continuous-time analogue of VAE bounds, with the bound defined on trajectory spaces. The running cost $\mathcal{R}^u_\mu$ penalizes deviation of a learned drift from the true marginal score (the pathwise optimal control), and the terminal cost employs the known generator prior density. Optimizing the ELBO by simulating sample paths and backpropagating through the running-cost integral recovers common losses (score-matching, denoising) used in diffusion models. The path-space ELBO also inherits the mode-seeking/mode-covering dichotomy from the directionality of the KL divergence. This framework unifies score-based diffusion models, VAEs, and Schrödinger-bridge methods under stochastic optimal control theory [2211.01364].

For deep Gaussian processes, recent formulations leverage Doob-bridged reference diffusions and Onsager–Machlup action regularization to derive strict path-space ELBOs and alternative objectives. For example, the "FFJORD log-det" and "OM-regularised CNF" ELBOs utilize probability-flow ODE variants and Onsager–Machlup actions as path priors. These formulations link the ELBO to negative log unnormalized path densities and, through the small-noise Freidlin–Wentzell LDP, to MAP path estimators under the formal tempered Doob-bridge posterior [2605.23434].

| Model Type                                | Path-Space ELBO           | Notable Features                                     |
|--------------------------------------------|---------------------------|------------------------------------------------------|
| Diffusion-based generative models [2211.01364] | $\mathbb{E}[\log p_{X_0}(Y_T^u) - \mathcal{R}_\mu^u(Y^u)]$ | Unifies score models, VAEs, and Schrödinger bridges   |
| Deep GP inference [2605.23434]                 | FFJORD/OM-regularised CNF | Deterministic samplers, Onsager–Machlup path regularization |

## 6. Assumptions, Implementation, and Empirical Perspectives

The path-space ELBO framework assumes Lipschitz continuity of drift and diffusion coefficients to ensure unique strong solutions and strictly positive densities in the corresponding Fokker–Planck PDEs. The volatility matrix $\sigma(t)$ is often chosen independent of $x$ (state) to simplify the Kolmogorov backward generator. The control function class is restricted to ensure Novikov’s condition and the applicability of the verification theorem. Boundary conditions require knowledge of the prior density at the trajectory endpoint [2211.01364].

In practical implementations, both stochastic and deterministic sampler paradigms are used. For instance, OM-Path (a deterministic sampler ODE) applies Onsager–Machlup regularization using bridge-marginal coefficients, and results empirically demonstrate domain-dependent advantages: on larger datasets with lower noise, pathwise regularized objectives outperform strict density-matching ELBOs, while stochastic formulations remain competitive in small-sample/high-noise settings [2605.23434].

## 7. Theoretical Insights and Unifying Principles

The path-space perspective on the ELBO reveals deep connections among generative modeling, stochastic control, and variational inference. Notably:
- The ELBO on path space is the negative of a KL divergence between trajectory measures.
- Formulations such as FFJORD log-det or OM-regularised ELBOs connect instantaneous change-of-variables with path-prior regularization.
- The Onsager–Machlup functional acts as the rate functional for the small-noise Freidlin–Wentzell LDP of the reference diffusion, making OM-regularized training amortized MAP estimation of the true path posterior rather than a variational density fit.
- The overall framework unifies Doob SDE anchoring, probability-flow ODEs, classical control theory, and Schrödinger bridge formulations.

A plausible implication is that path-space ELBOs provide a rigorous foundation to extend variational methods beyond densities over finite-dimensional variables, enabling principled learning and inference in diffusion-based and bridge-based generative models across domains [2211.01364][2605.23434].

Source: https://www.emergentmind.com/topics/path-space-evidence-lower-bound-elbo