Papers
Topics
Authors
Recent
Search
2000 character limit reached

Underdamped Langevin Diffusion

Updated 10 July 2026
  • Underdamped Langevin Diffusion is a phase-space process that updates both configuration and momentum, incorporating inertia, friction, and thermal noise.
  • It leverages a Gibbs stationary measure and hypoelliptic characteristics to achieve accelerated convergence and efficient Markov chain Monte Carlo sampling.
  • The framework extends to diverse applications, from anomalous transport in heterogeneous media to kinetic reaction-diffusion and momentum‐based stochastic optimization.

Underdamped Langevin diffusion is a phase-space diffusion that evolves both a configuration variable and a momentum or velocity variable, thereby retaining inertia, friction, and thermal noise. In its standard sampling form,

dxt=ptdt,dpt=f(xt)dtγptdt+2γdBt,dx_t = p_t\,dt,\qquad dp_t = -\nabla f(x_t)\,dt - \gamma p_t\,dt + \sqrt{2\gamma}\,dB_t,

or, more generally,

dxt=vtdt,dvt=γvtdtuf(xt)dt+2γudBt,dx_t = v_t\,dt,\qquad dv_t = -\gamma v_t\,dt - u\nabla f(x_t)\,dt + \sqrt{2\gamma u}\,dB_t,

and it admits the Gibbs stationary law

Π(x,p)exp ⁣(f(x)12p2),\Pi(x,p)\propto \exp\!\left(-f(x)-\frac{1}{2}\|p\|^2\right),

whose xx-marginal is the target distribution π(x)ef(x)\pi(x)\propto e^{-f(x)} (Zuo et al., 2024, Cheng et al., 2017, Foster et al., 2021). Because noise enters only through momentum, the process is hypoelliptic rather than elliptic, and this structural feature governs its convergence theory, transport asymptotics, thermodynamic properties, and numerical discretizations (Lee et al., 16 Mar 2025).

1. Phase-space formulation and invariant structure

The canonical physical form of the underdamped Langevin equation writes

dx=pmdt,dp=γpmdt+F(x,t)dt+dBt,d\boldsymbol{x} = \frac{\boldsymbol{p}}{m}\,dt,\qquad d\boldsymbol{p} = -\gamma \frac{\boldsymbol{p}}{m}\,dt + \boldsymbol{F}(\boldsymbol{x},t)\,dt + d\boldsymbol{B}_t,

with Brownian forcing of covariance 2Ddt2D\,dt, D=γTD=\gamma T, and corresponding Kramers dynamics for the phase-space density f(x,p,t)f(x,p,t) (Lyu et al., 2024). In the sampling literature, the same structure is commonly rewritten with ff as the negative log target density and with unit mass, yielding a Hamiltonian

dxt=vtdt,dvt=γvtdtuf(xt)dt+2γudBt,dx_t = v_t\,dt,\qquad dv_t = -\gamma v_t\,dt - u\nabla f(x_t)\,dt + \sqrt{2\gamma u}\,dB_t,0

and a compact drift-diffusion representation in dxt=vtdt,dvt=γvtdtuf(xt)dt+2γudBt,dx_t = v_t\,dt,\qquad dv_t = -\gamma v_t\,dt - u\nabla f(x_t)\,dt + \sqrt{2\gamma u}\,dB_t,1 variables (Zuo et al., 2024). The invariant measure is then Gibbsian on phase space and Gaussian in the momentum coordinate, a fact used both in MCMC analysis and in nonequilibrium thermodynamics (Zuo et al., 2024, Paraguassú et al., 2021).

The explicit momentum variable is the essential distinction from overdamped dynamics. It introduces inertial transport through dxt=vtdt,dvt=γvtdtuf(xt)dt+2γudBt,dx_t = v_t\,dt,\qquad dv_t = -\gamma v_t\,dt - u\nabla f(x_t)\,dt + \sqrt{2\gamma u}\,dB_t,2, suppresses purely diffusive random-walk motion, and places underdamped Langevin methods in close conceptual proximity to Hamiltonian Monte Carlo; one non-asymptotic analysis states that the underdamped Langevin MCMC scheme “can be viewed as a version of Hamiltonian Monte Carlo (HMC)” (Cheng et al., 2017). A plausible implication is that phase-space geometry, rather than position-space drift alone, is the correct organizing principle for understanding acceleration, metastability, and response in these systems.

2. Equilibration, transport coefficients, and singular limits

In generalized Langevin settings with colored noise and auxiliary variables, the same underdamped degeneracy persists in a stronger form: dissipation is degenerate and standard coercive arguments fail. For the quasi-Markovian GLE in a periodic potential, sharp longtime equilibration estimates are obtained by hypocoercivity, with exponential decay in a twisted Sobolev norm equivalent to a standard Sobolev norm:

dxt=vtdt,dvt=γvtdtuf(xt)dt+2γudBt,dx_t = v_t\,dt,\qquad dv_t = -\gamma v_t\,dt - u\nabla f(x_t)\,dt + \sqrt{2\gamma u}\,dB_t,3

for mean-zero dxt=vtdt,dvt=γvtdtuf(xt)dt+2γudBt,dx_t = v_t\,dt,\qquad dv_t = -\gamma v_t\,dt - u\nabla f(x_t)\,dt + \sqrt{2\gamma u}\,dB_t,4 (Pavliotis et al., 2020). The same work formulates effective diffusion through the Poisson problem

dxt=vtdt,dvt=γvtdtuf(xt)dt+2γudBt,dx_t = v_t\,dt,\qquad dv_t = -\gamma v_t\,dt - u\nabla f(x_t)\,dt + \sqrt{2\gamma u}\,dB_t,5

which is the basic homogenization object for long-time transport in periodic media (Pavliotis et al., 2020).

In the underdamped limit of that GLE, corresponding to dxt=vtdt,dvt=γvtdtuf(xt)dt+2γudBt,dx_t = v_t\,dt,\qquad dv_t = -\gamma v_t\,dt - u\nabla f(x_t)\,dt + \sqrt{2\gamma u}\,dB_t,6, the effective diffusion coefficient obeys

dxt=vtdt,dvt=γvtdtuf(xt)dt+2γudBt,dx_t = v_t\,dt,\qquad dv_t = -\gamma v_t\,dt - u\nabla f(x_t)\,dt + \sqrt{2\gamma u}\,dB_t,7

so diffusion diverges as friction vanishes, while the prefactor retains explicit dependence on the memory kernel and the periodic potential (Pavliotis et al., 2020). In the white-noise limit dxt=vtdt,dvt=γvtdtuf(xt)dt+2γudBt,dx_t = v_t\,dt,\qquad dv_t = -\gamma v_t\,dt - u\nabla f(x_t)\,dt + \sqrt{2\gamma u}\,dB_t,8, the prefactor converges to the classical underdamped Langevin value dxt=vtdt,dvt=γvtdtuf(xt)dt+2γudBt,dx_t = v_t\,dt,\qquad dv_t = -\gamma v_t\,dt - u\nabla f(x_t)\,dt + \sqrt{2\gamma u}\,dB_t,9, and numerical results further report

Π(x,p)exp ⁣(f(x)12p2),\Pi(x,p)\propto \exp\!\left(-f(x)-\frac{1}{2}\|p\|^2\right),0

confirming the theoretical expansion (Pavliotis et al., 2020). This suggests a robust separation between universal scaling in friction and non-universal dependence on temporal correlations of the noise.

3. Discretization, MCMC, and algorithmic variants

For strongly log-concave targets with smooth gradients, underdamped Langevin diffusion admits a sharper non-asymptotic MCMC theory than overdamped Langevin. One foundational result proves that a discretization of the underdamped process achieves Π(x,p)exp ⁣(f(x)12p2),\Pi(x,p)\propto \exp\!\left(-f(x)-\frac{1}{2}\|p\|^2\right),1-Wasserstein error Π(x,p)exp ⁣(f(x)12p2),\Pi(x,p)\propto \exp\!\left(-f(x)-\frac{1}{2}\|p\|^2\right),2 in

Π(x,p)exp ⁣(f(x)12p2),\Pi(x,p)\propto \exp\!\left(-f(x)-\frac{1}{2}\|p\|^2\right),3

steps, compared with

Π(x,p)exp ⁣(f(x)12p2),\Pi(x,p)\propto \exp\!\left(-f(x)-\frac{1}{2}\|p\|^2\right),4

for overdamped Langevin under the same smoothness and strong convexity assumptions (Cheng et al., 2017). The analysis is based on contraction of the continuous-time semigroup,

Π(x,p)exp ⁣(f(x)12p2),\Pi(x,p)\propto \exp\!\left(-f(x)-\frac{1}{2}\|p\|^2\right),5

followed by explicit control of discretization error (Cheng et al., 2017).

A parallel line of work studies geometric splitting methods for the discrete chain itself. The OBABO integrator is a symmetric Strang splitting for the kinetic Langevin SDE and requires only one gradient evaluation per iteration (Monmarché, 2020). For convex and smooth targets, the discrete-time chain enjoys dimension-free Wasserstein contraction and explicit total variation regularization, and the unadjusted chain has Wasserstein efficiency bounds of order Π(x,p)exp ⁣(f(x)12p2),\Pi(x,p)\propto \exp\!\left(-f(x)-\frac{1}{2}\|p\|^2\right),6 in the general case, Π(x,p)exp ⁣(f(x)12p2),\Pi(x,p)\propto \exp\!\left(-f(x)-\frac{1}{2}\|p\|^2\right),7 if the Hessian is Lipschitz, and Π(x,p)exp ⁣(f(x)12p2),\Pi(x,p)\propto \exp\!\left(-f(x)-\frac{1}{2}\|p\|^2\right),8 for separable targets (Monmarché, 2020). The same paper establishes non-asymptotic bounds for a Metropolis-adjusted version as well (Monmarché, 2020).

Higher-order gradient-only schemes exploit additional smoothness of the target. The shifted ODE method matches the leading stochastic Taylor terms of underdamped Langevin, uses only gradients, and becomes practical through a third-order Runge–Kutta implementation (SORT) or a fourth-order splitting implementation (SOFA); the reported per-step overhead for SORT is two additional gradient evaluations, and logistic-regression experiments show faster convergence than competing unadjusted Langevin algorithms (Foster et al., 2021). More recently, QUICSORT is presented as the first gradient-only ULD method with third-order convergence, with step complexities

Π(x,p)exp ⁣(f(x)12p2),\Pi(x,p)\propto \exp\!\left(-f(x)-\frac{1}{2}\|p\|^2\right),9

under, respectively, Lipschitz gradient, Lipschitz Hessian, and Lipschitz third derivative assumptions (Scott et al., 22 Aug 2025).

Several variants modify the basic dynamics while preserving the same invariant law. Gradient-adjusted underdamped Langevin dynamics add a term xx0 to the xx1 equation together with matching noise, retain the Gibbs stationary distribution, and for Gaussian targets yield discrete mixing time toward a biased target with square-root dependence on the condition number of the covariance matrix (Zuo et al., 2024). Skew-symmetric perturbations in both position and momentum likewise preserve the invariant Gibbs measure and can reduce asymptotic variance, especially for Gaussian targets, provided the perturbations are properly aligned with the observable and covariance structure (Duncan et al., 2017). Underdamped transitions have also been incorporated into variational inference: Langevin Diffusion Variational Inference numerically simulates the underdamped forward diffusion and its time reversal, augments backward transitions with a score network, and reports consistently stronger ELBOs than ULA, MCD, and UHA across logistic-regression and time-series tasks (Geffner et al., 2022).

4. Anomalous transport and heterogeneous environments

Underdamped Langevin diffusion has also served as a vehicle for anomalous transport in heterogeneous media. In an isothermal one-dimensional medium with spatially dependent diffusion coefficient xx2, analytical diffusion-equation solutions and numerical simulations of the underdamped Langevin equation agree at long times for xx3, yielding subdiffusion for xx4, normal diffusion for xx5, and superdiffusion for xx6 (Regev et al., 2016). The same study shows that for xx7 the diffusion equation predicts faster-than-ballistic spreading, but this is unphysical: in the full underdamped dynamics, motion cannot exceed the ballistic limit xx8, because particles eventually coast inertially once friction becomes negligible (Regev et al., 2016). A recurrent misconception is therefore that position-only diffusion equations remain valid when spatially varying friction destroys the time-scale separation needed for overdamped closure.

A different failure of the overdamped description appears in underdamped scaled Brownian motion, where temperature, damping, and diffusion coefficients depend explicitly on time,

xx9

For this model, the overdamped limit can fail to describe the long-time behavior, and for small π(x)ef(x)\pi(x)\propto e^{-f(x)}0 or the ultraslow case π(x)ef(x)\pi(x)\propto e^{-f(x)}1, inertial corrections remain significant over practical observation windows (Bodrova et al., 2016). The paper connects this behavior to free cooling granular gases and emphasizes persistent inertia in both ensemble and time-averaged statistics (Bodrova et al., 2016).

Non-Markovian timing mechanisms generate further anomalies. A weakly damped Langevin system coupled to an π(x)ef(x)\pi(x)\propto e^{-f(x)}2-dependent subordinator with π(x)ef(x)\pi(x)\propto e^{-f(x)}3 leads to a fractional Klein–Kramers equation, and its mean-squared displacement depends crucially on the two-point distribution of the inverse subordinator (Wang et al., 2018). For π(x)ef(x)\pi(x)\propto e^{-f(x)}4, the model yields super-ballistic diffusion,

π(x)ef(x)\pi(x)\propto e^{-f(x)}5

whereas for π(x)ef(x)\pi(x)\propto e^{-f(x)}6 it gives sub-ballistic superdiffusion,

π(x)ef(x)\pi(x)\propto e^{-f(x)}7

together with an ergodicity-breaking parameter

π(x)ef(x)\pi(x)\propto e^{-f(x)}8

These features motivate the description “Lévy-walk-like” (Wang et al., 2018).

Randomization of medium parameters produces another class of anomalies. In an underdamped Langevin system with time-dependent diffusivity π(x)ef(x)\pi(x)\propto e^{-f(x)}9 and a random relaxation timescale dx=pmdt,dp=γpmdt+F(x,t)dt+dBt,d\boldsymbol{x} = \frac{\boldsymbol{p}}{m}\,dt,\qquad d\boldsymbol{p} = -\gamma \frac{\boldsymbol{p}}{m}\,dt + \boldsymbol{F}(\boldsymbol{x},t)\,dt + d\boldsymbol{B}_t,0, a heavy-tailed distribution of dx=pmdt,dp=γpmdt+F(x,t)dt+dBt,d\boldsymbol{x} = \frac{\boldsymbol{p}}{m}\,dt,\qquad d\boldsymbol{p} = -\gamma \frac{\boldsymbol{p}}{m}\,dt + \boldsymbol{F}(\boldsymbol{x},t)\,dt + d\boldsymbol{B}_t,1 with finite mean suppresses the decay of the velocity correlation function and promotes diffusion (Chen et al., 2021). The ensemble MSD scales as dx=pmdt,dp=γpmdt+F(x,t)dt+dBt,d\boldsymbol{x} = \frac{\boldsymbol{p}}{m}\,dt,\qquad d\boldsymbol{p} = -\gamma \frac{\boldsymbol{p}}{m}\,dt + \boldsymbol{F}(\boldsymbol{x},t)\,dt + d\boldsymbol{B}_t,2 for fixed dx=pmdt,dp=γpmdt+F(x,t)dt+dBt,d\boldsymbol{x} = \frac{\boldsymbol{p}}{m}\,dt,\qquad d\boldsymbol{p} = -\gamma \frac{\boldsymbol{p}}{m}\,dt + \boldsymbol{F}(\boldsymbol{x},t)\,dt + d\boldsymbol{B}_t,3 and dx=pmdt,dp=γpmdt+F(x,t)dt+dBt,d\boldsymbol{x} = \frac{\boldsymbol{p}}{m}\,dt,\qquad d\boldsymbol{p} = -\gamma \frac{\boldsymbol{p}}{m}\,dt + \boldsymbol{F}(\boldsymbol{x},t)\,dt + d\boldsymbol{B}_t,4, as dx=pmdt,dp=γpmdt+F(x,t)dt+dBt,d\boldsymbol{x} = \frac{\boldsymbol{p}}{m}\,dt,\qquad d\boldsymbol{p} = -\gamma \frac{\boldsymbol{p}}{m}\,dt + \boldsymbol{F}(\boldsymbol{x},t)\,dt + d\boldsymbol{B}_t,5 after averaging over heavy-tailed dx=pmdt,dp=γpmdt+F(x,t)dt+dBt,d\boldsymbol{x} = \frac{\boldsymbol{p}}{m}\,dt,\qquad d\boldsymbol{p} = -\gamma \frac{\boldsymbol{p}}{m}\,dt + \boldsymbol{F}(\boldsymbol{x},t)\,dt + d\boldsymbol{B}_t,6, and as dx=pmdt,dp=γpmdt+F(x,t)dt+dBt,d\boldsymbol{x} = \frac{\boldsymbol{p}}{m}\,dt,\qquad d\boldsymbol{p} = -\gamma \frac{\boldsymbol{p}}{m}\,dt + \boldsymbol{F}(\boldsymbol{x},t)\,dt + d\boldsymbol{B}_t,7 in the special random-initialization regime dx=pmdt,dp=γpmdt+F(x,t)dt+dBt,d\boldsymbol{x} = \frac{\boldsymbol{p}}{m}\,dt,\qquad d\boldsymbol{p} = -\gamma \frac{\boldsymbol{p}}{m}\,dt + \boldsymbol{F}(\boldsymbol{x},t)\,dt + d\boldsymbol{B}_t,8 (Chen et al., 2021). The same model exhibits aging and broken ergodicity in the time-averaged MSD (Chen et al., 2021). This suggests that underdamped diffusion is particularly effective at disentangling inertia, quenched heterogeneity, and temporal nonstationarity.

5. Heat, entropy production, and periodically driven states

In stochastic thermodynamics, underdamped Langevin diffusion differs from the overdamped theory because heat includes explicit kinetic-energy boundary terms. For the free particle, linear potential, and harmonic potential in arbitrary dimension, the heat over dx=pmdt,dp=γpmdt+F(x,t)dt+dBt,d\boldsymbol{x} = \frac{\boldsymbol{p}}{m}\,dt,\qquad d\boldsymbol{p} = -\gamma \frac{\boldsymbol{p}}{m}\,dt + \boldsymbol{F}(\boldsymbol{x},t)\,dt + d\boldsymbol{B}_t,9 can be written as

2Ddt2D\,dt0

and exact characteristic functions and probability distributions are available in each case (Paraguassú et al., 2021). The free-particle heat law is Besselian, the free and harmonic cases are symmetric in 2Ddt2D\,dt1, and the linear potential produces an asymmetric distribution because the external field shifts the mean heat (Paraguassú et al., 2021). The same analysis emphasizes leptokurtic tails and “heat leakage” from kinetic fluctuations that are absent in overdamped reductions (Paraguassú et al., 2021).

Precision bounds for fluctuating currents must also be modified in the finite-mass regime. For steady-state underdamped Langevin dynamics, the overdamped thermodynamic uncertainty relation fails in general, and the relevant bound becomes

2Ddt2D\,dt2

where 2Ddt2D\,dt3 contains the overdamped entropy-production term together with two additional positive contributions associated with local mean acceleration and velocity fluctuations, plus a transient term 2Ddt2D\,dt4 (Dechant, 2022). The theory extends to magnetic fields, anisotropic temperature, and joint bounds for several observables, and in periodic potentials the correlation TUR can make the bound tight when correlations with auxiliary observables are incorporated (Dechant, 2022).

A separate modified TUR addresses entropy-production inference directly. For underdamped Langevin systems, one may define a cumulant current 2Ddt2D\,dt5 and a stochastic current 2Ddt2D\,dt6 so that

2Ddt2D\,dt7

The bound is saturable when 2Ddt2D\,dt8 is proportional to the irreversible current 2Ddt2D\,dt9, and the method requires only the damping-to-mass ratio D=γTD=\gamma T0 and the diffusion constant D=γTD=\gamma T1, rather than full knowledge of the force field or stationary density (Lyu et al., 2024). The same work states that the construction applies to both overdamped and underdamped Langevin dynamics, for arbitrary time-dependent protocols and nonsteady scenarios (Lyu et al., 2024).

Periodic driving produces explicitly time-dependent nonequilibrium asymptotic states. For a harmonic underdamped Langevin system with D=γTD=\gamma T2-periodic friction D=γTD=\gamma T3, stiffness D=γTD=\gamma T4, and diffusion D=γTD=\gamma T5, the asymptotic distribution is Gaussian with D=γTD=\gamma T6-periodic covariance matrix D=γTD=\gamma T7,

D=γTD=\gamma T8

and the full nonperturbative dynamics can be reduced via an D=γTD=\gamma T9 symmetry to a Hill equation with periodic coefficients (Awasthi et al., 2019). This framework yields exact or near-exact asymptotic distributions, two-time correlation functions, and linear responses to drift and diffusion perturbations (Awasthi et al., 2019).

6. Metastability, applications, and kinetic mean-field extensions

At low temperature, underdamped Langevin diffusion exhibits metastable transitions between basins of attraction. For a double-well Morse potential with metastable minimum f(x,p,t)f(x,p,t)0, global minimum f(x,p,t)f(x,p,t)1, and saddle f(x,p,t)f(x,p,t)2, the sharp Eyring–Kramers asymptotic for the mean transition time from f(x,p,t)f(x,p,t)3 to a neighborhood of f(x,p,t)f(x,p,t)4 is

f(x,p,t)f(x,p,t)5

with

f(x,p,t)f(x,p,t)6

The derivation circumvents traditional potential-theoretic tools, which are ill-defined for this hypoelliptic process (Lee et al., 16 Mar 2025).

Outside equilibrium theory, underdamped simulations have been used as direct computational tools for diffusion in layered or discontinuous media. In a one-dimensional two-layer model of a drug-eluting stent, the motion is simulated with the underdamped Langevin equation

f(x,p,t)f(x,p,t)7

using the G-JF integrator, the inertial convention for spatially varying friction, and a fictitious-mass device to enlarge the ballistic time without altering the long-time diffusive solution (Regev et al., 2018). A semi-permeable membrane at the interface is implemented by an exponential waiting time

f(x,p,t)f(x,p,t)8

and the reported finding is that the membrane has virtually no delay effect on the overall rate of delivery for realistic parameter values (Regev et al., 2018).

Underdamped Langevin dynamics also underlies kinetic reaction-diffusion limits. In particle-based reactive Langevin dynamics, single-particle motion is governed by

f(x,p,t)f(x,p,t)9

while reactions are driven by stochastic kernels and placement densities on phase space (Isaacson et al., 2 Jun 2026). In the large-population limit, the empirical phase-space measures converge to deterministic nonlocal kinetic reaction-diffusion partial integro-differential equations with hypoelliptic transport operators and reaction terms that retain both spatial and velocity structure (Isaacson et al., 2 Jun 2026). Relative to overdamped PBSRD models, the limiting equations are kinetic rather than purely configurational, and they preserve momentum-level information absent from position-only closures (Isaacson et al., 2 Jun 2026).

A final connection runs toward stochastic optimization. For stochastic gradient descent with momentum, quantitative error estimates to the underdamped Langevin diffusion are available in both ff0-Wasserstein and total variation distance (Guillin et al., 2024). Under the stated smoothness, moment, and learning-rate assumptions, the discrepancies scale as ff1 in ff2 and ff3 in total variation, with an additional ff4 term reflecting minibatch noise (Guillin et al., 2024). This suggests that underdamped Langevin diffusion functions not only as a sampling process and a physical transport model, but also as a mathematically controlled continuum surrogate for momentum-based stochastic optimization.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (20)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Underdamped Langevin Diffusion.