Papers
Topics
Authors
Recent
Search
2000 character limit reached

Lagrangian Map Distillation

Updated 9 January 2026
  • Lagrangian Map Distillation is a computational framework for learning deterministic mappings between probability measures using least-action principles.
  • It employs neural and spline amortization to bypass costly ODE solvers, enabling fast, one-shot inference in optimal transport and generative modeling.
  • The framework integrates dual, c-transform, and path-energy losses to achieve scalable, geometry-aware transport with state-of-the-art empirical performance.

Lagrangian Map Distillation (LMD) refers to a set of computational frameworks for learning deterministic, one-shot mappings and associated paths between probability measures or samples, subject to least-action or Lagrangian transport costs. LMD principles unlock efficient inference in optimal transport (OT) and generative flow models, notably by amortizing path-solving and optimization steps that would traditionally require expensive ODE solvers or iterative procedures at test time. Building on advances in neural parameterizations and amortized optimization, LMD methods are applicable to a variety of transport settings, including those with complex dynamics or geometric constraints, and furnish state-of-the-art generative models with reduced inference costs (Pooladian et al., 2024, Sabour et al., 17 Jun 2025).

1. Lagrangian Costs and Least-Action Formulations

Lagrangian costs generalize classical OT by incorporating the geometry, dynamics, and constraints of the underlying system via an action functional. For points x,y∈Rdx,y\in\mathbb{R}^d, the Lagrangian cost is

c(x,y)=inf⁡γ:γ(0)=x, γ(1)=y∫01L(γ(t),γ˙(t))dt,c(x, y) = \inf_{\gamma: \gamma(0) = x,\, \gamma(1) = y} \int_0^1 L(\gamma(t), \dot{\gamma}(t)) dt,

where the Lagrangian L(x,v)L(x, v) is typically chosen as one of:

  • Kinetic cost (“Benamou–Brenier”): L(x,v)=12∥v∥2L(x, v) = \frac{1}{2}\|v\|^2.
  • Kinetic + Potential (obstacles/barriers): L(x,v)=12∥v∥2−U(x)L(x, v) = \frac{1}{2}\|v\|^2 - U(x), with UU penalizing undesirable regions.
  • Position-dependent Riemannian metrics: L(x,v)=12v⊤A(x)vL(x, v) = \frac{1}{2} v^\top A(x) v, A(x)∈S++dA(x) \in S_{++}^d encoding non-Euclidean geometry.

The associated action functional J[γ]=∫01L(γ(t),γ˙(t))dtJ[\gamma] = \int_0^1 L(\gamma(t), \dot{\gamma}(t)) dt defines a geodesic with unique minimizer γ∗\gamma^* under mild regularity conditions, simultaneously yielding c(x,y)=inf⁡γ:γ(0)=x, γ(1)=y∫01L(γ(t),γ˙(t))dt,c(x, y) = \inf_{\gamma: \gamma(0) = x,\, \gamma(1) = y} \int_0^1 L(\gamma(t), \dot{\gamma}(t)) dt,0 and the optimal path (Pooladian et al., 2024).

2. LMD for Optimal Transport: Neural and Spline Amortization

The LMD recipe for OT underlies recent “neural optimal transport with Lagrangian costs” approaches. The goal is to distill in a single training stage:

  • A Kantorovich potential c(x,y)=inf⁡γ:γ(0)=x, γ(1)=y∫01L(γ(t),γ˙(t))dt,c(x, y) = \inf_{\gamma: \gamma(0) = x,\, \gamma(1) = y} \int_0^1 L(\gamma(t), \dot{\gamma}(t)) dt,1, parameterized by an MLP, for dual optimization.
  • An amortized map predictor c(x,y)=inf⁡γ:γ(0)=x, γ(1)=y∫01L(γ(t),γ˙(t))dt,c(x, y) = \inf_{\gamma: \gamma(0) = x,\, \gamma(1) = y} \int_0^1 L(\gamma(t), \dot{\gamma}(t)) dt,2, also an MLP, to approximate c(x,y)=inf⁡γ:γ(0)=x, γ(1)=y∫01L(γ(t),γ˙(t))dt,c(x, y) = \inf_{\gamma: \gamma(0) = x,\, \gamma(1) = y} \int_0^1 L(\gamma(t), \dot{\gamma}(t)) dt,3.
  • A spline path predictor c(x,y)=inf⁡γ:γ(0)=x, γ(1)=y∫01L(γ(t),γ˙(t))dt,c(x, y) = \inf_{\gamma: \gamma(0) = x,\, \gamma(1) = y} \int_0^1 L(\gamma(t), \dot{\gamma}(t)) dt,4, amortizing the coefficients for cubic spline representations of geodesic paths.

Training alternates three batch Monte Carlo losses:

  1. Dual objective (c(x,y)=inf⁡γ:γ(0)=x, γ(1)=y∫01L(γ(t),γ˙(t))dt,c(x, y) = \inf_{\gamma: \gamma(0) = x,\, \gamma(1) = y} \int_0^1 L(\gamma(t), \dot{\gamma}(t)) dt,5):

c(x,y)=inf⁡γ:γ(0)=x, γ(1)=y∫01L(γ(t),γ˙(t))dt,c(x, y) = \inf_{\gamma: \gamma(0) = x,\, \gamma(1) = y} \int_0^1 L(\gamma(t), \dot{\gamma}(t)) dt,6

where c(x,y)=inf⁡γ:γ(0)=x, γ(1)=y∫01L(γ(t),γ˙(t))dt,c(x, y) = \inf_{\gamma: \gamma(0) = x,\, \gamma(1) = y} \int_0^1 L(\gamma(t), \dot{\gamma}(t)) dt,7, with c(x,y)=inf⁡γ:γ(0)=x, γ(1)=y∫01L(γ(t),γ˙(t))dt,c(x, y) = \inf_{\gamma: \gamma(0) = x,\, \gamma(1) = y} \int_0^1 L(\gamma(t), \dot{\gamma}(t)) dt,8 obtained by warm-started L-BFGS from c(x,y)=inf⁡γ:γ(0)=x, γ(1)=y∫01L(γ(t),γ˙(t))dt,c(x, y) = \inf_{\gamma: \gamma(0) = x,\, \gamma(1) = y} \int_0^1 L(\gamma(t), \dot{\gamma}(t)) dt,9.

  1. L(x,v)L(x, v)0-transform amortization (L(x,v)L(x, v)1):

L(x,v)L(x, v)2

  1. Path-energy amortization (L(x,v)L(x, v)3):

L(x,v)L(x, v)4

with L(x,v)L(x, v)5, and L(x,v)L(x, v)6 given by the amortized spline.

Optimization proceeds by stochastic, highly vectorized updates (typically via Adam) of L(x,v)L(x, v)7. L-BFGS is used only during training to refine L(x,v)L(x, v)8 for the dual objective and quickly amortized out (Pooladian et al., 2024).

3. LMD in Continuous-Time Flow Map Distillation

In generative modeling, LMD appears within the “Align Your Flow” (AYF) framework as a variant of map distillation from score-based or flow-matching diffusion models. Continuous-time flow maps are neural networks L(x,v)L(x, v)9 that predict L(x,v)=12∥v∥2L(x, v) = \frac{1}{2}\|v\|^20 conditioned on L(x,v)=12∥v∥2L(x, v) = \frac{1}{2}\|v\|^21 and L(x,v)=12∥v∥2L(x, v) = \frac{1}{2}\|v\|^22, such that L(x,v)=12∥v∥2L(x, v) = \frac{1}{2}\|v\|^23 and L(x,v)=12∥v∥2L(x, v) = \frac{1}{2}\|v\|^24. The parametrization is

L(x,v)=12∥v∥2L(x, v) = \frac{1}{2}\|v\|^25

where L(x,v)=12∥v∥2L(x, v) = \frac{1}{2}\|v\|^26 approximates the score or velocity field (Sabour et al., 17 Jun 2025).

The LMD loss in AYF is

L(x,v)=12∥v∥2L(x, v) = \frac{1}{2}\|v\|^27

with L(x,v)=12∥v∥2L(x, v) = \frac{1}{2}\|v\|^28 and the ODE operator propagating the teacher’s output backwards from L(x,v)=12∥v∥2L(x, v) = \frac{1}{2}\|v\|^29 to L(x,v)=12∥v∥2−U(x)L(x, v) = \frac{1}{2}\|v\|^2 - U(x)0. In the L(x,v)=12∥v∥2−U(x)L(x, v) = \frac{1}{2}\|v\|^2 - U(x)1 limit, this enforces end-point consistency with the velocity field, serving as a Lagrangian constraint.

The training loop involves:

  • Sampling L(x,v)=12∥v∥2−U(x)L(x, v) = \frac{1}{2}\|v\|^2 - U(x)2 from data, L(x,v)=12∥v∥2−U(x)L(x, v) = \frac{1}{2}\|v\|^2 - U(x)3 from noise, forming L(x,v)=12∥v∥2−U(x)L(x, v) = \frac{1}{2}\|v\|^2 - U(x)4 for L(x,v)=12∥v∥2−U(x)L(x, v) = \frac{1}{2}\|v\|^2 - U(x)5.
  • Constructing autoguided teacher velocities L(x,v)=12∥v∥2−U(x)L(x, v) = \frac{1}{2}\|v\|^2 - U(x)6 as convex combinations of the main and weak teachers.
  • Computing the student prediction L(x,v)=12∥v∥2−U(x)L(x, v) = \frac{1}{2}\|v\|^2 - U(x)7 and the target L(x,v)=12∥v∥2−U(x)L(x, v) = \frac{1}{2}\|v\|^2 - U(x)8 as above.
  • Applying consistency and linearity regularization during training.
  • Using AdamW optimizers with large-batch GPU vectorization (Sabour et al., 17 Jun 2025).

4. Computational Advantages and Inference Procedures

LMD achieves fast, single-forward-pass inference of deterministic transport maps or generative flows without requiring inner optimization or ODE solves at test time:

  • For OT (static map): L(x,v)=12∥v∥2−U(x)L(x, v) = \frac{1}{2}\|v\|^2 - U(x)9 via MLP, paths via UU0.
  • For flow distillation: UU1 gives UU2 in one evaluation for arbitrary UU3.

Amortized neural predictors eliminate expensive iterative procedures at inference. In OT, spline basis evaluation yields fast geodesic reconstruction. In flow models, batch inference is supported by compact UNet/MLP architectures (Pooladian et al., 2024, Sabour et al., 17 Jun 2025).

The following table summarizes key computational elements:

Setting Main Predictors Runtime Inference
OT (Pooladian et al., 2024) UU4, UU5 One MLP + spline eval
AYF (Sabour et al., 17 Jun 2025) UU6 One UNet/MLP pass

This approach is especially advantageous in high-throughput or large-scale applications, as demonstrated in low-dimensional OT examples and in state-of-the-art generative modeling on ImageNet (Pooladian et al., 2024, Sabour et al., 17 Jun 2025).

5. Architectural, Regularization, and Training Considerations

The neural components are designed for efficiency:

  • MLP parameterizations for UU7 and UU8 use four layers of 64 leaky ReLU units; splines via a UU9 MLP.
  • Backpropagation through splines is employed for updating L(x,v)=12v⊤A(x)vL(x, v) = \frac{1}{2} v^\top A(x) v0.
  • Loss alternation is utilized for joint optimization: maximizing the dual, regressing to amortized L(x,v)=12v⊤A(x)vL(x, v) = \frac{1}{2} v^\top A(x) v1-transforms, and minimizing path energies.
  • Warm-started L-BFGS (with backtracking line search) is used during training, not inference, to stabilize dual optimization.

In AYF, tangent warm-up, autoguidance (convex combinations of strong/weak teachers), and time embedding stabilization are critical for stable training and state-of-the-art sample quality. Adversarial fine-tuning with a StyleGAN2-based discriminator further improves sample FID without significantly harming recall (Sabour et al., 17 Jun 2025).

6. Theoretical Guarantees and Empirical Performance

Theoretical analysis in the generative context establishes that flow maps generalized through LMD maintain stability as the number of function evaluations (NFE) varies, in contrast to standard consistency models where error compounds as NFE increases. In OT, the LMD approach guarantees recovery of action-minimizing maps under standard convexity and smoothness assumptions of the Lagrangian.

Empirical benchmarks show:

  • OT with general Lagrangians: Efficient geodesic computation and map extraction even for non-Euclidean or constrained systems, outperforming classical approaches in scalability and flexibility (Pooladian et al., 2024).
  • ImageNet 64×64, 512×512 (AYF): 2–4 step FID scores set new state-of-the-art for few-step image generation. For example, AYF (no GAN) achieves FID 1.25 at 2 steps, 1.15 at 4 steps; adversarial tuning further reduces FID, with only a minor decrease in recall (Sabour et al., 17 Jun 2025).
  • Text-to-image: AYF-variant showed user preference of ≈60% against LoRA-CM baselines in a 47-person study.

A practical implication is that LMD-based methods provide fast, scalable, geometry- or constraint-aware mappings and generations, efficiently amortizing complex computations while maintaining generation quality even with very few neural evaluations.

7. Practical Guidelines and Recommendations

For OT tasks with general Lagrangian costs:

  • Use cubic spline path parameterizations and MLP amortization for both maps and paths.
  • Train with batched dual, c-transform, and path energy losses, alternating updates.

For generative flows:

  • AYF recommends Eulerian Map Distillation (EMD) for stability over LMD losses.
  • Adopt teacher autoguidance with L(x,v)=12v⊤A(x)vL(x, v) = \frac{1}{2} v^\top A(x) v2.
  • Use L(x,v)=12v⊤A(x)vL(x, v) = \frac{1}{2} v^\top A(x) v3 parameterization; time embeddings via MLP.
  • Apply tangent warm-up (linearity ramping) and tangent normalization.
  • Adversarial fine-tuning is optional but effective for maximizing sample quality at very low NFE.

These procedures are supported by publicly released implementations for reproducibility (Pooladian et al., 2024). The frameworks are equally applicable to low-dimensional OT (with general geometric constraints) as well as high-dimensional generative modeling benchmarks (Pooladian et al., 2024, Sabour et al., 17 Jun 2025).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (2)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Lagrangian Map Distillation (LMD).