Papers
Topics
Authors
Recent
Search
2000 character limit reached

Marginal Flow: Optimal Transport & Density Estimation

Updated 14 July 2026
  • Marginal Flow is a family of mathematical methods that use marginal distributions as constraints to guide changes in couplings, with applications in optimal transport and density estimation.
  • In optimal transport, neural-ODE updates adjust couplings via marginal-preserving interior flows, ensuring endpoint distributions remain unchanged while reducing cost.
  • Extensions include continuous-time formulations, parameter marginalization, and applications in field theory, offering robust insights into distribution evolution across varied domains.

“Marginal Flow” is not a single standardized object but a family of constructions in which the central mathematical role is played by marginal distributions: preserving them, matching them at multiple times, transporting between them, or parameterizing densities through marginalization. In optimal transport, the term denotes a cost-specific, marginal-preserving neural-ODE update that moves inside the feasible set of couplings (Liu, 2022). In density estimation, it denotes a model of the form

qθ(x)=q(xw)qθ(w)dw,q_\theta(x)=\int q(x\mid w)\,q_\theta(w)\,dw,

where latent parameters are marginalized rather than optimized as fixed mixture components (Negri et al., 30 Sep 2025). Related usages appear in multi-marginal Schrödinger bridges, continuum-marginal optimal transport, bridge-aware samplers, and several field-theoretic settings involving marginal or marginally irrelevant operators (Chen et al., 2019, Nakano, 27 Apr 2026, Park, 2021). This suggests that the term is best understood through recurring structural motifs rather than a single canonical definition.

1. Marginals as constraints, invariants, and dynamical observables

A recurring formal idea is that a flow acts on couplings or path measures while the observable constraints are marginal laws. In the optimal-transport usage of “Marginal Flow,” one starts from a differentiable stochastic process {Xt}t[0,1]\{X_t\}_{t\in[0,1]} with velocity field

vtX(z)=E[X˙tXt=z],v_t^{\mathbf X}(z)=\mathbb{E}[\dot X_t \mid X_t=z],

whose marginals $\rho_t=\law(X_t)$ satisfy

tρt+ ⁣ ⁣(vtXρt)=0.\partial_t \rho_t + \nabla\!\cdot\!\big(v_t^{\mathbf X}\rho_t\big)=0.

A vector field rtr_t is X\mathbf X-marginal-preserving if

0tE[h(Xs)rs(Xs)]ds=0hCc,\int_0^t \mathbb{E}\big[\nabla h(X_s)\cdot r_s(X_s)\big]\,ds = 0 \quad \forall h\in \mathcal C_c,

equivalently,

(rtρt)=0.\nabla\cdot\big(r_t\rho_t\big)=0.

The key theorem states that two processes with the same initial law have identical marginals for all tt iff the difference of their velocity fields is marginal-preserving (Liu, 2022).

In continuum-marginal optimal transport, marginals are not merely endpoint constraints but a full time-continuous input {Xt}t[0,1]\{X_t\}_{t\in[0,1]}0. The admissible velocity field {Xt}t[0,1]\{X_t\}_{t\in[0,1]}1 must reproduce every marginal through the weak continuity equation

{Xt}t[0,1]\{X_t\}_{t\in[0,1]}2

and the optimization problem is

{Xt}t[0,1]\{X_t\}_{t\in[0,1]}3

Here the marginal flow itself is the object being recovered (Nakano, 27 Apr 2026).

A related continuum viewpoint appears in the Sinkhorn/IPFP limit. As {Xt}t[0,1]\{X_t\}_{t\in[0,1]}4 and {Xt}t[0,1]\{X_t\}_{t\in[0,1]}5, the one-marginals produced by the discrete entropic OT iterations converge to an absolutely continuous curve {Xt}t[0,1]\{X_t\}_{t\in[0,1]}6 in {Xt}t[0,1]\{X_t\}_{t\in[0,1]}7, called the Sinkhorn flow (Deb et al., 2023). Across these formulations, “marginal flow” refers less to a particular architecture than to a dynamical law constrained or characterized by marginals.

2. Marginal-preserving optimal transport and interior updates

The most explicit algorithmic use of the name is the cost-specific extension of rectified flow in “Rectified Flow: A Marginal Preserving Approach to Optimal Transport” (Liu, 2022). The problem is the Monge–Kantorovich program

{Xt}t[0,1]\{X_t\}_{t\in[0,1]}8

with {Xt}t[0,1]\{X_t\}_{t\in[0,1]}9 convex. The method starts from any valid coupling, often vtX(z)=E[X˙tXt=z],v_t^{\mathbf X}(z)=\mathbb{E}[\dot X_t \mid X_t=z],0, forms the linear interpolation

vtX(z)=E[X˙tXt=z],v_t^{\mathbf X}(z)=\mathbb{E}[\dot X_t \mid X_t=z],1

and learns a new velocity restricted to the structured class

vtX(z)=E[X˙tXt=z],v_t^{\mathbf X}(z)=\mathbb{E}[\dot X_t \mid X_t=z],2

where vtX(z)=E[X˙tXt=z],v_t^{\mathbf X}(z)=\mathbb{E}[\dot X_t \mid X_t=z],3 minimizes

vtX(z)=E[X˙tXt=z],v_t^{\mathbf X}(z)=\mathbb{E}[\dot X_t \mid X_t=z],4

The updated coupling is obtained from the neural ODE

vtX(z)=E[X˙tXt=z],v_t^{\mathbf X}(z)=\mathbb{E}[\dot X_t \mid X_t=z],5

The defining feature is the “interior approach”: the iterate remains feasible by construction. The update changes the coupling while preserving all time marginals, rather than enforcing constraints through dual variables or penalties. The velocity admits a Bregman/Helmholtz-like decomposition

vtX(z)=E[X˙tXt=z],v_t^{\mathbf X}(z)=\mathbb{E}[\dot X_t \mid X_t=z],6

where the residual is exactly the marginal-preserving component removable without changing the marginals. The cost-improvement theorem states

vtX(z)=E[X˙tXt=z],v_t^{\mathbf X}(z)=\mathbb{E}[\dot X_t \mid X_t=z],7

and the fixed-point criterion is

vtX(z)=E[X˙tXt=z],v_t^{\mathbf X}(z)=\mathbb{E}[\dot X_t \mid X_t=z],8

For quadratic cost,

vtX(z)=E[X˙tXt=z],v_t^{\mathbf X}(z)=\mathbb{E}[\dot X_t \mid X_t=z],9

the construction reduces to standard rectified flow. A common misconception is that “marginal-preserving” means the process is unchanged. In fact, the entire purpose is to alter the coupling while keeping the prescribed marginals fixed (Liu, 2022).

3. Multi-marginal, bridge, and continuum-time formulations

Several works generalize marginal flow from two endpoint constraints to multi-time or all-time settings. In the multi-marginal Schrödinger bridge problem for inertial particles, the state evolves by

$\rho_t=\law(X_t)$0

but only positional marginals

$\rho_t=\law(X_t)$1

are observed. The solution is the path-space law $\rho_t=\law(X_t)$2 minimizing relative entropy with respect to the inertial Brownian prior, equivalently the optimal control problem

$\rho_t=\law(X_t)$3

subject to

$\rho_t=\law(X_t)$4

In density form, the formulation becomes a Benamou–Brenier-like variational problem, and after the substitution $\rho_t=\law(X_t)$5, the cost acquires a Fisher information term in the velocity variable (Chen et al., 2019).

For high-dimensional snapshot data at irregular time points, Multi-Marginal Stochastic Flow Matching (MMSFM) models a stochastic flow

$\rho_t=\law(X_t)$6

with the decomposition

$\rho_t=\law(X_t)$7

The method constructs overlapping windows of consecutive marginals, uses transport splines to build Gaussian conditional paths

$\rho_t=\law(X_t)$8

and trains flow and score networks jointly by a simulation-free objective. The paper emphasizes spline reparameterization over actual time intervals and stratified time sampling as the mechanisms that handle irregular snapshot timing (Lee et al., 6 Aug 2025).

ALI-CFM introduces adversarially learnt interpolants

$\rho_t=\law(X_t)$9

and matches intermediate snapshot distributions through a GAN-style objective so that

tρt+ ⁣ ⁣(vtXρt)=0.\partial_t \rho_t + \nabla\!\cdot\!\big(v_t^{\mathbf X}\rho_t\big)=0.0

at observed times. The learned interpolants are then marginalized by conditional flow matching to obtain a time-dependent vector field (Kviman et al., 1 Oct 2025).

A further continuum limit is the Sinkhorn flow, which is both a limit of entropic OT iterations and an example of Wasserstein mirror gradient flow. Its continuity equation is

tρt+ ⁣ ⁣(vtXρt)=0.\partial_t \rho_t + \nabla\!\cdot\!\big(v_t^{\mathbf X}\rho_t\big)=0.1

with mirror-gradient velocity

tρt+ ⁣ ⁣(vtXρt)=0.\partial_t \rho_t + \nabla\!\cdot\!\big(v_t^{\mathbf X}\rho_t\big)=0.2

and potential evolution governed by the parabolic Monge–Ampère equation

tρt+ ⁣ ⁣(vtXρt)=0.\partial_t \rho_t + \nabla\!\cdot\!\big(v_t^{\mathbf X}\rho_t\big)=0.3

The paper also proves exponential convergence under a logarithmic Sobolev inequality and a uniform lower bound on mirror curvature (Deb et al., 2023).

4. Density estimation, tail control, and parameter marginalization

In density modeling, “Marginal Flow” denotes a different construction: a conditional density with marginalized latent parameters. The framework defines

tρt+ ⁣ ⁣(vtXρt)=0.\partial_t \rho_t + \nabla\!\cdot\!\big(v_t^{\mathbf X}\rho_t\big)=0.4

implemented in practice as

tρt+ ⁣ ⁣(vtXρt)=0.\partial_t \rho_t + \nabla\!\cdot\!\big(v_t^{\mathbf X}\rho_t\big)=0.5

The latent parameters are sampled by

tρt+ ⁣ ⁣(vtXρt)=0.\partial_t \rho_t + \nabla\!\cdot\!\big(v_t^{\mathbf X}\rho_t\big)=0.6

so tρt+ ⁣ ⁣(vtXρt)=0.\partial_t \rho_t + \nabla\!\cdot\!\big(v_t^{\mathbf X}\rho_t\big)=0.7 need not be invertible and the method imposes no bijectivity constraint. The paper emphasizes exact density evaluation, efficient sampling, forward-KL or reverse-KL training, architectural freedom, and the ability to model lower-dimensional manifolds by choosing tρt+ ⁣ ⁣(vtXρt)=0.\partial_t \rho_t + \nabla\!\cdot\!\big(v_t^{\mathbf X}\rho_t\big)=0.8 in tρt+ ⁣ ⁣(vtXρt)=0.\partial_t \rho_t + \nabla\!\cdot\!\big(v_t^{\mathbf X}\rho_t\big)=0.9 with rtr_t0 (Negri et al., 30 Sep 2025).

A related but older use of “marginal flow” appears in copula-and-marginal generative flows. There the joint law is decomposed as

rtr_t1

with a copula flow rtr_t2 and a marginal flow rtr_t3 applied componentwise: rtr_t4 The univariate marginal flow is a hybrid inverse-CDF construction,

rtr_t5

so exact tail behavior is encoded directly where prior knowledge exists, while the interior is learned (Wiese et al., 2019).

Tail control is also the theme of Marginal Tail-Adaptive Normalizing Flows. The core theorem states that, under the stated conditions and the copula assumption, if the first rtr_t6 marginals are light-tailed and the remaining marginals are heavy-tailed, the transformed variable has the same marginal tail behavior as the base. The architecture therefore uses a factorized base distribution with Gaussian marginals for light-tailed coordinates and Student-rtr_t7 marginals for heavy-tailed coordinates, together with tail-preserving permutations and block-triangular linear layers

rtr_t8

The method is designed to preserve mixed-tailed structure rather than merely global heaviness (Laszkiewicz et al., 2022).

5. Marginal-conditioned sampling and few-step distillation

A further line of work uses marginal information at inference time rather than training-time density construction. In bridge-aware discretization, the relevant quantity is the difference between conditional bridge geometry and marginal population flow. The conditional–marginal entropy-rate scheduler is based on

rtr_t9

followed by CDF inversion

X\mathbf X0

For Gaussian Brownian bridges this rate is closed-form and U-shaped, motivating boundary-heavy nonuniform grids. The reported low-budget gains include an 18.1% 10-step ODE-Heun MMD improvement over linear and a paired 22.7% SDE-Heun improvement; on CIFAR-10, the entropic schedule gives the best tested five-step FID, X\mathbf X1, versus X\mathbf X2 for linear and X\mathbf X3 for cosine (Trentini et al., 15 May 2026).

In Flow LLMs, the denoiser block X\mathbf X4 is exactly a posterior marginal probability over clean tokens: X\mathbf X5 The marginal-conditioned bridge sampler replaces a simplex-valued conditional mean endpoint by a sampled one-hot endpoint from the factorized posterior

X\mathbf X6

and then uses the exact Ornstein–Uhlenbeck bridge kernel. Under exact posterior marginals, the endpoint approximation error is exactly the conditional multi-information among token positions, and the induced one-step bridge kernel preserves all token-wise posterior-predictive marginals while losing only the residual cross-position dependence (Azangulov et al., 13 May 2026).

In few-step 3D generation, MDT-dist defines “Marginal-Data Transport” as transport from a noisy marginal X\mathbf X7 back to data. The primary objective is

X\mathbf X8

which is converted into Velocity Matching and Velocity Distillation. On TRELLIS, the method reduces the sampling steps of each flow transformer from 25 to 1 or 2, achieving X\mathbf X9s (1 step 0tE[h(Xs)rs(Xs)]ds=0hCc,\int_0^t \mathbb{E}\big[\nabla h(X_s)\cdot r_s(X_s)\big]\,ds = 0 \quad \forall h\in \mathcal C_c,0 2) and 0tE[h(Xs)rs(Xs)]ds=0hCc,\int_0^t \mathbb{E}\big[\nabla h(X_s)\cdot r_s(X_s)\big]\,ds = 0 \quad \forall h\in \mathcal C_c,1s (2 steps 0tE[h(Xs)rs(Xs)]ds=0hCc,\int_0^t \mathbb{E}\big[\nabla h(X_s)\cdot r_s(X_s)\big]\,ds = 0 \quad \forall h\in \mathcal C_c,2 2) latency with 0tE[h(Xs)rs(Xs)]ds=0hCc,\int_0^t \mathbb{E}\big[\nabla h(X_s)\cdot r_s(X_s)\big]\,ds = 0 \quad \forall h\in \mathcal C_c,3 and 0tE[h(Xs)rs(Xs)]ds=0hCc,\int_0^t \mathbb{E}\big[\nabla h(X_s)\cdot r_s(X_s)\big]\,ds = 0 \quad \forall h\in \mathcal C_c,4 speedup on A800 (Zhou et al., 4 Sep 2025).

6. Marginal flow in field theory and critical phenomena

In high-energy theory and critical phenomena, “marginal flow” refers to renormalization-group behavior generated by marginal or marginally irrelevant operators rather than probability transport. In holographic RG, a classically marginal operator with 0tE[h(Xs)rs(Xs)]ds=0hCc,\int_0^t \mathbb{E}\big[\nabla h(X_s)\cdot r_s(X_s)\big]\,ds = 0 \quad \forall h\in \mathcal C_c,5 has vanishing classical beta function,

0tE[h(Xs)rs(Xs)]ds=0hCc,\int_0^t \mathbb{E}\big[\nabla h(X_s)\cdot r_s(X_s)\big]\,ds = 0 \quad \forall h\in \mathcal C_c,6

but quantum corrections produce

0tE[h(Xs)rs(Xs)]ds=0hCc,\int_0^t \mathbb{E}\big[\nabla h(X_s)\cdot r_s(X_s)\big]\,ds = 0 \quad \forall h\in \mathcal C_c,7

so a classically marginal deformation can become truly marginal, marginally relevant, or marginally irrelevant. The holographic dual is a 5D Einstein-scalar model with massless bulk scalar and domain-wall metric

0tE[h(Xs)rs(Xs)]ds=0hCc,\int_0^t \mathbb{E}\big[\nabla h(X_s)\cdot r_s(X_s)\big]\,ds = 0 \quad \forall h\in \mathcal C_c,8

The trivial branch 0tE[h(Xs)rs(Xs)]ds=0hCc,\int_0^t \mathbb{E}\big[\nabla h(X_s)\cdot r_s(X_s)\big]\,ds = 0 \quad \forall h\in \mathcal C_c,9, (rtρt)=0.\nabla\cdot\big(r_t\rho_t\big)=0.0 is pure AdS and corresponds to a truly marginal deformation, whereas the nontrivial scalar profile encodes a quantum-corrected RG flow (Park, 2021).

A structurally different use appears in the “triple (rtρt)=0.\nabla\cdot\big(r_t\rho_t\big)=0.1-like flow,” where the one-parameter deformation

(rtρt)=0.\nabla\cdot\big(r_t\rho_t\big)=0.2

organizes irrelevant, marginal, and relevant branches. The marginal point is (rtρt)=0.\nabla\cdot\big(r_t\rho_t\big)=0.3, for which

(rtρt)=0.\nabla\cdot\big(r_t\rho_t\big)=0.4

reproducing the root-(rtρt)=0.\nabla\cdot\big(r_t\rho_t\big)=0.5 / ModMax branch. The paper states: “However, the only marginal theory is ModMax, and it is unique” (Babaei-Aghbolagh et al., 30 May 2026).

In endpoint criticality, marginal flow refers to a marginally irrelevant scaling field that produces logarithmically slow renormalization and delays asymptotic scaling. For the square-lattice (rtρt)=0.\nabla\cdot\big(r_t\rho_t\big)=0.6–(rtρt)=0.\nabla\cdot\big(r_t\rho_t\big)=0.7 Ising endpoint, the generalized field-mixing framework uses one relevant field plus one marginally irrelevant one, with four-state Potts exponents

(rtρt)=0.\nabla\cdot\big(r_t\rho_t\big)=0.8

so the scaling variable is

(rtρt)=0.\nabla\cdot\big(r_t\rho_t\big)=0.9

The paper emphasizes that negative Binder dips and bimodal histograms are not decisive by themselves, because they are compatible with Potts criticality plus strong finite-size corrections (Sun et al., 25 May 2026).

Boundary conformal field theory supplies yet another marginal notion. In the presence of a boundary, Herzog and Shamir propose deformations of the form

tt0

which remain compatible with the preserved tt1 symmetry when tt2 is protected. Here the boundary compensates for the mismatch between tt3 and tt4, generalizing the notion of exact marginality beyond translation-invariant bulk deformations (Herzog et al., 2019).

Across these literatures, the common denominator is precise but abstract: a “marginal flow” is a dynamics whose defining object is not the full microscopic configuration alone but the evolution, preservation, or deformation of marginals. What changes from domain to domain is whether those marginals are endpoint laws in optimal transport, all-time probability distributions, posterior token marginals, noisy generative intermediates, or RG couplings associated with marginal operators.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Marginal Flow.