Papers
Topics
Authors
Recent
Search
2000 character limit reached

Monge–Kantorovich Transportation Distance

Updated 10 July 2026
  • The Monge–Kantorovich transportation distance is defined as the optimal cost of redistributing mass between probability measures subject to prescribed marginal constraints.
  • It contrasts the classical Monge map with Kantorovich’s relaxed transport plans, enabling mass splitting and facilitating dual formulations like the Wasserstein-1 distance.
  • Extensions include one-dimensional, circular, matrix-valued, quantum, and noncommutative variants, applied in image analysis, statistical inference, and advanced optimal transport metrics.

The Monge–Kantorovich transportation distance is the optimal transport cost obtained by minimizing an average transportation cost over all couplings with prescribed marginals. In the classical setting, if (X,d)(X,d) is a metric space and μ,μ~\mu,\tilde\mu are probability measures on XX, the order-1 Monge–Kantorovich, or Wasserstein-1, distance is defined by

W(μ,μ~)=infπX×Xd(x,y)dπ(x,y),W(\mu,\tilde\mu)=\inf_{\pi}\int_{X\times X} d(x,y)\,d\pi(x,y),

where the infimum runs over all couplings π\pi of μ\mu and μ~\tilde\mu; more generally, one minimizes cdπ\int c\,d\pi for a prescribed cost cc, and with quadratic ground cost one obtains the L2L^2 Wasserstein distance (Martinetti, 2012, Snow et al., 2018).

1. Classical formulations and duality

The classical theory distinguishes between Monge’s original problem and Kantorovich’s relaxation. Monge seeks a transport map μ,μ~\mu,\tilde\mu0 such that

μ,μ~\mu,\tilde\mu1

and minimizes

μ,μ~\mu,\tilde\mu2

Kantorovich replaces the map μ,μ~\mu,\tilde\mu3 by a transport plan μ,μ~\mu,\tilde\mu4, a nonnegative measure on μ,μ~\mu,\tilde\mu5, with marginal constraints

μ,μ~\mu,\tilde\mu6

and minimizes

μ,μ~\mu,\tilde\mu7

The passage from μ,μ~\mu,\tilde\mu8 to μ,μ~\mu,\tilde\mu9 allows mass splitting, whereas Monge’s formulation requires a single-valued map, so every unit of mass at a source point must move to exactly one target point; in particular, mass cannot split. This distinction is central in image comparison and in other settings where one-to-one matching is too restrictive (Snow et al., 2018).

For power costs, the standard Wasserstein family is

XX0

and in the quadratic case

XX1

The quadratic case is also the regime in which Brenier-type structure appears: under suitable conditions the optimal transport map is the gradient of a convex potential, and the Monge–Ampère equation becomes a natural analytic representation of the same transport geometry (Snow et al., 2018).

For XX2, Kantorovich duality rewrites the infimum over couplings as a supremum over XX3-Lipschitz test functions: XX4 In the commutative spectral-triple setting on a complete manifold, this dual formula coincides with Connes’ spectral distance because the operator norm of the commutator is exactly the Lipschitz norm. This observation supplies a direct bridge between transport duality and metric notions from noncommutative geometry (Martinetti, 2012).

2. One-dimensional structure and circular transport

On the real line, the transportation cost has a particularly explicit representation. For distributions XX5 with quantile functions XX6,

XX7

This quantile formula underlies the asymptotic theory of empirical transport costs developed for XX8 on XX9. In the two-sample setting, with independent empirical cdfs W(μ,μ~)=infπX×Xd(x,y)dπ(x,y),W(\mu,\tilde\mu)=\inf_{\pi}\int_{X\times X} d(x,y)\,d\pi(x,y),0 and W(μ,μ~)=infπX×Xd(x,y)dπ(x,y),W(\mu,\tilde\mu)=\inf_{\pi}\int_{X\times X} d(x,y)\,d\pi(x,y),1, the paper proves a central limit theorem for W(μ,μ~)=infπX×Xd(x,y)dπ(x,y),W(\mu,\tilde\mu)=\inf_{\pi}\int_{X\times X} d(x,y)\,d\pi(x,y),2 under finite W(μ,μ~)=infπX×Xd(x,y)dπ(x,y),W(\mu,\tilde\mu)=\inf_{\pi}\int_{X\times X} d(x,y)\,d\pi(x,y),3-moments and continuity of quantile functions, together with a consistent variance estimator and a studentized similarity test. The analysis is explicitly for the nondegenerate case W(μ,μ~)=infπX×Xd(x,y)dπ(x,y),W(\mu,\tilde\mu)=\inf_{\pi}\int_{X\times X} d(x,y)\,d\pi(x,y),4, while the null case W(μ,μ~)=infπX×Xd(x,y)dπ(x,y),W(\mu,\tilde\mu)=\inf_{\pi}\int_{X\times X} d(x,y)\,d\pi(x,y),5 is known to exhibit different asymptotic behavior (Barrio et al., 2018).

On the unit circle, the absence of a canonical origin changes the transport geometry. For a geodesic distance

W(μ,μ~)=infπX×Xd(x,y)dπ(x,y),W(\mu,\tilde\mu)=\inf_{\pi}\int_{X\times X} d(x,y)\,d\pi(x,y),6

and a cost of the form

W(μ,μ~)=infπX×Xd(x,y)dπ(x,y),W(\mu,\tilde\mu)=\inf_{\pi}\int_{X\times X} d(x,y)\,d\pi(x,y),7

with W(μ,μ~)=infπX×Xd(x,y)dπ(x,y),W(\mu,\tilde\mu)=\inf_{\pi}\int_{X\times X} d(x,y)\,d\pi(x,y),8 increasing and convex, the circular Monge–Kantorovich problem reduces to a one-dimensional problem after cutting the circle at a suitable point. If W(μ,μ~)=infπX×Xd(x,y)dπ(x,y),W(\mu,\tilde\mu)=\inf_{\pi}\int_{X\times X} d(x,y)\,d\pi(x,y),9 and π\pi0 are cumulative distribution functions and π\pi1, then

π\pi2

For the geodesic cost itself, π\pi3, this simplifies to

π\pi4

For discrete circular histograms, the same quantity becomes

π\pi5

where π\pi6 is a median of the cumulative differences. The paper presents this as the Circular Earth Mover’s Distance and emphasizes that it can be computed in linear time (0906.5499).

These one-dimensional and circular formulas are significant because they replace a general linear program by explicit cumulative or quantile expressions. A plausible implication is that the most transparent formulas for Monge–Kantorovich distance arise when the ambient geometry supplies a total order or can be reduced to one.

3. Discrete, graph, and PDE formulations

In discrete image comparison, grayscale images can be treated as densities concentrated at pixel centers: π\pi7 with equal total mass

π\pi8

The discrete Monge–Kantorovich problem then becomes

π\pi9

subject to

μ\mu0

This is a standard linear programming problem, and in the cited image-comparison study each image is normalized so that total pixel sum μ\mu1 before the transport distance is used in a 1-nearest-neighbour classifier (Snow et al., 2018).

For μ\mu2-based transport on a bounded convex domain, a different formulation uses the Monge–Kantorovich PDE system

μ\mu3

and the Beckmann problem

μ\mu4

The paper proposes a dynamic model

μ\mu5

and uses it as a numerical route to the μ\mu6 transport density and the Wasserstein-1 cost

μ\mu7

The numerical implementation is based on low-order finite elements and a two-mesh stabilization in which μ\mu8 and μ\mu9 (Facca et al., 2017).

On a finite metric space, the Kantorovich problem can be written on a transportation polytope of nonnegative tables with prescribed row and column sums. The paper on finite-space OT interprets feasible couplings as contingency tables and studies “moves” preserving the marginals; it then proposes a simulated-annealing MCMC over feasible tables. In the graph-metric setting, a different exact reformulation is available: if μ~\tilde\mu0 is induced by a weighted connected graph μ~\tilde\mu1, then

μ~\tilde\mu2

where the minimum runs over spanning trees. Once an optimal spanning tree is known, the paper derives an explicit Kantorovich potential from the imbalanced cumulative mass

μ~\tilde\mu3

and constructs an optimal plan by dynamic programming on the tree (Pistone et al., 2020, Bigot et al., 13 Jan 2026).

4. Statistical, imaging, and data-analytic uses

The cited image-comparison study treats the Monge–Kantorovich transportation distance as a geometry-aware image metric. Images are converted into normalized mass distributions over pixel locations, and the cost is the minimum total squared travel cost needed to redistribute one image’s mass into the other’s. In a 1-nearest-neighbour classifier on MNIST, the paper compares Euclidean distance, Tangent Space Distance, the Kantorovich transport distance, and a PDE-based Monge transport distance. Both optimal transport formulations outperform Euclidean distance and Tangent Space Distance, and the Kantorovich LP and PDE methods give very similar classification performance (Snow et al., 2018).

In statistics on the real line, the transport cost becomes an inferential quantity. For μ~\tilde\mu4, the asymptotic variance

μ~\tilde\mu5

supports confidence intervals and a two-sample similarity test for

μ~\tilde\mu6

The same paper uses this machinery for fairness assessment by comparing the conditional laws of a classifier score given a protected attribute, and interprets decreasing Wasserstein distance between those conditional score distributions as a reduction in disparity (Barrio et al., 2018).

Quadratic optimal transport maps also support a center-outward statistical geometry. With a reference distribution μ~\tilde\mu7, usually the spherical uniform distribution μ~\tilde\mu8 on the unit ball, the Brenier map μ~\tilde\mu9 pushes cdπ\int c\,d\pi0 to a target distribution cdπ\int c\,d\pi1, while cdπ\int c\,d\pi2 pushes cdπ\int c\,d\pi3 back to cdπ\int c\,d\pi4. This yields Monge–Kantorovich depth, quantiles, ranks, and signs. In the spherical-reference case, the MK rank of cdπ\int c\,d\pi5 is cdπ\int c\,d\pi6, the MK sign is cdπ\int c\,d\pi7, and the MK depth is

cdπ\int c\,d\pi8

The paper emphasizes that MK depth reduces to halfspace depth for spherical or elliptical families, but can account for non-convex features of the target distribution in more general settings (Chernozhukov et al., 2014).

A different statistical reinterpretation embeds transport costs of the form

cdπ\int c\,d\pi9

into a statistical manifold built from the natural exponential family generated by a base measure cc0. In that setting, generalized transport potentials on cc1 become ordinary supergradients of an exponentially concave function on a convex space of probability measures absolutely continuous with respect to cc2. This identifies a “universal geometry” for a class of Monge–Kantorovich problems that includes the quadratic Gaussian case (Pal, 2017).

5. Relative, matrix-valued, quantum, and noncommutative variants

Several cited works generalize the Monge–Kantorovich framework by changing either the admissible measures or the transported objects. A relative version is formulated on a metric pair cc3, where the closed subset cc4 acts as a mass reservoir. The modified kernel

cc5

allows comparison of measures with different total mass through transport to and from cc6. The corresponding relative Wasserstein-1 problem is

cc7

and the relative Kantorovich–Rubinstein duality becomes

cc8

This is a geometric, reservoir-based unbalanced transport model rather than an unbalanced model based on KL, TV, or Hellinger penalties (Bubenik et al., 2024).

For matrix-valued densities, the transport plan is no longer a scalar joint density but a positive semidefinite operator on a tensor-product space, and marginalization is implemented by partial trace. One formulation introduces a cost

cc9

where the additional Frobenius term penalizes rotation between normalized source and target matrix orientations. The unrestricted cost is symmetric, nonnegative, and definite, but does not satisfy the triangle inequality in general; a restricted version yields a genuine metric after taking the square root (Ning et al., 2013). A related noncommutative dynamic formulation replaces scalar continuity equations by

L2L^20

and defines a matrix Wasserstein distance

L2L^21

This formulation supports convexity, strong duality, and constant-speed geodesics in the matrix setting (Chen et al., 2017).

Quantum variants replace classical couplings by bipartite density operators or bipartite quantum states with fixed marginals. One paper defines

L2L^22

for density operators L2L^23 on L2L^24, proves a Kantorovich-type duality theorem, and derives a gradient-type structure

L2L^25

for optimal quantum couplings (Caglioti et al., 2021). Another paper studies finite-dimensional density matrices and defines

L2L^26

with a distinguished cost

L2L^27

equal to the projector onto the antisymmetric subspace. In all dimensions this yields a semidistance, and for qubits its square root is a genuine metric; the associated SWAP-fidelity satisfies

L2L^28

The paper interprets L2L^29 as a quantum analogue of a Wasserstein-2 distance on the Bloch ball (Friedland et al., 2021).

In noncommutative geometry, the dual point of view becomes primary. For a spectral triple μ,μ~\mu,\tilde\mu00, the paper defines

μ,μ~\mu,\tilde\mu01

where μ,μ~\mu,\tilde\mu02 consists of elements whose evaluations on pure states are μ,μ~\mu,\tilde\mu03-Lipschitz for the spectral distance μ,μ~\mu,\tilde\mu04. One always has

μ,μ~\mu,\tilde\mu05

and equality is proved on pure states, on convex segments generated by two pure states, and on the whole state space of μ,μ~\mu,\tilde\mu06 (Martinetti, 2012).

6. Multistochastic, causal, geometric, and higher-order developments

The classical constraint of fixing one-dimensional marginals can be replaced by stronger projection constraints. In the multistochastic μ,μ~\mu,\tilde\mu07-Monge–Kantorovich problem, one fixes all μ,μ~\mu,\tilde\mu08-dimensional coordinate projections of a probability measure on μ,μ~\mu,\tilde\mu09 and minimizes or maximizes μ,μ~\mu,\tilde\mu10. In the model case

μ,μ~\mu,\tilde\mu11

with all pairwise projections equal to Lebesgue measure on μ,μ~\mu,\tilde\mu12, the paper proves that the minimizer is supported on

μ,μ~\mu,\tilde\mu13

where μ,μ~\mu,\tilde\mu14 is bitwise xor, and that the support is the Sierpiński tetrahedron. It also constructs an explicit dual optimizer for this problem (Gladkov et al., 2018).

A different structural constraint is causality. For filtered Polish spaces μ,μ~\mu,\tilde\mu15 and μ,μ~\mu,\tilde\mu16, a coupling μ,μ~\mu,\tilde\mu17 is causal if the conditional law of the second component up to time μ,μ~\mu,\tilde\mu18 is measurable with respect to the source filtration up to time μ,μ~\mu,\tilde\mu19. The paper denotes the causal set by

μ,μ~\mu,\tilde\mu20

and shows that with trivial filtrations one recovers the usual optimal transport problem, whereas in stochastic settings the resulting causal Monge–Kantorovich problems are related to stochastic control, stochastic calculus, martingales, and strong solutions of SDEs (Lassalle, 2013).

On general geodesic spaces, the Monge problem for cost μ,μ~\mu,\tilde\mu21 can be reduced to one-dimensional transport problems along transport geodesics. In a non-branching geodesic space, the transport set is partitioned into rays, and under non-degeneracy conditions on the source measure the conditional measures on those rays are continuous or absolutely continuous with respect to one-dimensional Hausdorff measure. This permits construction of a Monge transport map by solving one-dimensional transport problems on each ray and gluing them together. The paper also emphasizes that in this setting μ,μ~\mu,\tilde\mu22-cyclical monotonicity is not sufficient for optimality (Bianchini et al., 2011).

The transport problem itself can also be lifted to Wasserstein space. Given probability laws

μ,μ~\mu,\tilde\mu23

on the space of probability measures over a smooth compact manifold μ,μ~\mu,\tilde\mu24, the paper studies

μ,μ~\mu,\tilde\mu25

Under a Rademacher property for a reference measure on μ,μ~\mu,\tilde\mu26, absolute continuity of the source law with respect to that reference measure, and concentration of the source law on absolutely continuous measures on μ,μ~\mu,\tilde\mu27, the optimal plan exists, is unique, and is induced by a measurable map

μ,μ~\mu,\tilde\mu28

An extension to costs of the form μ,μ~\mu,\tilde\mu29, with μ,μ~\mu,\tilde\mu30 strictly increasing and strictly convex, is also proved (Emami et al., 2024).

Finally, recent quantitative analysis addresses stability and second variation of the quadratic Monge–Kantorovich distance. For non-degenerate Hölder densities on uniformly convex domains, the cited paper proves

μ,μ~\mu,\tilde\mu31

where μ,μ~\mu,\tilde\mu32, and also derives an explicit second-variation formula. If μ,μ~\mu,\tilde\mu33, μ,μ~\mu,\tilde\mu34, and μ,μ~\mu,\tilde\mu35, then

μ,μ~\mu,\tilde\mu36

This gives a precise Hessian-type formula for the quadratic transport cost around a regular optimal transport configuration (Caja-Lopez et al., 22 May 2026).

These generalizations indicate that the Monge–Kantorovich transportation distance is not restricted to balanced scalar transport on Euclidean domains. The same primal–dual structure persists, with substantial modifications, for random measures, higher-order projection constraints, causal filtrations, matrix-valued densities, quantum states, and noncommutative state spaces.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (19)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Monge-Kantorovich Transportation Distance.