Monge–Kantorovich Transportation Distance
- The Monge–Kantorovich transportation distance is defined as the optimal cost of redistributing mass between probability measures subject to prescribed marginal constraints.
- It contrasts the classical Monge map with Kantorovich’s relaxed transport plans, enabling mass splitting and facilitating dual formulations like the Wasserstein-1 distance.
- Extensions include one-dimensional, circular, matrix-valued, quantum, and noncommutative variants, applied in image analysis, statistical inference, and advanced optimal transport metrics.
The Monge–Kantorovich transportation distance is the optimal transport cost obtained by minimizing an average transportation cost over all couplings with prescribed marginals. In the classical setting, if is a metric space and are probability measures on , the order-1 Monge–Kantorovich, or Wasserstein-1, distance is defined by
where the infimum runs over all couplings of and ; more generally, one minimizes for a prescribed cost , and with quadratic ground cost one obtains the Wasserstein distance (Martinetti, 2012, Snow et al., 2018).
1. Classical formulations and duality
The classical theory distinguishes between Monge’s original problem and Kantorovich’s relaxation. Monge seeks a transport map 0 such that
1
and minimizes
2
Kantorovich replaces the map 3 by a transport plan 4, a nonnegative measure on 5, with marginal constraints
6
and minimizes
7
The passage from 8 to 9 allows mass splitting, whereas Monge’s formulation requires a single-valued map, so every unit of mass at a source point must move to exactly one target point; in particular, mass cannot split. This distinction is central in image comparison and in other settings where one-to-one matching is too restrictive (Snow et al., 2018).
For power costs, the standard Wasserstein family is
0
and in the quadratic case
1
The quadratic case is also the regime in which Brenier-type structure appears: under suitable conditions the optimal transport map is the gradient of a convex potential, and the Monge–Ampère equation becomes a natural analytic representation of the same transport geometry (Snow et al., 2018).
For 2, Kantorovich duality rewrites the infimum over couplings as a supremum over 3-Lipschitz test functions: 4 In the commutative spectral-triple setting on a complete manifold, this dual formula coincides with Connes’ spectral distance because the operator norm of the commutator is exactly the Lipschitz norm. This observation supplies a direct bridge between transport duality and metric notions from noncommutative geometry (Martinetti, 2012).
2. One-dimensional structure and circular transport
On the real line, the transportation cost has a particularly explicit representation. For distributions 5 with quantile functions 6,
7
This quantile formula underlies the asymptotic theory of empirical transport costs developed for 8 on 9. In the two-sample setting, with independent empirical cdfs 0 and 1, the paper proves a central limit theorem for 2 under finite 3-moments and continuity of quantile functions, together with a consistent variance estimator and a studentized similarity test. The analysis is explicitly for the nondegenerate case 4, while the null case 5 is known to exhibit different asymptotic behavior (Barrio et al., 2018).
On the unit circle, the absence of a canonical origin changes the transport geometry. For a geodesic distance
6
and a cost of the form
7
with 8 increasing and convex, the circular Monge–Kantorovich problem reduces to a one-dimensional problem after cutting the circle at a suitable point. If 9 and 0 are cumulative distribution functions and 1, then
2
For the geodesic cost itself, 3, this simplifies to
4
For discrete circular histograms, the same quantity becomes
5
where 6 is a median of the cumulative differences. The paper presents this as the Circular Earth Mover’s Distance and emphasizes that it can be computed in linear time (0906.5499).
These one-dimensional and circular formulas are significant because they replace a general linear program by explicit cumulative or quantile expressions. A plausible implication is that the most transparent formulas for Monge–Kantorovich distance arise when the ambient geometry supplies a total order or can be reduced to one.
3. Discrete, graph, and PDE formulations
In discrete image comparison, grayscale images can be treated as densities concentrated at pixel centers: 7 with equal total mass
8
The discrete Monge–Kantorovich problem then becomes
9
subject to
0
This is a standard linear programming problem, and in the cited image-comparison study each image is normalized so that total pixel sum 1 before the transport distance is used in a 1-nearest-neighbour classifier (Snow et al., 2018).
For 2-based transport on a bounded convex domain, a different formulation uses the Monge–Kantorovich PDE system
3
and the Beckmann problem
4
The paper proposes a dynamic model
5
and uses it as a numerical route to the 6 transport density and the Wasserstein-1 cost
7
The numerical implementation is based on low-order finite elements and a two-mesh stabilization in which 8 and 9 (Facca et al., 2017).
On a finite metric space, the Kantorovich problem can be written on a transportation polytope of nonnegative tables with prescribed row and column sums. The paper on finite-space OT interprets feasible couplings as contingency tables and studies “moves” preserving the marginals; it then proposes a simulated-annealing MCMC over feasible tables. In the graph-metric setting, a different exact reformulation is available: if 0 is induced by a weighted connected graph 1, then
2
where the minimum runs over spanning trees. Once an optimal spanning tree is known, the paper derives an explicit Kantorovich potential from the imbalanced cumulative mass
3
and constructs an optimal plan by dynamic programming on the tree (Pistone et al., 2020, Bigot et al., 13 Jan 2026).
4. Statistical, imaging, and data-analytic uses
The cited image-comparison study treats the Monge–Kantorovich transportation distance as a geometry-aware image metric. Images are converted into normalized mass distributions over pixel locations, and the cost is the minimum total squared travel cost needed to redistribute one image’s mass into the other’s. In a 1-nearest-neighbour classifier on MNIST, the paper compares Euclidean distance, Tangent Space Distance, the Kantorovich transport distance, and a PDE-based Monge transport distance. Both optimal transport formulations outperform Euclidean distance and Tangent Space Distance, and the Kantorovich LP and PDE methods give very similar classification performance (Snow et al., 2018).
In statistics on the real line, the transport cost becomes an inferential quantity. For 4, the asymptotic variance
5
supports confidence intervals and a two-sample similarity test for
6
The same paper uses this machinery for fairness assessment by comparing the conditional laws of a classifier score given a protected attribute, and interprets decreasing Wasserstein distance between those conditional score distributions as a reduction in disparity (Barrio et al., 2018).
Quadratic optimal transport maps also support a center-outward statistical geometry. With a reference distribution 7, usually the spherical uniform distribution 8 on the unit ball, the Brenier map 9 pushes 0 to a target distribution 1, while 2 pushes 3 back to 4. This yields Monge–Kantorovich depth, quantiles, ranks, and signs. In the spherical-reference case, the MK rank of 5 is 6, the MK sign is 7, and the MK depth is
8
The paper emphasizes that MK depth reduces to halfspace depth for spherical or elliptical families, but can account for non-convex features of the target distribution in more general settings (Chernozhukov et al., 2014).
A different statistical reinterpretation embeds transport costs of the form
9
into a statistical manifold built from the natural exponential family generated by a base measure 0. In that setting, generalized transport potentials on 1 become ordinary supergradients of an exponentially concave function on a convex space of probability measures absolutely continuous with respect to 2. This identifies a “universal geometry” for a class of Monge–Kantorovich problems that includes the quadratic Gaussian case (Pal, 2017).
5. Relative, matrix-valued, quantum, and noncommutative variants
Several cited works generalize the Monge–Kantorovich framework by changing either the admissible measures or the transported objects. A relative version is formulated on a metric pair 3, where the closed subset 4 acts as a mass reservoir. The modified kernel
5
allows comparison of measures with different total mass through transport to and from 6. The corresponding relative Wasserstein-1 problem is
7
and the relative Kantorovich–Rubinstein duality becomes
8
This is a geometric, reservoir-based unbalanced transport model rather than an unbalanced model based on KL, TV, or Hellinger penalties (Bubenik et al., 2024).
For matrix-valued densities, the transport plan is no longer a scalar joint density but a positive semidefinite operator on a tensor-product space, and marginalization is implemented by partial trace. One formulation introduces a cost
9
where the additional Frobenius term penalizes rotation between normalized source and target matrix orientations. The unrestricted cost is symmetric, nonnegative, and definite, but does not satisfy the triangle inequality in general; a restricted version yields a genuine metric after taking the square root (Ning et al., 2013). A related noncommutative dynamic formulation replaces scalar continuity equations by
0
and defines a matrix Wasserstein distance
1
This formulation supports convexity, strong duality, and constant-speed geodesics in the matrix setting (Chen et al., 2017).
Quantum variants replace classical couplings by bipartite density operators or bipartite quantum states with fixed marginals. One paper defines
2
for density operators 3 on 4, proves a Kantorovich-type duality theorem, and derives a gradient-type structure
5
for optimal quantum couplings (Caglioti et al., 2021). Another paper studies finite-dimensional density matrices and defines
6
with a distinguished cost
7
equal to the projector onto the antisymmetric subspace. In all dimensions this yields a semidistance, and for qubits its square root is a genuine metric; the associated SWAP-fidelity satisfies
8
The paper interprets 9 as a quantum analogue of a Wasserstein-2 distance on the Bloch ball (Friedland et al., 2021).
In noncommutative geometry, the dual point of view becomes primary. For a spectral triple 00, the paper defines
01
where 02 consists of elements whose evaluations on pure states are 03-Lipschitz for the spectral distance 04. One always has
05
and equality is proved on pure states, on convex segments generated by two pure states, and on the whole state space of 06 (Martinetti, 2012).
6. Multistochastic, causal, geometric, and higher-order developments
The classical constraint of fixing one-dimensional marginals can be replaced by stronger projection constraints. In the multistochastic 07-Monge–Kantorovich problem, one fixes all 08-dimensional coordinate projections of a probability measure on 09 and minimizes or maximizes 10. In the model case
11
with all pairwise projections equal to Lebesgue measure on 12, the paper proves that the minimizer is supported on
13
where 14 is bitwise xor, and that the support is the Sierpiński tetrahedron. It also constructs an explicit dual optimizer for this problem (Gladkov et al., 2018).
A different structural constraint is causality. For filtered Polish spaces 15 and 16, a coupling 17 is causal if the conditional law of the second component up to time 18 is measurable with respect to the source filtration up to time 19. The paper denotes the causal set by
20
and shows that with trivial filtrations one recovers the usual optimal transport problem, whereas in stochastic settings the resulting causal Monge–Kantorovich problems are related to stochastic control, stochastic calculus, martingales, and strong solutions of SDEs (Lassalle, 2013).
On general geodesic spaces, the Monge problem for cost 21 can be reduced to one-dimensional transport problems along transport geodesics. In a non-branching geodesic space, the transport set is partitioned into rays, and under non-degeneracy conditions on the source measure the conditional measures on those rays are continuous or absolutely continuous with respect to one-dimensional Hausdorff measure. This permits construction of a Monge transport map by solving one-dimensional transport problems on each ray and gluing them together. The paper also emphasizes that in this setting 22-cyclical monotonicity is not sufficient for optimality (Bianchini et al., 2011).
The transport problem itself can also be lifted to Wasserstein space. Given probability laws
23
on the space of probability measures over a smooth compact manifold 24, the paper studies
25
Under a Rademacher property for a reference measure on 26, absolute continuity of the source law with respect to that reference measure, and concentration of the source law on absolutely continuous measures on 27, the optimal plan exists, is unique, and is induced by a measurable map
28
An extension to costs of the form 29, with 30 strictly increasing and strictly convex, is also proved (Emami et al., 2024).
Finally, recent quantitative analysis addresses stability and second variation of the quadratic Monge–Kantorovich distance. For non-degenerate Hölder densities on uniformly convex domains, the cited paper proves
31
where 32, and also derives an explicit second-variation formula. If 33, 34, and 35, then
36
This gives a precise Hessian-type formula for the quadratic transport cost around a regular optimal transport configuration (Caja-Lopez et al., 22 May 2026).
These generalizations indicate that the Monge–Kantorovich transportation distance is not restricted to balanced scalar transport on Euclidean domains. The same primal–dual structure persists, with substantial modifications, for random measures, higher-order projection constraints, causal filtrations, matrix-valued densities, quantum states, and noncommutative state spaces.