Papers
Topics
Authors
Recent
Search
2000 character limit reached

Convex Distance Operator Transport (CDOT)

Updated 15 July 2026
  • CDOT is a convex optimal transport framework that aligns probability distributions by jointly preserving feature correspondence and intrinsic geometric structure.
  • It replaces pairwise distance matching with an aggregated operator-level regularization, offering robustness and global optimality in heterogeneous settings.
  • The method yields an explicit coupling with favorable empirical performance compared to GW-based methods, particularly in domains like graphs, manifolds, and connectomes.

Convex Distance Operator Transport (CDOT) is a convex optimal transport framework for aligning probability distributions across heterogeneous domains by jointly preserving feature correspondence and intrinsic geometric structure. It was introduced as the first convex optimal transport framework on population-level attributed compact metric-measure spaces that simultaneously aligns features and intrinsic geometry while outputting an explicit coupling. Its defining move is to replace Gromov--Wasserstein-style pairwise distance matching by an operator-level regularization built from distance operators and conditional expectation operators, yielding a discrepancy that is a valid pseudometric on the space of attributed compact metric-measure spaces (Chung et al., 1 Jun 2026).

1. Formal setting and motivation

CDOT is formulated for attributed compact metric-measure spaces

X=(X,dX,PX,fX),Y=(Y,dY,PY,fY),\mathfrak{X}=(\mathcal{X},d_{\mathcal{X}},\mathbb{P}_X,f_{\mathcal{X}}), \qquad \mathfrak{Y}=(\mathcal{Y},d_{\mathcal{Y}},\mathbb{P}_Y,f_{\mathcal{Y}}),

where (X,dX)(\mathcal{X},d_{\mathcal{X}}) and (Y,dY)(\mathcal{Y},d_{\mathcal{Y}}) are compact metric spaces, PX,PY\mathbb{P}_X,\mathbb{P}_Y are probability measures, and fX,fYf_{\mathcal{X}},f_{\mathcal{Y}} are feature maps into a compact feature space MRk\mathcal{M}\subset\mathbb{R}^k. The optimization variable is a coupling

πΠ(PX,PY),\pi\in \Pi(\mathbb{P}_X,\mathbb{P}_Y),

and the goal is to align the two spaces by jointly preserving feature correspondence across domains and intrinsic geometry within each domain (Chung et al., 1 Jun 2026).

This problem class is motivated by heterogeneous settings in which a direct cross-domain ground cost c(x,y)c(x,y) is unavailable or unreliable. Graphs, manifolds, connectomes, shapes, and point clouds with unrelated coordinates fall into this regime. Classical Wasserstein transport presupposes a meaningful cross-space cost, whereas Gromov--Wasserstein (GW) circumvents that requirement by comparing within-space distances,

RGW,2(π)=Eππ[dX(X,X)dY(Y,Y)2].\mathcal{R}_{\mathrm{GW},2}(\pi) = \mathbb{E}_{\pi\otimes\pi} \big[ |d_{\mathcal{X}}(X,X')-d_{\mathcal{Y}}(Y,Y')|^2 \big].

Fused GW augments this with a feature term, but the dependence on ππ\pi\otimes\pi makes GW and FGW non-convex. According to the CDOT formulation, the source of difficulty is not only non-convexity in an abstract sense, but the rigidity of edge-by-edge pairwise distance consistency. CDOT instead aligns aggregated distance profiles, which the paper states improves robustness to local geometric variations (Chung et al., 1 Jun 2026).

A recurring misconception is that CDOT is intended to recover exact measure-preserving isometries. The formulation is weaker: it preserves geometry at the operator level. This distinction is central to both its convexity and its scope.

2. Operator formulation

The feature term uses the squared Euclidean discrepancy

(X,dX)(\mathcal{X},d_{\mathcal{X}})0

The geometric term is defined through linear operators on (X,dX)(\mathcal{X},d_{\mathcal{X}})1 spaces. For a compact metric-measure space (X,dX)(\mathcal{X},d_{\mathcal{X}})2, the distance operator is

(X,dX)(\mathcal{X},d_{\mathcal{X}})3

For a coupling (X,dX)(\mathcal{X},d_{\mathcal{X}})4, the conditional expectation operator is

(X,dX)(\mathcal{X},d_{\mathcal{X}})5

The structural regularization is the Hilbert--Schmidt commutator penalty

(X,dX)(\mathcal{X},d_{\mathcal{X}})6

The full CDOT objective is

(X,dX)(\mathcal{X},d_{\mathcal{X}})7

and the CDOT problem is

(X,dX)(\mathcal{X},d_{\mathcal{X}})8

The term (X,dX)(\mathcal{X},d_{\mathcal{X}})9 expresses approximate intertwining of the intrinsic distance operators through the coupling-induced conditional expectation operator (Chung et al., 1 Jun 2026).

The paper also gives a kernel representation of the structural term: (Y,dY)(\mathcal{Y},d_{\mathcal{Y}})0 where

(Y,dY)(\mathcal{Y},d_{\mathcal{Y}})1

and (Y,dY)(\mathcal{Y},d_{\mathcal{Y}})2. This makes explicit that CDOT compares conditional expected distance profiles rather than all pairwise distances. The geometric information is therefore aggregated before comparison, rather than tensorized as in GW (Chung et al., 1 Jun 2026).

3. Convexity, pseudometric structure, and relation to Gromov--Wasserstein

The central optimization theorem states that the CDOT problem admits a minimizer and that the objective is convex in the coupling (Y,dY)(\mathcal{Y},d_{\mathcal{Y}})3. The reason is structural: (Y,dY)(\mathcal{Y},d_{\mathcal{Y}})4 is linear in (Y,dY)(\mathcal{Y},d_{\mathcal{Y}})5, the map (Y,dY)(\mathcal{Y},d_{\mathcal{Y}})6 is affine, the commutator (Y,dY)(\mathcal{Y},d_{\mathcal{Y}})7 is affine in (Y,dY)(\mathcal{Y},d_{\mathcal{Y}})8, and the squared Hilbert--Schmidt norm is convex. Convexity therefore holds on the transport polytope (Y,dY)(\mathcal{Y},d_{\mathcal{Y}})9 itself, not merely after relaxation (Chung et al., 1 Jun 2026).

The induced discrepancy is

PX,PY\mathbb{P}_X,\mathbb{P}_Y0

For each PX,PY\mathbb{P}_X,\mathbb{P}_Y1, this is a pseudometric on the space of attributed compact metric-measure spaces: non-negativity, identity on the diagonal, symmetry, and triangle inequality all hold. The triangle inequality uses gluing of couplings, Minkowski for the feature term, the operator contraction property PX,PY\mathbb{P}_X,\mathbb{P}_Y2, and composition of conditional expectation operators (Chung et al., 1 Jun 2026).

The reason it is only a pseudometric is equally important. Zero discrepancy does not imply that the two spaces are identical or measure-preserving isometric. It implies only the existence of a coupling that preserves both the feature maps and the distance operators at the operator level. The authors describe this as closer to a fractional structural equivalence. This weaker notion of equivalence is one of the main conceptual tradeoffs of CDOT.

The relation to GW is sharpened by the dispersion decomposition. Defining

PX,PY\mathbb{P}_X,\mathbb{P}_Y3

the paper proves

PX,PY\mathbb{P}_X,\mathbb{P}_Y4

This identifies the dispersion term as the geometric source of the extra rigidity and non-convexity in GW. CDOT keeps only the convex structural component PX,PY\mathbb{P}_X,\mathbb{P}_Y5, while GW also penalizes conditional variance and thus favors concentrated, nearly deterministic couplings (Chung et al., 1 Jun 2026).

Two special cases are immediate. When PX,PY\mathbb{P}_X,\mathbb{P}_Y6, CDOT reduces to feature-based OT. When PX,PY\mathbb{P}_X,\mathbb{P}_Y7, it becomes pure operator-geometry matching,

PX,PY\mathbb{P}_X,\mathbb{P}_Y8

4. Finite-sample formulation and optimization

For samples

PX,PY\mathbb{P}_X,\mathbb{P}_Y9

the empirical measures are

fX,fYf_{\mathcal{X}},f_{\mathcal{Y}}0

The feature cost matrix and normalized distance matrices are

fX,fYf_{\mathcal{X}},f_{\mathcal{Y}}1

fX,fYf_{\mathcal{X}},f_{\mathcal{Y}}2

The empirical objective becomes

fX,fYf_{\mathcal{X}},f_{\mathcal{Y}}3

fX,fYf_{\mathcal{X}},f_{\mathcal{Y}}4

fX,fYf_{\mathcal{X}},f_{\mathcal{Y}}5

Hence empirical CDOT is a convex quadratic program over the transport polytope (Chung et al., 1 Jun 2026).

The paper uses a globally convergent Frank--Wolfe algorithm. Its gradient is

fX,fYf_{\mathcal{X}},f_{\mathcal{Y}}6

At iteration fX,fYf_{\mathcal{X}},f_{\mathcal{Y}}7, the linear minimization oracle is

fX,fYf_{\mathcal{X}},f_{\mathcal{Y}}8

followed by

fX,fYf_{\mathcal{X}},f_{\mathcal{Y}}9

A lazy-gradient variant is possible because the gradient map is affine in MRk\mathcal{M}\subset\mathbb{R}^k0: MRk\mathcal{M}\subset\mathbb{R}^k1 This reduces gradient-update cost from cubic to quadratic, although the per-iteration complexity is still dominated by the linear OT subproblem (Chung et al., 1 Jun 2026).

The convergence statement is the standard MRk\mathcal{M}\subset\mathbb{R}^k2 Frank--Wolfe primal error rate,

MRk\mathcal{M}\subset\mathbb{R}^k3

and the non-asymptotic population risk bound decomposes into optimization and statistical errors: MRk\mathcal{M}\subset\mathbb{R}^k4 Under i.i.d. sampling and MRk\mathcal{M}\subset\mathbb{R}^k5, the paper states risk consistency. The assumptions are non-atomicity, no ties, Lipschitz feature maps, and a stability condition on an optimal population coupling with conditional Lipschitz continuity in MRk\mathcal{M}\subset\mathbb{R}^k6 (Chung et al., 1 Jun 2026).

5. Empirical behavior, practical scope, and limitations

The empirical evaluation covers synthetic point clouds, brain connectomes, and graph classification. On synthetic MRk\mathcal{M}\subset\mathbb{R}^k7D clustered point clouds in MRk\mathcal{M}\subset\mathbb{R}^k8, with four square regions, region-label features, MRk\mathcal{M}\subset\mathbb{R}^k9, and πΠ(PX,PY),\pi\in \Pi(\mathbb{P}_X,\mathbb{P}_Y),0, CDOT achieves the best reported MSE across all sample sizes. For example, at πΠ(PX,PY),\pi\in \Pi(\mathbb{P}_X,\mathbb{P}_Y),1, the reported MSE is πΠ(PX,PY),\pi\in \Pi(\mathbb{P}_X,\mathbb{P}_Y),2 for CDOT versus πΠ(PX,PY),\pi\in \Pi(\mathbb{P}_X,\mathbb{P}_Y),3 for FGW and πΠ(PX,PY),\pi\in \Pi(\mathbb{P}_X,\mathbb{P}_Y),4 for EFGW; at πΠ(PX,PY),\pi\in \Pi(\mathbb{P}_X,\mathbb{P}_Y),5, the reported values are πΠ(PX,PY),\pi\in \Pi(\mathbb{P}_X,\mathbb{P}_Y),6, πΠ(PX,PY),\pi\in \Pi(\mathbb{P}_X,\mathbb{P}_Y),7, and πΠ(PX,PY),\pi\in \Pi(\mathbb{P}_X,\mathbb{P}_Y),8, respectively. The appendix also reports strong sensitivity of EFGW to entropic regularization πΠ(PX,PY),\pi\in \Pi(\mathbb{P}_X,\mathbb{P}_Y),9, whereas CDOT avoids that tuning parameter (Chung et al., 1 Jun 2026).

On OASIS-3 brain connectome matching, the dataset consists of 696 structural brain networks with 170 nodes and 6 anatomical labels. With diffusion distance, the reported CDOT score is c(x,y)c(x,y)0, versus c(x,y)c(x,y)1 for FGW and c(x,y)c(x,y)2 for best-tuned EFGW. With geodesic distance, FGW is reported slightly better, with FGW c(x,y)c(x,y)3 and CDOT c(x,y)c(x,y)4. The interpretation given in the paper is that geodesic distance is rigid and preserves pairwise contrast, which helps FGW, while diffusion distance smooths local pairwise structure in a way that suits CDOT’s aggregated-operator geometry (Chung et al., 1 Jun 2026).

For graph classification on MUTAG, IMDB-BINARY, PROTEINS, NCI1, and ENZYMES, CDOT is reported best on all listed datasets. The reported accuracies are c(x,y)c(x,y)5 on MUTAG, c(x,y)c(x,y)6 on IMDB-BINARY, c(x,y)c(x,y)7 on PROTEINS, c(x,y)c(x,y)8 on NCI1, and c(x,y)c(x,y)9 on ENZYMES. Even on IMDB-BINARY, where there are no node features and RGW,2(π)=Eππ[dX(X,X)dY(Y,Y)2].\mathcal{R}_{\mathrm{GW},2}(\pi) = \mathbb{E}_{\pi\otimes\pi} \big[ |d_{\mathcal{X}}(X,X')-d_{\mathcal{Y}}(Y,Y')|^2 \big].0, the paper reports that CDOT still performs best, which suggests that the operator geometry alone can be useful (Chung et al., 1 Jun 2026).

The formulation is most appropriate when source and target lie in heterogeneous spaces, geometry matters, exact one-to-one isometry is too rigid, an explicit transport plan is required, and global optimality is preferable to local stationary behavior. The same source also states its limitations: zero discrepancy induces a weaker equivalence notion than GW, scalability is currently moderate because each Frank--Wolfe iteration still solves a standard OT subproblem, the metric structure must be predefined rather than learned jointly, and the exact zero-discrepancy equivalence class remains open (Chung et al., 1 Jun 2026).

Several earlier transport frameworks were described in the supplied literature as directly relevant to a CDOT-like viewpoint, even though they do not use the term. They clarify which parts of CDOT are genuinely new and which reflect broader developments in convex or structured transport.

Related work Shared element with CDOT Key distinction
CROT (Nielsen et al., 2018) OT over structured objects via a ground discrepancy on conditionals Transports latent marginals rather than metric-measure operators
Semi-discrete DOT (Gu et al., 2013) Convex variational structure, operator-like mass-matching map Euclidean quadratic power-diagram setting
Generic convex-regularized OT (Tsutsui, 2020) Convex objective design, Bregman geometry, generalized Sinkhorn Regularizes couplings rather than aligning distance operators
Unbalanced OT (Chizat et al., 2015) Convex dynamic/static formulations under operator constraints Handles transport with mass creation/destruction
Cone-compatible ordered OT (Luo et al., 3 Jun 2026) Compatibility of convex costs with structured transport operators Requires cone-chain support and partial-order structure
Dynamic OT on surfaces (Chen et al., 10 Jun 2025) Explicit operator discretization of convex dynamic OT Computes Benamou--Brenier RGW,2(π)=Eππ[dX(X,X)dY(Y,Y)2].\mathcal{R}_{\mathrm{GW},2}(\pi) = \mathbb{E}_{\pi\otimes\pi} \big[ |d_{\mathcal{X}}(X,X')-d_{\mathcal{Y}}(Y,Y')|^2 \big].1 on meshes rather than heterogeneous-space CDOT

"On The Chain Rule Optimal Transport Distance" defines

RGW,2(π)=Eππ[dX(X,X)dY(Y,Y)2].\mathcal{R}_{\mathrm{GW},2}(\pi) = \mathbb{E}_{\pi\otimes\pi} \big[ |d_{\mathcal{X}}(X,X')-d_{\mathcal{Y}}(Y,Y')|^2 \big].2

which transports latent marginals while using a discrepancy between conditionals as the ground cost. In the supplied material this is identified as a clean prototype of transport over structured objects. Its most CDOT-relevant theorem is the upper-bound principle

RGW,2(π)=Eππ[dX(X,X)dY(Y,Y)2].\mathcal{R}_{\mathrm{GW},2}(\pi) = \mathbb{E}_{\pi\otimes\pi} \big[ |d_{\mathcal{X}}(X,X')-d_{\mathcal{Y}}(Y,Y')|^2 \big].3

whenever the ground distance RGW,2(π)=Eππ[dX(X,X)dY(Y,Y)2].\mathcal{R}_{\mathrm{GW},2}(\pi) = \mathbb{E}_{\pi\otimes\pi} \big[ |d_{\mathcal{X}}(X,X')-d_{\mathcal{Y}}(Y,Y')|^2 \big].4 is jointly convex, together with metric inheritance when RGW,2(π)=Eππ[dX(X,X)dY(Y,Y)2].\mathcal{R}_{\mathrm{GW},2}(\pi) = \mathbb{E}_{\pi\otimes\pi} \big[ |d_{\mathcal{X}}(X,X')-d_{\mathcal{Y}}(Y,Y')|^2 \big].5 is a metric (Nielsen et al., 2018).

"Variational Principles for Minkowski Type Problems, Discrete Optimal Transport, and Discrete Monge-Ampere Equations" develops the semi-discrete convex potential

RGW,2(π)=Eππ[dX(X,X)dY(Y,Y)2].\mathcal{R}_{\mathrm{GW},2}(\pi) = \mathbb{E}_{\pi\otimes\pi} \big[ |d_{\mathcal{X}}(X,X')-d_{\mathcal{Y}}(Y,Y')|^2 \big].6

whose gradient induces power-diagram cells RGW,2(π)=Eππ[dX(X,X)dY(Y,Y)2].\mathcal{R}_{\mathrm{GW},2}(\pi) = \mathbb{E}_{\pi\otimes\pi} \big[ |d_{\mathcal{X}}(X,X')-d_{\mathcal{Y}}(Y,Y')|^2 \big].7 satisfying prescribed mass constraints. The energy RGW,2(π)=Eππ[dX(X,X)dY(Y,Y)2].\mathcal{R}_{\mathrm{GW},2}(\pi) = \mathbb{E}_{\pi\otimes\pi} \big[ |d_{\mathcal{X}}(X,X')-d_{\mathcal{Y}}(Y,Y')|^2 \big].8 is convex, strictly convex on the normalized subspace, and its Hessian is a sparse graph-Laplacian-type matrix derived from face integrals. In the supplied interpretation, this is a foundational semi-discrete convex transport blueprint that later CDOT-style viewpoints can be read against (Gu et al., 2013).

"Optimal transport problems regularized by generic convex functions" replaces entropic regularization by a generic strictly convex smooth RGW,2(π)=Eππ[dX(X,X)dY(Y,Y)2].\mathcal{R}_{\mathrm{GW},2}(\pi) = \mathbb{E}_{\pi\otimes\pi} \big[ |d_{\mathcal{X}}(X,X')-d_{\mathcal{Y}}(Y,Y')|^2 \big].9, studies the resulting Bregman geometry, and derives the generalized stationarity condition

ππ\pi\otimes\pi0

The quadratic case

ππ\pi\otimes\pi1

yields the sparse thresholded form

ππ\pi\otimes\pi2

This is not CDOT by name, but it is a general convex-regularized OT template with a clear operator-geometric interpretation (Tsutsui, 2020).

Other adjacent developments broaden the contextual field. Unbalanced OT constructs convex dynamic and static formulations linked by a generalized continuity equation with source

ππ\pi\otimes\pi3

and semi-couplings on weighted points, showing how convex operator constraints can define distances beyond balanced transport (Chizat et al., 2015). Cone-compatible Monge geometry shows that a convex quadratic cost ππ\pi\otimes\pi4 supports a high-dimensional monotone transport operator exactly when

ππ\pi\otimes\pi5

thereby identifying compatibility conditions under which convex costs and structured order yield closed-form transport on cone chains (Luo et al., 3 Jun 2026). Dynamic OT on surfaces, finally, reformulates Benamou--Brenier transport on triangulated meshes as a linear second-order cone program built from explicit differential and interpolation operators, which is closely aligned with an operator-based convex transport perspective even though it solves ππ\pi\otimes\pi6 rather than heterogeneous-space CDOT (Chen et al., 10 Jun 2025).

Taken together, these works suggest that CDOT occupies a specific point within a broader research program: transport defined or approximated through convex structure, higher-level objects, or explicit operators, but specialized to the heterogeneous-domain alignment problem by the commutator penalty

ππ\pi\otimes\pi7

That specialization is what distinguishes CDOT from both GW-style metric alignment and generic convex OT regularization (Chung et al., 1 Jun 2026).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Convex Distance Operator Transport (CDOT).