Parameterized Gromov–Wasserstein Distances
- Parameterized Gromov–Wasserstein distances are families of optimal transport discrepancies that incorporate adjustable parameters in distortion kernels, admissible couplings, and computational relaxations.
- They enable versatile control through objective-side, feasible-set, and computational hierarchy parameterizations, thereby enhancing metric flexibility and invariance.
- These approaches improve computational tractability and support advanced applications ranging from time-varying network models to Gaussian mixture comparisons.
Parameterized Gromov–Wasserstein distances are families of GW-type discrepancies in which a parameter changes the object being optimized, rather than merely changing a numerical solver. In the classical quadratic formulation, for metric measure spaces and , the GW loss of a coupling is
and the distance is
Across the literature, “parameterized” has several non-equivalent meanings: the parameter can modify the distortion kernel, the codomain in which pairwise relations live, the admissible coupling family, the amount of entropic or semidefinite relaxation, the sampled structural resolution, or the underlying object class itself, such as time-indexed or random families of networks (Chowdhury et al., 2021, Bauer et al., 2024, Gómez et al., 26 Sep 2025).
1. Classical framework and the main axes of parameterization
Discrete GW is typically written as a nonconvex quadratic program over a coupling matrix with prescribed marginals, and exact computation is NP-hard; naive evaluation is , while improved implementations remain about (Chowdhury et al., 2021, Tran et al., 13 Feb 2025). This computational difficulty is one reason parameterized variants appear in several forms: some preserve the original distortion but restrict couplings, some preserve the metric meaning but regularize or relax the optimization, and some replace the scalar edge kernel by richer structured data.
A first axis is objective-side parameterization. The -family changes both the outer aggregation exponent and the way distances are transformed before comparison (Arya et al., 2023). The -GW framework replaces real-valued pairwise relations by kernels with values in an arbitrary metric space , so that discrepancies are measured by 0 rather than absolute value on 1 (Bauer et al., 2024). Fused and augmented formulations introduce weights that trade structure against node features, edge features, or raw coordinate alignment (Yang et al., 2023, Demetci et al., 2023).
A second axis is feasible-set parameterization. Quantized GW restricts transport to couplings compatible with pointed partitions and representatives (Chowdhury et al., 2021). Supervised GW restricts admissible joint pairings by inserting 2-entries into the fourth-order cost tensor, thereby forbidding simultaneous activation of incompatible coupling entries (Cang et al., 2024).
A third axis is computational hierarchy parameterization. Entropic GW introduces the regularization parameter 3 (Zhang et al., 2022, Rioux et al., 2023). Sum-of-Squares relaxations are indexed by hierarchy order 4 and induce proxy distances 5 (Tran et al., 13 Feb 2025). Distance-Matrix Wasserstein is indexed by finite-subspace order 6 and further by empirical sample size 7, number of slicing directions 8, and multi-scale weights (Xu et al., 14 May 2026). Lower-bound sliced constructions based on local distance distributions are parameterized by feature/structure balance 9, quadrature size 0, and slicing budget 1 (Piening et al., 4 Aug 2025).
2. Distortion-side parameterizations
The explicit two-parameter family 2 compares 3-th powers of distances and then aggregates the resulting discrepancy in 4. For 5,
6
This family recovers standard GW at 7 and the ultrametric variant at 8; for Euclidean spheres, the special choice 9 is exactly computable, and the optimal coupling is the equatorial coupling (Arya et al., 2023).
A more structural generalization is the 0-Gromov–Wasserstein distance. A 1-network is a measure space endowed with a kernel 2, where 3 is an arbitrary metric space. The corresponding distance is
4
This subsumes classical GW, Wasserstein distance, ultrametric GW, 5-GW, fused GW, fused network GW, spectral GW variants, and dynamic metric-space constructions. The framework proves that 6 is a metric on 7-networks modulo weak isomorphism, and it transfers separability, completeness, and geodesicity from 8 to the induced GW space (Bauer et al., 2024).
Objective-side parameterization also appears as interpolation. Augmented Gromov–Wasserstein is
9
so 0 interpolates between GW and CO-Optimal Transport. The intended effect is to control how rigid or invariant the comparison should be: 1 approaches GW’s isometry-invariant structural regime, while 2 approaches a more coordinate- and feature-aware regime. The formulation admits solutions, converges to COOT as 3 and to GW as 4, and satisfies a relaxed triangle inequality (Demetci et al., 2023).
3. Coupling restrictions, quantization, and supervised admissibility
Quantized Gromov–Wasserstein parameterizes GW by a pointed partition
5
where each block 6 has representative 7. The admissible couplings are the quantization couplings
8
combining a global representative-level coupling 9 with local block couplings 0. The resulting distance is
1
Because 2, qGW is an upper-bound family for ordinary GW. At the same time, it is a genuine metric on finite pointed mm-spaces up to pointed isomorphism. The parameterization is by the number of blocks 3, the partitions themselves, the representatives, and the local matching rule (Chowdhury et al., 2021).
The same paper develops quantitative control of this restriction. Quantized eccentricity
4
governs the approximation error between the full space and its 5-point quantization, and Theorem 5 states that if every block has diameter at most 6, then
7
Algorithmically, qGW performs a GW solve on representatives and local one-dimensional optimal transport problems around anchors. The expected complexity is
8
and choosing 9 yields iterative cost 0. The formulation was demonstrated at scales containing over 1M points (Chowdhury et al., 2021).
Supervised GW imposes a different coupling-space parameterization. It threshold-cuts the fourth-order cost tensor by
2
so 3 is a hard tolerance on distance distortion. An 4-entry means that 5 and 6 cannot both be positive. The induced fourth-order exclusions are reduced to entrywise zero constraints on 7 through a graph construction and a minimal-vertex-cover heuristic. The approximate entropic objective adds
8
where 9 encourages maximal transported mass and 0 regularizes the solver. When the tensor has no 1-entries, the formulation degenerates to PGW, and in the balanced case to ordinary GW (Cang et al., 2024).
4. Regularized, relaxed, and proxy hierarchies
The entropic family
2
is the most prominent regularized parameterization. Here 3 controls the trade-off between the quadratic GW objective and KL smoothing. The paper on duality and sample complexity derives a variational representation over an auxiliary matrix 4, turning EGW into a family of entropic OT problems with cost 5. This yields an explicit dependence on 6, proves approximation and continuity results as 7, and establishes two-sample rates 8 for standard GW and 9 for EGW (Zhang et al., 2022).
A companion algorithmic analysis studies the same quadratic entropic family through the variational objective
0
Its gradient is
1
where 2 is the EOT coupling. The paper proves 3-smoothness, weak convexity, and strict convexity under the quantitative condition
4
and gives accelerated gradient methods with Sinkhorn-based inexact oracle guarantees. It also proves that stationary points of the EGW variational problem converge to stationary points of the unregularized variational GW problem as 5 (Rioux et al., 2023).
Semidefinite parameterization appears in the Sum-of-Squares hierarchy. For discrete GW, the hierarchy order 6 indexes reduced Schmüdgen-type and Putinar-type SDPs whose values are lower bounds to the exact discrete GW optimum. The paper defines
7
a computable order-8 proxy for the distortion distance. These 9 are pseudo-metrics, satisfy the triangle inequality via an SOS analogue of the gluing lemma, and converge upward to the exact value, with convergence rate 0 (Tran et al., 13 Feb 2025).
Another hierarchy is Distance-Matrix Wasserstein: 1 where 2 is the law of the random 3-point distance matrix induced by i.i.d. sampling from 4. The parameter 5 is the sampled subspace size. The paper proves
6
and
7
so 8 as 9. Sliced and multi-scale variants introduce further parameters 00, 01, and weights 02, and for 03 the resulting sliced multi-scale dissimilarities yield positive-definite exponential kernels (Xu et al., 14 May 2026).
A related lower-bound family starts from Mémoli’s third lower bound and then slices only after embedding local distance distributions into Euclidean space by quadrature-sampled quantiles. For 04, the paper defines 05 and 06, with parameters 07 for structure/feature balance, quadrature size 08, positive weights 09, and number of slicing directions 10. These quantities are pseudo-metrics, bound FGW from below in the exact-quadrature regime, and interpolate between sliced Wasserstein and TLB-type structure comparison while preserving isometry invariance at the structural level (Piening et al., 4 Aug 2025).
5. Parameterized families of networks
The most literal formalization of a parameterized GW distance is the pm-net framework. A parameterized measure network is a quintuple
11
where 12 assigns a bounded measurable kernel 13 to each parameter 14. This covers time-varying metrics, time-varying weighted graphs, heat-kernel families, random graph models, and random metric-space models (Gómez et al., 26 Sep 2025).
When both objects share the same parameter space 15, the standard cost structure is
16
and the parameterized GW distance is the infimum of this quantity over node couplings 17. The same node coupling must explain the whole family of kernels, which is the central modeling constraint (Gómez et al., 26 Sep 2025).
When parameter spaces differ, a second coupling 18 is introduced: 19 This is a genuine two-level OT problem: 20 aligns nodes, 21 aligns parameters. The paper proves that the resulting constructions are pseudometrics, that optimal couplings exist, and that distance zero is equivalent to an appropriate isomorphism notion based on stabilization and structure-preserving maps (Gómez et al., 26 Sep 2025).
The same framework connects back to earlier theories. For fixed parameter space and 22, the distance is exactly a 23-GW distance with 24. It also yields a lower bound
25
where 26 is the distribution of the parameter slices 27 in classical GW space. A further lower bound uses the distribution of global edge-weight distributions, and in random graph models this specializes to stable control by the distribution of total numbers of edges (Gómez et al., 26 Sep 2025).
6. Structured data, graph-specific variants, and model-based parameterizations
Fused Network Gromov–Wasserstein introduces an explicit three-way graph objective
28
where the local cost combines node-feature discrepancy, edge-feature discrepancy, and structural discrepancy with weights 29, 30, and 31, respectively. In the discrete case, this is a quadratic OT problem over the node coupling 32, optimized by Frank–Wolfe. The formulation is a relaxed metric: it satisfies
33
with exact triangle inequality only when 34. Its parameters explicitly distribute modeling emphasis across node attributes, edge attributes, and graph structure (Yang et al., 2023).
Augmented GW belongs to the same class of structure-feature interpolations but is motivated by invariance control. By blending GW with COOT through 35, it makes the distance less permissive than pure GW and allows prior information to be injected through the feature alignment term. The paper’s experiments show that supervision on features improves sample alignment and supervision on samples improves feature alignment, reflecting the joint optimization over sample and feature couplings (Demetci et al., 2023).
A different parameterization compresses the underlying measure space itself. For Gaussian mixture models
36
Mixture Gromov–Wasserstein is
37
This is GW on the compressed component space, with Gaussian components as atoms and 38 as the intra-space geometry. Embedded Wasserstein and mixture embedded Wasserstein add optimization over isometric embeddings 39, with 40 on a Stiefel manifold, thereby supporting comparisons across different Euclidean dimensions and enabling recovery of transport plans between GMMs (Salmona et al., 2023).
A cautionary special case comes from one-dimensional deterministic matching. For sorted point sets 41, 42 and powered costs 43, the induced permutation problem
44
is not, in general, solved by the identity or anti-identity permutation once 45. Thus even in one dimension, GW-type deterministic correspondences need not be monotone, and “sorted-to-sorted” or “sorted-to-reversed” is not a general theory (Beinert et al., 2022).
7. Statistical regimes, computational practice, and conceptual boundaries
Parameterized GW constructions differ not only mathematically but also statistically. For quadratic Euclidean GW, empirical convergence under compact support is now complemented by unbounded-support results: under finite polynomial moments, the plug-in estimator 46 of 47 attains the same benchmark rate as in the compactly supported case, namely
48
up to the stated logarithmic factors, and matching minimax lower bounds hold up to logarithms. The proof is organized through a parameterized OT representation
49
which is described as penalized Wasserstein alignment and includes GW itself via the matrix parameter 50 (Kato et al., 6 Aug 2025).
In practice, nonconvexity makes algorithmic parameters part of the usable definition. In assignment- and QAP-style settings, the same paper trail shows three distinct operational levers: entropic regularization 51 in EGW, structure-feature weighting 52 in FGW, and the number of random feasible starts 53 in multi-initialization GW. On the reported CQAP instances, GW-MultiInit achieved the best approximate objective values, EGW with 54 gave the best regularized compromise, and FGW with 55 was the best among the tested fused settings, while initialization itself emerged as a first-class parameter of the nonconvex search (Seyedi et al., 4 Sep 2025).
Across the literature, a sharp conceptual boundary separates metric generalizations from tractable proxies. 56-GW and qGW define bona fide metrics on quotient or pointed spaces (Bauer et al., 2024, Chowdhury et al., 2021). SOS distances 57, DMW58, and sliced TLB-based constructions are instead lower-bound or proxy families that trade exactness for tractability (Tran et al., 13 Feb 2025, Xu et al., 14 May 2026, Piening et al., 4 Aug 2025). Entropic GW is a regularized discrepancy whose debiased form is often used in practice because the raw entropic objective does not vanish on isomorphic spaces (Rioux et al., 2023).
The unifying theme is that parameterization changes where complexity is placed. A parameter may enter the distortion kernel, the relation codomain, the admissible couplings, the mass constraints, the relaxation hierarchy, the sampled structural resolution, or the object class being compared. Consequently, “parameterized Gromov–Wasserstein distance” denotes not a single construction but a research program: to expose interpretable degrees of freedom in GW while preserving, approximating, or selectively sacrificing metric structure, invariance, and computational feasibility.