Papers
Topics
Authors
Recent
Search
2000 character limit reached

Additive Optimal Transport Regression

Updated 15 December 2025
  • Additive optimal transport regression is a framework that replaces linear shifts with compositions of geodesic transport maps to model responses in general metric spaces.
  • It decomposes multivariate Euclidean predictors into univariate transport maps applied to the Fréchet mean, addressing the curse of dimensionality.
  • The ADOPT model ensures estimation consistency via a transport backfitting algorithm and has practical applications in SPD matrices and fMRI connectivity analysis.

Additive optimal transport regression is a model class for analyzing regression problems where the predictor variables are Euclidean and the response lies in a general geodesic metric space. The central innovation is the replacement of additive linear shifts (which require vector space structure) with compositions of optimal geodesic transport maps, thus enabling additive modeling for manifold- or distribution-valued responses without embedding or vectorization. The most developed framework for this approach is ADOPT (Additive Optimal Transport Regression), which systematically extends additive regression concepts through the geometry of geodesic metric spaces and optimal transport (Song et al., 8 Dec 2025).

1. Mathematical Formulation and Problem Setting

Consider observations (Xi,Yi)(X_i, Y_i), i=1,,ni=1,\dots, n, with Xi=(Xi1,...,Xip)RpX_i = (X_{i1}, ..., X_{ip}) \in \mathbb{R}^p—Euclidean predictors—and YiY_i valued in a bounded, separable, uniquely geodesic metric space (M,d)(\mathcal{M}, d). Unlike classical regression, where the conditional mean E[YX=x]\mathbb{E}[Y\,|\,X=x] is well defined, the lack of vector addition in general metric spaces is overcome by using the Fréchet conditional mean: μ(x):=argminmM  E[d2(Y,m)X=x].\mu_\oplus(x) := \arg\min_{m\in\mathcal M}\; \mathbb E\big[d^2(Y, m)\mid X = x\big]. The objective is to model μ(x)\mu_\oplus(x) flexibly for multivariate (X1,...,Xp)(X_1, ..., X_p) while maintaining interpretability and avoiding the curse of dimensionality (Song et al., 8 Dec 2025).

2. Additive Structure via Geodesic Transport

In Euclidean additive models, the regression function is parameterized as m(x)β0+jgj(xj)m(x) \approx \beta_0 + \sum_j g_j(x_j). In the ADOPT paradigm, addition is replaced by composition of transport maps, each encoding the influence of i=1,,ni=1,\dots, n0 on the response in the geometry of i=1,,ni=1,\dots, n1.

  • Define i=1,,ni=1,\dots, n2 as the (unconditional) Fréchet mean.
  • For each predictor i=1,,ni=1,\dots, n3, introduce a transport-valued function i=1,,ni=1,\dots, n4, with i=1,,ni=1,\dots, n5 a geodesic transport map on i=1,,ni=1,\dots, n6.
  • The combined transport is given by composition i=1,,ni=1,\dots, n7, i.e., i=1,,ni=1,\dots, n8, acting on i=1,,ni=1,\dots, n9 as Xi=(Xi1,...,Xip)RpX_i = (X_{i1}, ..., X_{ip}) \in \mathbb{R}^p0.
  • The ADOPT model is:

Xi=(Xi1,...,Xip)RpX_i = (X_{i1}, ..., X_{ip}) \in \mathbb{R}^p1

where Xi=(Xi1,...,Xip)RpX_i = (X_{i1}, ..., X_{ip}) \in \mathbb{R}^p2 is a small random perturbation map.

This construction enables interpretability: each Xi=(Xi1,...,Xip)RpX_i = (X_{i1}, ..., X_{ip}) \in \mathbb{R}^p3 explicitly describes how Xi=(Xi1,...,Xip)RpX_i = (X_{i1}, ..., X_{ip}) \in \mathbb{R}^p4 alters the conditional mean via geodesic transport from the Fréchet mean along the geometry of Xi=(Xi1,...,Xip)RpX_i = (X_{i1}, ..., X_{ip}) \in \mathbb{R}^p5.

3. Optimal Geodesic Transports and Their Composition

In one-dimensional Xi=(Xi1,...,Xip)RpX_i = (X_{i1}, ..., X_{ip}) \in \mathbb{R}^p6 Wasserstein space, the optimal transport from Xi=(Xi1,...,Xip)RpX_i = (X_{i1}, ..., X_{ip}) \in \mathbb{R}^p7 to Xi=(Xi1,...,Xip)RpX_i = (X_{i1}, ..., X_{ip}) \in \mathbb{R}^p8 is given by quantile mapping. For general geodesic spaces, ADOPT posits the existence of a ternary map Xi=(Xi1,...,Xip)RpX_i = (X_{i1}, ..., X_{ip}) \in \mathbb{R}^p9 such that

YiY_i0

which moves YiY_i1 along the unique geodesic from YiY_i2 to YiY_i3. The sum of transports is defined by map composition: YiY_i4 This composition is associative and produces the effect of sequentially transporting the base point YiY_i5 according to each coordinate's effect.

4. Estimation via Transport Backfitting Algorithm

Estimation in ADOPT is based on a transport backfitting scheme, analogous to classical additive model backfitting but operating on transport maps and Fréchet means.

Algorithmic steps:

  • Initialization: Compute YiY_i6. Set YiY_i7 for all YiY_i8.
  • Iterative update for each coordinate YiY_i9:
    • Form partial transport–residuals by undoing the effect of other (M,d)(\mathcal{M}, d)0 and of the observed (M,d)(\mathcal{M}, d)1 relative to the overall mean.
    • Fit (M,d)(\mathcal{M}, d)2 by local Fréchet regression of these residuals against (M,d)(\mathcal{M}, d)3.
    • Center (normalize) (M,d)(\mathcal{M}, d)4 and use it to update (M,d)(\mathcal{M}, d)5 via geodesic transport.
  • Update fitted responses (M,d)(\mathcal{M}, d)6 and iterate until (M,d)(\mathcal{M}, d)7.

Every (M,d)(\mathcal{M}, d)8 is estimated univariately, which mitigates the curse of dimensionality and makes the procedure practical for moderate to large (M,d)(\mathcal{M}, d)9.

5. Theoretical Guarantees

Under the key assumptions below, the transport backfitting estimates converge to the true transport maps:

  • Kernel regularity: smoothness of predictor densities and uniqueness/stability of Fréchet minimizers [(A1)-(A3)].
  • Small perturbation error and Lipschitz property of E[YX=x]\mathbb{E}[Y\,|\,X=x]0 [(A4)-(A5)].
  • Transport maps as Fréchet-type perturbations [(A6)].

Main result: For any fixed E[YX=x]\mathbb{E}[Y\,|\,X=x]1 and anchor E[YX=x]\mathbb{E}[Y\,|\,X=x]2, the fitted E[YX=x]\mathbb{E}[Y\,|\,X=x]3 satisfies

E[YX=x]\mathbb{E}[Y\,|\,X=x]4

uniformly over compact predictor domains, as E[YX=x]\mathbb{E}[Y\,|\,X=x]5, E[YX=x]\mathbb{E}[Y\,|\,X=x]6, E[YX=x]\mathbb{E}[Y\,|\,X=x]7 (Song et al., 8 Dec 2025).

6. Applications and Examples

The ADOPT framework accommodates responses including:

  • SPD (symmetric positive-definite) matrix regression using the log-Cholesky metric:

E[YX=x]\mathbb{E}[Y\,|\,X=x]8

where E[YX=x]\mathbb{E}[Y\,|\,X=x]9 is the Cholesky decomposition, and μ(x):=argminmM  E[d2(Y,m)X=x].\mu_\oplus(x) := \arg\min_{m\in\mathcal M}\; \mathbb E\big[d^2(Y, m)\mid X = x\big].0 is split into diagonal/off-diagonal parts.

  • Correlation matrix analysis from fMRI data: Used to regress μ(x):=argminmM  E[d2(Y,m)X=x].\mu_\oplus(x) := \arg\min_{m\in\mathcal M}\; \mathbb E\big[d^2(Y, m)\mid X = x\big].1 correlation matrices of brain connectivity on biomedical predictors (e.g., cerebrospinal amyloid-β, diagnostic stage, p-Tau), revealing interpretable patterns in network connectivity.
  • Metrics with explicit geodesics and transports (e.g., the affine-invariant metric for SPD matrices and Wasserstein metrics for probability distributions).
  • The method is intrinsic and does not require embedding or vectorizing the metric space responses.

7. Advantages, Limitations, and Open Problems

ADOPT inherits the interpretability and statistical appeal of additive models while allowing responses in arbitrary geodesic metric spaces.

  • Advantages:
    • Each μ(x):=argminmM  E[d2(Y,m)X=x].\mu_\oplus(x) := \arg\min_{m\in\mathcal M}\; \mathbb E\big[d^2(Y, m)\mid X = x\big].2 is interpretable and univariate, allowing visualization and diagnostics.
    • The model is intrinsic to μ(x):=argminmM  E[d2(Y,m)X=x].\mu_\oplus(x) := \arg\min_{m\in\mathcal M}\; \mathbb E\big[d^2(Y, m)\mid X = x\big].3, avoiding possibly-distorting embeddings.
    • Efficiently mitigates dimensionality issues by decoupling predictor effects.
  • Challenges:
    • Computational scalability for large μ(x):=argminmM  E[d2(Y,m)X=x].\mu_\oplus(x) := \arg\min_{m\in\mathcal M}\; \mathbb E\big[d^2(Y, m)\mid X = x\big].4 or complex metric spaces.
    • Bandwidth selection and regularization remain semi-manual.
    • Extension to spaces lacking unique geodesics (e.g., spheres) is non-trivial.
    • High-order interaction modeling and dimension-reduction schemes are underexplored.

Further empirical evaluation, theoretical guarantees under weaker conditions, and generalizations to broader classes of metric-valued regression remain substantial avenues of research (Song et al., 8 Dec 2025).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Additive Optimal Transport Regression.