Papers
Topics
Authors
Recent
Search
2000 character limit reached

Optimal Transport Inversion Pipeline

Updated 7 July 2026
  • OTIP is a family of inversion architectures that replaces traditional residual matching with transport structure for improved robustness and convergence.
  • Seismic OTIP employs the quadratic Wasserstein metric in PDE-constrained full-waveform inversion, offering enhanced convexity and noise resilience compared to L2-based methods.
  • In image editing and cost inference, OTIP frameworks utilize zero-shot and Bayesian approaches to guide latent adjustments and recover transport cost structures.

Searching arXiv for recent and foundational OTIP-related papers across seismic inversion, rectified-flow inversion, and inverse OT. {"query":"all:\"Optimal Transport Inversion Pipeline\" OR id:(Lupascu et al., 4 Aug 2025) OR id:(Engquist et al., 2016) OR id:(Yang et al., 2016) OR id:(Engquist et al., 2018)","max_results":10} Searching arXiv for optimal transport inversion papers and adjacent inverse-OT theory. {"query":"cat:math.OC OR cat:cs.CV OR cat:physics.geo-ph AND (\"optimal transport inversion\" OR \"full waveform inversion\" Wasserstein OR \"inverse optimal transport\")","max_results":10} Optimal Transport Inversion Pipeline (OTIP) denotes an inversion workflow in which optimal-transport geometry is used to compare predicted and observed objects, regularize reverse trajectories, or infer latent transport costs. In the literature, the phrase appears explicitly as a zero-shot framework for rectified-flow image inversion and editing (Lupascu et al., 4 Aug 2025), while closely related seismic papers by Engquist, Froese, Yang, and collaborators effectively instantiate OTIP as a PDE-constrained full-waveform inversion pipeline driven by the quadratic Wasserstein metric rather than pointwise L2L_2 discrepancy (Engquist et al., 2016). This suggests a broader usage: OTIP is best understood as a family of inversion architectures in which transport structure replaces or augments classical residual matching.

1. Terminological scope and development

In seismic inversion, the foundational move was to replace the conventional least-squares full-waveform inversion objective with the quadratic Wasserstein distance W22W_2^2. The 2016 paper by Engquist, Froese, and Yang formulates full-waveform inversion as

v~=argminvd(f(v),g),\tilde v = \arg\min_v d(f(v),g),

with

d(f,g)=W22(f,g),d(f,g)=W_2^2(f,g),

and proves advantageous properties related to convexity and robustness to noise (Engquist et al., 2016). A later seismic paper embeds this misfit inside the adjoint-state method, contrasts trace-by-trace and global Wasserstein objectives, and applies the resulting workflow to the Camembert, Marmousi, and 2004 BP models (Yang et al., 2016). Subsequent work makes normalization a first-class design issue, arguing that the success or failure of OT-based inversion depends critically on how signed oscillatory seismic traces are mapped into nonnegative equal-mass densities (Engquist et al., 2018). The 2020 paper “Optimal Transport Based Seismic Inversion: Beyond Cycle Skipping” sharpens the convexity theory and argues that OT can recover geophysical properties from domains where no seismic waves travel through, particularly in salt-inclusion imaging (Engquist et al., 2020).

A distinct usage emerges in generative modeling. The 2025 paper “Transport-Guided Rectified Flow Inversion: Improved Image Editing Using Optimal Transport Theory” introduces OTIP explicitly as a training-free, zero-shot framework for rectified-flow inversion and editing, where OT guidance is injected into the reverse denoising trajectory (Lupascu et al., 4 Aug 2025). Related rectified-flow theory frames OT as a marginal-preserving flow problem over continuous distributions and provides a convex-cost flow construction whose fixed points coincide with cc-optimal couplings (Liu, 2022).

Inverse-OT papers extend the scope further. They treat inversion not as recovery of a physical field from waveforms, but as recovery of an unknown transport cost from observed plans, marginals, values, or potentials. In this setting, OTIP becomes a pipeline for cost inference, posterior sampling, or identifiability analysis rather than a waveform-misfit engine (Stuart et al., 2019, Chiu et al., 2021, González-Sanz et al., 2024).

2. Seismic OTIP as a PDE-constrained inversion workflow

In seismic full-waveform inversion, the unknown is the subsurface wave velocity field or slowness parameter, and the forward model is the acoustic wave equation. A standard OTIP in this setting consists of model initialization, forward wave simulation, trace extraction, transformation of synthetic and observed data into admissible transport densities, OT-misfit evaluation, derivative computation with respect to predicted data, adjoint-state backpropagation, and gradient-based model update (Yang et al., 2016).

Two misfit organizations dominate. The first is trace-by-trace 1D Wasserstein inversion,

J1(m)=r=1RW22(f(xr,t;m),g(xr,t)),J_1(m)=\sum_{r=1}^R W_2^2\big(f(x_r,t;m),g(x_r,t)\big),

where each receiver trace is treated independently in time. The second is global multidimensional Wasserstein inversion,

J2(m)=W22(f(xr,t;m),g(xr,t)),J_2(m)=W_2^2\big(f(x_r,t;m),g(x_r,t)\big),

where the full receiver-time gather is treated as one density (Yang et al., 2016). The trace-by-trace formulation uses the explicit 1D formula

W22(f,g)=0T0tG1(F(t))2f(t)dt,W_2^2(f,g)=\int_0^{T_0} |t-G^{-1}(F(t))|^2 f(t)\,dt,

while the global formulation computes W2W_2 through a Monge–Ampère equation with transport map T(x)=u(x)T(x)=\nabla u(x) (Engquist et al., 2016).

The principal theoretical rationale is that W22W_2^20 is convex with respect to physically relevant deformations. For shifted densities W22W_2^21, the seismic OT papers show that W22W_2^22 is convex in the shift parameter. They also establish convexity with respect to dilation parameters and partial amplitude changes, together with a 1D noise result

W22W_2^23

contrasting with

W22W_2^24

for W22W_2^25-type discrepancy (Engquist et al., 2016). In the later “Beyond Cycle Skipping” analysis, these convexity results are sharpened to joint strict convexity under simultaneous translation and dilation (Engquist et al., 2020).

The practical consequence is that OTIP changes the adjoint source. Under W22W_2^26-FWI the adjoint source is essentially the residual W22W_2^27; under OTIP it becomes W22W_2^28, derived from the Wasserstein geometry. The remainder of the PDE-constrained machinery remains standard: W22W_2^29 This architectural compatibility is one of the reasons OTIP could be inserted into existing adjoint-state FWI codes (Yang et al., 2016, Engquist et al., 2018).

3. Representation and normalization as the central engineering layer

A recurring OTIP difficulty is that raw seismic traces are signed oscillatory signals, whereas standard OT requires nonnegative densities of equal total mass. The earliest seismic formulation addresses this by splitting traces into positive and negative parts,

v~=argminvd(f(v),g),\tilde v = \arg\min_v d(f(v),g),0

and using

v~=argminvd(f(v),g),\tilde v = \arg\min_v d(f(v),g),1

This preserves the basic OT theory but is awkward for large-scale adjoint differentiation (Engquist et al., 2016, Engquist et al., 2018).

Later seismic work develops a broader normalization class

v~=argminvd(f(v),g),\tilde v = \arg\min_v d(f(v),g),2

with v~=argminvd(f(v),g),\tilde v = \arg\min_v d(f(v),g),3 positive, one-to-one, and v~=argminvd(f(v),g),\tilde v = \arg\min_v d(f(v),g),4. Three explicit choices are linear, exponential, and softplus scaling: v~=argminvd(f(v),g),\tilde v = \arg\min_v d(f(v),g),5 The same paper argues that a good normalization must satisfy positivity and mass balance, but also differentiability, signal-structure preservation, and stable behavior under missing arrivals or partial waveform mismatch (Engquist et al., 2018). In “Beyond Cycle Skipping,” linear normalization with positive constant v~=argminvd(f(v),g),\tilde v = \arg\min_v d(f(v),g),6 is shown to induce a Huber-type behavior—quadratic for small shifts and effectively linear for large ones—and to smooth the optimal map and the Fréchet gradient (Engquist et al., 2020).

The same theme appears in a different form in “Geophysical Inversion and Optimal Transport.” Instead of shifting raw traces upward, that paper maps each waveform into a 2D fingerprint density in the time-amplitude plane. Time is normalized to v~=argminvd(f(v),g),\tilde v = \arg\min_v d(f(v),g),7, amplitude is mapped to v~=argminvd(f(v),g),\tilde v = \arg\min_v d(f(v),g),8 with an arctangent transform, the nearest-distance field to the piecewise-linear waveform is computed, and a positive ridge density is built as

v~=argminvd(f(v),g),\tilde v = \arg\min_v d(f(v),g),9

The misfit is then a weighted sum of 1D Wasserstein distances of the time and amplitude marginals,

d(f,g)=W22(f,g),d(f,g)=W_2^2(f,g),0

This construction avoids positive/negative splitting and remains differentiable through the full chain rule (Sambridge et al., 2022).

The rectified-flow OTIP uses a different representation problem. There the source image d(f,g)=W22(f,g),d(f,g)=W_2^2(f,g),1 is encoded as d(f,g)=W22(f,g),d(f,g)=W_2^2(f,g),2, standard RF inversion produces a starting latent d(f,g)=W22(f,g),d(f,g)=W_2^2(f,g),3, and OT enters as a latent displacement direction

d(f,g)=W22(f,g),d(f,g)=W_2^2(f,g),4

The target image is set in practice to the source image itself, d(f,g)=W22(f,g),d(f,g)=W_2^2(f,g),5, so the OT guidance acts as a structure-preserving pull toward the reference latent while prompt-conditioned rectified-flow dynamics continue to drive semantic editing (Lupascu et al., 4 Aug 2025). That paper is explicit that this is OT-inspired rather than exact OT: no primal coupling, no Kantorovich dual potential, and no online Monge solve are performed.

4. Rectified-flow OTIP and trajectory-guided inversion

In rectified-flow image editing, OTIP is an inference-time controller layered on top of a pretrained RF backbone. The base model is FLUX, the number of denoising steps is 28, and the update is explicit Euler (Lupascu et al., 4 Aug 2025). The core continuous-time equation is

d(f,g)=W22(f,g),d(f,g)=W_2^2(f,g),6

with cosine schedule

d(f,g)=W22(f,g),d(f,g)=W_2^2(f,g),7

In discrete form the algorithm computes prompt-conditioned and reference-conditioned RF velocities, blends them using controller guidance d(f,g)=W22(f,g),d(f,g)=W_2^2(f,g),8, clips the OT direction at threshold d(f,g)=W22(f,g),d(f,g)=W_2^2(f,g),9, and updates

cc0

The reported recommended settings include cc1, cc2, cc3, and guidance scale cc4, although ablations in the paper also report best performance around cc5 for age editing and cc6 in a specific face-editing setting (Lupascu et al., 4 Aug 2025).

The RF OTIP literature frames the main problem as the fidelity–editability trade-off. Very faithful inversion can become semantically rigid, while looser inversion drifts away from the source manifold. OTIP addresses this by treating OT guidance as an additive correction to the reverse trajectory rather than as a replacement for prompt-conditioned generation. The paper reports reconstruction LPIPS cc7, SSIM cc8, face recognition distance cc9, and CLIP-I J1(m)=r=1RW22(f(xr,t;m),g(xr,t)),J_1(m)=\sum_{r=1}^R W_2^2\big(f(x_r,t;m),g(x_r,t)\big),0 on SFHQ, compared with RF-Inversion values J1(m)=r=1RW22(f(xr,t;m),g(xr,t)),J_1(m)=\sum_{r=1}^R W_2^2\big(f(x_r,t;m),g(x_r,t)\big),1, J1(m)=r=1RW22(f(xr,t;m),g(xr,t)),J_1(m)=\sum_{r=1}^R W_2^2\big(f(x_r,t;m),g(x_r,t)\big),2, J1(m)=r=1RW22(f(xr,t;m),g(xr,t)),J_1(m)=\sum_{r=1}^R W_2^2\big(f(x_r,t;m),g(x_r,t)\big),3, and J1(m)=r=1RW22(f(xr,t;m),g(xr,t)),J_1(m)=\sum_{r=1}^R W_2^2\big(f(x_r,t;m),g(x_r,t)\big),4, respectively (Lupascu et al., 4 Aug 2025). On LSUN-Bedroom and LSUN-Church stroke-to-image tasks, the same paper reports reconstruction-loss improvements of J1(m)=r=1RW22(f(xr,t;m),g(xr,t)),J_1(m)=\sum_{r=1}^R W_2^2\big(f(x_r,t;m),g(x_r,t)\big),5 and J1(m)=r=1RW22(f(xr,t;m),g(xr,t)),J_1(m)=\sum_{r=1}^R W_2^2\big(f(x_r,t;m),g(x_r,t)\big),6 over RF-Inversion, and on semantic face editing it reports an 11.2% improvement in identity preservation and a 1.6% enhancement in perceptual quality, with runtime essentially unchanged (Lupascu et al., 4 Aug 2025).

The broader rectified-flow OT theory is more exacting. “Rectified Flow: A Marginal Preserving Approach to Optimal Transport” studies convex-cost OT between continuous distributions and iteratively learns neural ODEs that preserve source and target marginals while monotonically decreasing the chosen convex cost (Liu, 2022). Its J1(m)=r=1RW22(f(xr,t;m),g(xr,t)),J_1(m)=\sum_{r=1}^R W_2^2\big(f(x_r,t;m),g(x_r,t)\big),7-rectified flow solves a sequence of unconstrained regressions and yields fixed points that are exactly J1(m)=r=1RW22(f(xr,t;m),g(xr,t)),J_1(m)=\sum_{r=1}^R W_2^2\big(f(x_r,t;m),g(x_r,t)\big),8-optimal couplings. This provides a theoretical background for transport-guided inversion, although the image-editing OTIP uses a cheaper closed-form displacement rule rather than the full J1(m)=r=1RW22(f(xr,t;m),g(xr,t)),J_1(m)=\sum_{r=1}^R W_2^2\big(f(x_r,t;m),g(x_r,t)\big),9-rectified construction.

5. Inverse optimal transport, identifiability, and cost inference

A different OTIP interpretation arises when the unknown is not a physical field or a latent code, but the transport cost itself. “Inverse optimal transport” formulates a Bayesian discrete inverse problem in which a noisy transport plan is observed and the latent variables are unknown marginals and cost parameters. The forward model is

J2(m)=W22(f(xr,t;m),g(xr,t)),J_2(m)=W_2^2\big(f(x_r,t;m),g(x_r,t)\big),0

and the observation model is

J2(m)=W22(f(xr,t;m),g(xr,t)),J_2(m)=W_2^2\big(f(x_r,t;m),g(x_r,t)\big),1

Using normalization maps J2(m)=W22(f(xr,t;m),g(xr,t)),J_2(m)=W_2^2\big(f(x_r,t;m),g(x_r,t)\big),2, J2(m)=W22(f(xr,t;m),g(xr,t)),J_2(m)=W_2^2\big(f(x_r,t;m),g(x_r,t)\big),3, J2(m)=W22(f(xr,t;m),g(xr,t)),J_2(m)=W_2^2\big(f(x_r,t;m),g(x_r,t)\big),4, the posterior becomes

J2(m)=W22(f(xr,t;m),g(xr,t)),J_2(m)=W_2^2\big(f(x_r,t;m),g(x_r,t)\big),5

Inference is performed by Random-Walk Metropolis-within-Gibbs, requiring only a forward OT solver and random-number generation (Stuart et al., 2019).

Discrete inverse-OT theory shows that identifiability is subtle. “Discrete Probabilistic Inverse Optimal Transport” emphasizes that in entropy-regularized discrete OT, a coupling does not determine a unique cost matrix; instead costs are identified only up to cross-ratio-equivalent perturbations, equivalently row/column additive shifts J2(m)=W22(f(xr,t;m),g(xr,t)),J_2(m)=W_2^2\big(f(x_r,t;m),g(x_r,t)\big),6 (Chiu et al., 2021). “Identifiability of the Optimal Transport Cost on Finite Spaces” turns this into necessary and sufficient finite-space criteria. With plans alone, exact cost recovery is impossible in general because

J2(m)=W22(f(xr,t;m),g(xr,t)),J_2(m)=W_2^2\big(f(x_r,t;m),g(x_r,t)\big),7

and only the equivalence class in the quotient space can be expected. With plans plus potentials, exact recovery holds if and only if the union of observed supports covers the entire cost matrix. With plans plus total costs, exact recovery reduces to LP uniqueness or, in a simpler sufficient condition, linear independence of enough face vertices (González-Sanz et al., 2024).

For OTIP design, these results imply that inverse-OT pipelines should not assume that an observed plan uniquely determines a cost. They should instead expose gauge freedoms, use structural priors such as Toeplitz, graph-based, Monge, or symmetric zero-diagonal classes, and return posterior uncertainty or quotient-space representations when the inverse problem is underdetermined (Stuart et al., 2019, González-Sanz et al., 2024). This is a direct theoretical constraint on any OTIP intended for cost inference.

6. Numerical engines, empirical behavior, and limitations

OTIP performance depends strongly on the forward OT solver. In seismic FWI, trace-by-trace 1D J2(m)=W22(f(xr,t;m),g(xr,t)),J_2(m)=W_2^2\big(f(x_r,t;m),g(x_r,t)\big),8 is attractive because it uses the exact monotone-rearrangement formula and has runtime close to J2(m)=W22(f(xr,t;m),g(xr,t)),J_2(m)=W_2^2\big(f(x_r,t;m),g(x_r,t)\big),9-FWI; one paper reports less than 1.1 times the W22(f,g)=0T0tG1(F(t))2f(t)dt,W_2^2(f,g)=\int_0^{T_0} |t-G^{-1}(F(t))|^2 f(t)\,dt,0 runtime for trace-by-trace OTIP, versus about 3 to 4 times for global Monge–Ampère OTIP (Yang et al., 2016). The global solver uses an almost-monotone finite-difference discretization of the Monge–Ampère equation, a filtered scheme,

W22(f,g)=0T0tG1(F(t))2f(t)dt,W_2^2(f,g)=\int_0^{T_0} |t-G^{-1}(F(t))|^2 f(t)\,dt,1

and Newton’s method for the resulting nonlinear system (Engquist et al., 2016, Yang et al., 2016). These papers also note that once the Monge–Ampère Jacobian is available, the Fréchet derivative with respect to predicted data can be computed at relatively low extra cost (Engquist et al., 2016).

In discrete OT, forward-solver engineering has become a separate concern. “A Study of Performance of Optimal Transport” reports that combinatorial methods such as network simplex and augmenting-path algorithms can consistently outperform Sinkhorn and Greenkhorn, even in low-accuracy regimes, with up to orders-of-magnitude speedups (Dong et al., 2020). For assignment-style OT it introduces batched KM as a practical combinatorial variant. By contrast, “A Truncated Newton Method for Optimal Transport” targets weakly regularized entropic OT, solving the dual with a GPU-parallel truncated Newton method inside an annealed continuation framework. It reports 1–5 seconds to 6-decimal precision on W22(f,g)=0T0tG1(F(t))2f(t)dt,W_2^2(f,g)=\int_0^{T_0} |t-G^{-1}(F(t))|^2 f(t)\,dt,2 problems, stability to roughly 9-decimal precision, and a large experiment with W22(f,g)=0T0tG1(F(t))2f(t)dt,W_2^2(f,g)=\int_0^{T_0} |t-G^{-1}(F(t))|^2 f(t)\,dt,3 solved approximately under weak entropic regularization (Kemertas et al., 2 Apr 2025). This suggests that OTIP implementations with repeated nearby OT solves can benefit from warm-started high-precision entropic solvers when exact LP OT is too expensive.

Empirically, the seismic OTIP literature reports consistent improvements over W22(f,g)=0T0tG1(F(t))2f(t)dt,W_2^2(f,g)=\int_0^{T_0} |t-G^{-1}(F(t))|^2 f(t)\,dt,4-based inversion. In the Camembert model, both trace-by-trace and global W22(f,g)=0T0tG1(F(t))2f(t)dt,W_2^2(f,g)=\int_0^{T_0} |t-G^{-1}(F(t))|^2 f(t)\,dt,5 converge in 10 l-BFGS iterations, while W22(f,g)=0T0tG1(F(t))2f(t)dt,W_2^2(f,g)=\int_0^{T_0} |t-G^{-1}(F(t))|^2 f(t)\,dt,6 converges to a local minimum after 100 iterations; in the same study, global W22(f,g)=0T0tG1(F(t))2f(t)dt,W_2^2(f,g)=\int_0^{T_0} |t-G^{-1}(F(t))|^2 f(t)\,dt,7 is more accurate than trace-by-trace W22(f,g)=0T0tG1(F(t))2f(t)dt,W_2^2(f,g)=\int_0^{T_0} |t-G^{-1}(F(t))|^2 f(t)\,dt,8, with about 25% lower W22(f,g)=0T0tG1(F(t))2f(t)dt,W_2^2(f,g)=\int_0^{T_0} |t-G^{-1}(F(t))|^2 f(t)\,dt,9 model error (Yang et al., 2016). In full Marmousi, trace-by-trace W2W_20 reduces relative misfit to 0.1 in 20 iterations, whereas W2W_21 converges slowly to a local minimum; in a hybrid test, continuing with W2W_22 yields relative error W2W_23, while switching from a W2W_24 warm start to W2W_25 yields W2W_26 (Yang et al., 2016). In the salt-model experiments of “Beyond Cycle Skipping,” OTIP is reported to recover much of the salt body while W2W_27 remains poorer and can produce geologically misleading artifacts (Engquist et al., 2020). In source inversion based on waveform fingerprints, Wasserstein-driven optimization reaches 77% convergence versus 41% for W2W_28 in source-location recovery, and 54% versus 35% in coupled source-location plus moment-tensor inversion (Sambridge et al., 2022).

Several limitations recur across domains. Seismic papers repeatedly state that normalization remains the central unresolved issue: theoretically favorable transforms do not always perform best in practice, and practical transforms can weaken the original convexity guarantees (Engquist et al., 2018, Engquist et al., 2020). Global Monge–Ampère OT is expensive, sensitive to data smoothness, and numerically more delicate than 1D tracewise OT (Yang et al., 2016). The earliest OT-FWI paper derives the derivative of the OT misfit with respect to predicted data, but not a full large-scale PDE-constrained adjoint implementation for distributed velocity fields (Engquist et al., 2016). The waveform-fingerprint approach is differentiable and efficient, but it does not define a true likelihood because it does not incorporate a statistical noise model (Sambridge et al., 2022). The rectified-flow OTIP is explicit that its OT step is approximate rather than exact, and that extreme semantic transformations may conflict with the method’s structure-preserving bias (Lupascu et al., 4 Aug 2025). In inverse-OT settings, non-identifiability is intrinsic unless additional information or priors are supplied (Chiu et al., 2021, González-Sanz et al., 2024).

Taken together, these works define OTIP not as one fixed algorithm but as a transport-centered inversion doctrine. In seismic imaging, OTIP is a Wasserstein-misfit and adjoint-state pipeline designed to enlarge the basin of attraction, smooth gradients, and recover low-wavenumber structure. In rectified-flow image editing, it is a transport-guided reverse solver that stabilizes inversion trajectories without retraining the generative model. In inverse OT, it is a cost-recovery pipeline whose success is governed by gauge freedoms, support coverage, and polyhedral identifiability. Across all three interpretations, the decisive ingredients are the same: how transport structure is represented, how positivity and mass constraints are enforced, how the transport computation is solved numerically, and what information is actually identifiable from the observed data.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Optimal Transport Inversion Pipeline (OTIP).