---
title: Procrustes Matching Problem
url: https://www.emergentmind.com/topics/procrustes-matching-problem
type: topic
---

# Procrustes Matching Problem

Searching arXiv for recent and foundational papers on the Procrustes matching problem and major variants.
Searching arXiv: "Procrustes matching problem".
The Procrustes matching problem is a family of alignment problems in which one seeks a transformation of a dataset, or of one point set relative to another, that minimizes a prescribed discrepancy while respecting geometric or structural constraints. In its classical form, congruent Procrustes analysis aims to find the best matching between two point sets through rotation, reflection and translation; in matrix form, it is a matrix approximation problem of the type $\min_{A\in\mathcal S}\|AX-B\|_F^2$. The modern literature extends this template to rigid and similarity alignment, unknown correspondences and permutation recovery, robust power-$1$ objectives, hyperbolic isometries, positive semidefinite and non-symmetric semidefinite constraints, and high-dimensional embedding alignment [2102.03723], [2304.14961], [2207.08592], [2405.14532].

## 1. Canonical formulations

In Euclidean space, the classical point-set formulation seeks an isometry $T$ minimizing
\[
\widehat{T} = \arg\min_{T \in \mathcal{T}} \sum_{n=1}^{N} \|z_n - T(z'_n)\|_2^2,
\]
where $\mathcal T$ is the group of isometries, typically including rotations, translations, and possibly reflections and scalings [2102.03723]. A closely related matrix formulation asks, given matrices $X,B\in\mathbb R^{n\times m}$, for a matrix $A$ in a structured class $\mathcal S$ minimizing $\|AX-B\|_F^2$; changing $\mathcal S$ yields orthogonal, unitary, positive semidefinite, or non-symmetric positive semidefinite variants [2104.06201].

When correspondences are known and the feasible set is orthogonal, the problem has a closed-form solution via singular value decomposition. In the orthogonal Procrustes problem, the optimal transformation is obtained from the SVD of the cross term, with the standard form $Q^\star = UV^\top$ after decomposing the appropriate matrix product [1805.11222]. This SVD-based solvability is one reason the Procrustes problem appears repeatedly as a subroutine in larger procedures, including sparse PCA on the Stiefel manifold and alternating methods for unsupervised alignment [1602.03992].

| Formulation | Representative objective | Decision variables |
|---|---|---|
| Orthogonal Procrustes | $\min_{Q \in \mathcal{O}_D}\|\!QX-Y\|\_F$ | orthogonal matrix |
| Procrustes matching | $d^2(P,Q)=\min_{X,R}\|RP-QX\|_F^2$ | orthogonal matrix, permutation matrix |
| PSDP / NSPSDP | $\min \|AX-B\|_F^2$ with $A \succeq 0$ or $A+A^\top \succeq 0$ | structured matrix |
| Procrustes-Wasserstein | $\min_{P\in S_n,\;Q\in O(d)} \|QX-YP\|_F^2$ | orthogonal matrix, permutation matrix |

This range of formulations shows that “Procrustes” is not limited to labeled rigid point-set alignment. The term also covers constrained matrix approximation, inverse problems with algebraic structure, and unsupervised alignment in which the transformation and the correspondence are both unknown [2304.14961].

## 2. Correspondence uncertainty and combinatorial matching

A central extension is the Procrustes matching problem in which the labeling itself must be recovered. For point clouds $P,Q\in\mathbb R^{d\times n}$, the exact formulation is
\[
d^2(P, Q) = \min_{X, R} \| RP - QX \|_F^2,
\]
with $R\in O(d)$ and $X\in \Pi_n$ a permutation matrix [1606.01548]. In this setting, the optimization is non-convex because the orthogonal and permutation constraints are coupled.

The exact decision version is computationally equivalent to the graph isomorphism problem [1606.01548]. That equivalence places Procrustes matching within the same combinatorial landscape as graph matching and quadratic assignment. For exact PM problems, a convex semidefinite relaxation, PM-SDP, has been analyzed using a moment interpretation. For generic input shapes that are asymmetric or bilaterally symmetric, the relaxation returns a correct solution of PM. For symmetric shapes, however, PM has multiple solutions; the non-convex set of optimal PM solutions is strictly contained in the convex set of optimal PM-SDP solutions, and the true PM solutions are precisely the extreme points of the PM-SDP solution set [1606.01548]. This resolves a common misconception that a convex relaxation must return a discrete matching whenever the objective value is exact.

Unsupervised alignment of embeddings produces a related formulation. Wasserstein Procrustes jointly estimates an orthogonal matrix and a permutation matrix through
\[
\min_{\mathbf{Q} \in \mathcal{O}_d}\min_{\mathbf{P} \in \mathcal{P}_n} \| \mathbf{XQ} - \mathbf{PY} \|_2^2,
\]
which is the squared Wasserstein distance between the aligned clouds under one-to-one matching [1805.11222]. A later planted-model treatment writes the joint maximum-likelihood problem as
\[
(\hat P, \hat Q) \in \argmin_{P\in S_n,\,Q\in O(d)} \|Q X - Y P\|_F^2,
\]
and studies performance through Euclidean transport cost rather than only exact overlap [2405.14532]. In low dimension, that distinction is decisive: small transport cost may be achievable even when exact overlap is not.

Unknown correspondences also arise in contour registration. For two-dimensional statistical shape models, a dynamic time warping strategy was proposed for Procrustes registration without correspondences. The method alternates between order-preserving correspondence estimation and weighted Procrustes registration, handles open or closed contours, and extends generalized Procrustes analysis to ensembles without pre-assigned landmarks [1911.11431]. In that setting, the ordering along the contour is part of the structure and is enforced by the dynamic-programming constraint set.

## 3. Robust, probabilistic, and high-dimensional variants

The least-squares objective is sensitive to outliers, so a robust variant replaces the sum of squared distances with a power-$1$ objective:
\[
\min_{R \in O(d),\; t \in \mathbb{R}^d} \sum_{i=1}^n \|R p(i) + t - q(i)\|.
\]
Unlike the least-squares problem, no closed-form solution is known for this robust objective [2207.08592]. A symmetrized convex relaxation introduces
\[
E_p(A) = \left( \sum_{i=1}^n \left( \|A p(i) - q(i)\|^p + \|A^\top q(i) - p(i)\|^p \right) \right)^{1/p},
\]
then relaxes orthogonality, solves a convex program, and projects back to $O(d)$ [2207.08592]. For $p=2$, the resulting method has a $\sqrt{2}$-factor approximation guarantee, and under the dominance-of-inliers property or its affine analogue it exactly recovers the true transformation from correspondences contaminated by outliers [2207.08592].

In statistical alignment of high-dimensional matrices, Goodall’s perturbation model expresses each observation as
\[
\boldsymbol{X}_i= \alpha_i (\boldsymbol{M} + \boldsymbol{E}_i)\boldsymbol{R}_i^\top + \mathbf{1}_n^\top \boldsymbol{t}_i,
\]
with orthogonal $\boldsymbol R_i$ and matrix-normal perturbations [2008.04631]. The ProMises model places a matrix von Mises–Fisher prior on $\boldsymbol R_i$, yielding a conjugate posterior and a maximum a posteriori estimator obtained by SVD of $\boldsymbol{X}_i^\top \Sigma_n^{-1} \boldsymbol{M} \Sigma_m^{-1} + k\boldsymbol F$ [2008.04631]. This addresses the non-identifiability and interpretability problems of the original perturbation model. For the high-dimensional regime $n\ll m$, the Efficient ProMises model reduces the SVD from $m\times m$ to $n\times n$, lowering the computational complexity from $\mathcal O(m^3)$ to $\mathcal O(mn^2)$ [2008.04631].

The fitted Procrustes transformations can themselves be turned into similarity measures. Two distances have been proposed for between-matrices similarity: a residual-based distance,
\[
d_{Re}(\hat X_i,\hat X_j)=\|\hat X_i-\hat X_j\|_F^2,
\]
and a rotational-based distance,
\[
d_{Ro}(\hat R_i,\hat R_j)=\|\hat R_i-\hat R_j\|_F^2 = 2m - 2\operatorname{tr}(\hat R_i^\top \hat R_j),
\]
which distinguish differences remaining after alignment from differences in the rotations required to reach a common space [2301.06164].

A related recent line studies interoperability between embedding models. If two sets of embeddings have approximately preserved Gram matrices, then there exists an orthogonal alignment with provably small Procrustes error; the alignment is again obtained by SVD and used as a Procrustes post-processing step [2510.13406]. This suggests that orthogonal alignment is not merely a fitting device but also a geometry-preserving interface between separately trained representation spaces.

## 4. Hyperbolic Procrustes analysis

The hyperbolic generalization replaces Euclidean geometry with the Lorentz, or hyperboloid, model. The ambient space is
\[
\mathbb{H}^d = \{ x \in \mathbb{R}^{d+1} : [x, x] = -1,\ x_1 > 0 \},
\]
with Lorentzian inner product
\[
[x, x'] = x^\top H x', \qquad
H = \begin{pmatrix}-1 & 0^\top \\ 0 & I_d \end{pmatrix},
\]
and hyperbolic distance
\[
d(x, x') = \mathrm{acosh}(-[x, x']) .
\]
The hyperbolic Procrustes problem seeks the optimal hyperbolic isometry aligning two point sets in $\mathbb H^d$ [2102.03723].

In the hyperboloid model, any isometry can be written as
\[
T(x) = R_U R_b x,
\]
where $U\in \mathbb O(d)$ and $R_b$ is a translation operator parameterized by $b\in \mathbb R^d$ [2102.03723]. The centering step is therefore more intricate than in Euclidean space. The hyperbolic centroid of $x_1,\ldots,x_N\in\mathbb H^d$ is defined by
\[
m_x := \frac{1}{\sqrt{-[\overline{x_n}, \overline{x_n}]}}\, \overline{(x_n)},
\]
and translation by $-m_x$ centers the point set in the sense that
\[
\overline{ R_{-m_x} x_n } = 0
\]
[2102.03723].

For noise-free matched data related by a hyperbolic isometry, the optimal map is obtained in two stages: center both sets by their hyperbolic centroids, then solve for the optimal rotation by SVD. If $U_l \Sigma U_r^\top$ is the singular value decomposition of
\[
(R_{-m_x} X) W (R_{-m_{x'}} X')^\top,
\]
then the optimal orthogonal factor is
\[
\widehat{U} = U_l U_r^\top,
\]
and the full solution is
\[
T = T_{m_x} \circ T_U \circ T_{-m_{x'}} .
\]
This reproduces the Euclidean “center, then rotate” pattern, but with Lorentzian centroid and hyperbolic isometries in place of Euclidean translations and rotations [2102.03723].

Under measurement noise, the closed-form solution is approximate. The paper compares a closed-form estimator $R_P$, direct gradient descent $R_{\mathrm{GD}}$, and a hybrid $R_{\mathrm{GD+P}}$ initialized from the closed-form map. The discrepancy metric is
\[
e(X, \overline{X}) := \frac{1}{Nd} \sum_{n=1}^N d(x_n, \overline{x}_n),
\]
and the empirical analysis reports that the proposed method is fast and robust, with significantly fewer outlier cases than vanilla gradient descent, while the hybrid yields the best results [2102.03723].

## 5. Constrained matrix variants

A major branch of the subject treats Procrustes problems as constrained matrix approximation. The positive semidefinite Procrustes problem (PSDP) is
\[
\inf_{A \in \mathcal{S}^n_{\succeq}} \|AX - B\|_F^2,
\]
where $\mathcal S^n_{\succeq}$ is the cone of real symmetric positive semidefinite matrices [1612.01354]. A key semi-analytical result reduces the problem, after an SVD of $X$, to a smaller PSDP in dimension $r=\operatorname{rank}(X)$:
\[
\inf_{A \in \mathcal{S}^n_{\succeq}} \|AX - B\|_F^2
=
\min_{A_{11} \in \mathcal{S}^r_{\succeq}} \|A_{11}\Sigma_1 - U_1^T B V_1\|_F^2 + \|B V_2\|_F^2 .
\]
This reduction yields a complete characterization of the solution set, identifies when the infimum is not attained, and supports a fast gradient method with linear convergence on the reduced problem [1612.01354].

The non-symmetric positive semidefinite Procrustes problem (NSPSDP) replaces symmetry by the condition $A+A^\top \succeq 0$:
\[
\inf_{A \in \mathcal{N}_{\succeq}^{S_n}} \|AX - B\|_F^2,
\qquad
\mathcal{N}_{\succeq}^{S_n} = \{ A \in \mathbb{R}^{n \times n} : A + A^T \succeq 0\}.
\]
As in the symmetric case, the problem can be reduced to a smaller diagonal full-rank instance, then solved by a fast gradient method whose main step is projection onto the NSPSD cone. The same framework extends to the complex case through an overparametrized real embedding [2104.06201].

A more general conic-optimization view writes Procrustes problems as
\[
\min_{X \in \mathcal{P}} \| \mathcal{L}(X) \|,
\]
with $\mathcal L$ linear, $\mathcal P$ a feasible set defined by linear, quadratic, or semidefinite constraints, and the norm chosen from Frobenius, $l_1$, $l_2$, or $l_\infty$ [2304.14961]. This covers balanced and unbalanced weighted Procrustes problems, partially specified targets, oblique constraints $\operatorname{diag}(X^\top X)=\mathbf 1$, projection problems such as $CX-XA$, and two-sided formulations. The reformulation produces a rank-constrained semidefinite program, so convex relaxations, trace or log-det heuristics, convex iteration, and bisection become available tools [2304.14961].

The Procrustes subproblem also appears inside other non-convex optimizations. In orthogonal sparse PCA, an MM surrogate reduces each iteration to the rectangular Procrustes problem
\[
\max_{X^\top X = I_q} \operatorname{Tr}(Y^\top X),
\]
whose solution is $X^\star = V_L V_R^\top$ from the SVD of $Y$ [1602.03992]. By contrast, in regularized multivariate analysis, replacing the second step of an alternating scheme by orthogonal Procrustes is not optimal from the perspective of the overall MVA method. An eigenvalue-based update preserves the uncorrelation of the extracted features and recovers the original unregularized solution as the regularization vanishes [1605.02674]. This is a methodological caveat: a Procrustes step may be algebraically convenient without being statistically or geometrically appropriate for the larger problem.

Structured inverse eigenvalue problems furnish another specialized variant. For normal (skew) $J$-Hamiltonian and normal $J$-symplectic matrices, the Procrustes problem asks for the matrix in the structured solution set of $AX=XD$ closest in Frobenius norm to a prescribed $A_0$. The paper solves these problems by block-diagonalizing $J$, deriving explicit blockwise minimizers, and using the Cayley transform in the $J$-symplectic case [2401.13121].

## 6. Statistical limits, Bayesian formulations, and large-scale applications

Recent work has made the statistical structure of Procrustes matching explicit. In a high-dimensional planted model,
\[
Y_i = \rho Q X_{\pi(i)} + \sqrt{1-\rho^2} Z_i,
\]
with unknown permutation $\pi$, unknown orthogonal matrix $Q$, and Gaussian noise, the task is exact recovery of $\pi$ from two sets of Gaussian vectors [2607.08538]. A polynomial-time algorithm based on weighted counts of “wide” trees succeeds with high probability for $d\ge \operatorname{polylog}(n)$ whenever
\[
\rho^2 > \sqrt{\alpha},
\qquad \alpha \approx 0.338,
\]
where $\alpha$ is Otter’s tree-counting constant [2607.08538]. The same paper gives an information-theoretic guarantee for exact recovery when
\[
\rho^2 \gtrsim \max\left\{\frac{\log n}{d},\; \sqrt{\frac{\log n}{n}}\right\},
\]
and a low-degree analysis suggesting that the condition $\rho^2 > \sqrt{\alpha}$ is necessary for tree-counting algorithms [2607.08538]. This establishes a statistical-computational gap in the high-dimensional regime.

The Procrustes-Wasserstein literature studies a related planted model through the Euclidean transport cost. In the high-dimensional regime $d\gg \log n$, recovery of both the orthogonal map and the relabeling is information-theoretically possible as soon as $\sigma \to 0$; in the low-dimensional regime $d\ll \log n$, small transport cost can still be achieved when exact overlap is much harder [2405.14532]. The Ping-Pong algorithm alternates a closed-form orthogonal Procrustes update with a linear assignment step, initialized by a Frank-Wolfe convex relaxation, and is analyzed for one-step recovery under sufficiently small noise [2405.14532].

Probabilistic formulations also appear in applied reconstruction. In unposed 3D Gaussian Splatting, alignment of large submaps is formulated as a probabilistic Procrustes problem over similarity transformations:
\[
\min_{s, R, \mathbf{t}, \gamma}
\sum_{l=1}^N \gamma_l \| s R \mathbf{p}_l + \mathbf{t} - \mathbf{q}_l \|^2
+
\epsilon \sum_{l=1}^N \gamma_l \ln \gamma_l,
\]
with soft correspondence weights $\gamma_l$ and a soft dustbin mechanism for uncertain matches [2507.18541]. The method alternates a Kabsch-Umeyama closed-form initialization with entropy-regularized weight updates and gradient-based refinement of pose parameters. This suggests a convergence of Procrustes ideas with probabilistic coupling and large-scale scene reconstruction.

A Bayesian perspective predates many of these developments. For unlabelled point sets, Bayesian matching has been developed under two likelihoods: a Procrustes size-and-shape model, which optimizes out rotation and translation before modeling residuals, and a full configuration model, which treats rotation and translation as parameters and samples them by MCMC [1009.3072]. An improved Procrustes algorithm introduces occasional large jumps during burn-in—nearness, rotation, translation, and flip moves—to improve convergence. The two models are linked by a Laplace approximation, under which the Procrustes posterior can be viewed as an approximation to the configuration posterior after marginalizing nuisance alignment parameters [1009.3072].

Across these lines of work, the main technical divide is between settings where closed-form SVD solves the entire problem and settings where SVD solves only a subproblem. Robust objectives, unknown correspondences, symmetry-induced multiplicity, and structured feasibility sets typically require convex relaxation, alternating minimization, first-order methods, semidefinite programming, or Bayesian computation [2207.08592], [1606.01548], [1612.01354], [1009.3072]. A plausible implication is that the Procrustes matching problem is best understood not as a single algorithm, but as a unifying optimization paradigm for alignment under geometry, correspondence, and structure.

Source: https://www.emergentmind.com/topics/procrustes-matching-problem