High-Dimensional Procrustes Matching via Tree Counts
Published 9 Jul 2026 in stat.ML, cs.IT, cs.LG, and math.ST | (2607.08538v1)
Abstract: Suppose we observe two sets of n Gaussian vectors in R<sup>d, with the promise that, after applying a permutation of [n] and a rotation of R<sup>d, the two sets are ρ-correlated. The Procrustes matching problem asks us to recover the unknown permutation of [n] that aligns the two sets. The problem is well-studied in the low-dimensional regime d=O(logn), but the high-dimensional regime d≫logn has remained largely uncharted: prior matching guarantees require nearly perfect correlation ρ=1−o(1), even for information-theoretic recovery. Our main result is a polynomial-time algorithm for exact recovery at constant correlation. The algorithm works by computing and comparing weighted counts of a specially chosen family of ``wide'' trees. So long as d≥polylog(n), the algorithm succeeds with high probability for any $ρ<sup>2>\sqrtα$, where α≈0.338 is Otter's tree-counting constant. We complement this algorithmic result with an improved information-theoretic guarantee, showing that exact recovery is possible when ρ<sup>2</sup>≳maxlogn/d,logn/n. We also carry out a low-degree advantage calculation, which suggests that the condition $ρ<sup>2</sup> > \sqrtα$ is necessary for any tree-counting algorithm.
The paper introduces a polynomial-time algorithm that recovers the hidden permutation exactly using weighted tree-count features in a high-dimensional setting with constant correlation.
It employs a novel mapping of data points to polynomial feature vectors using wide trees and efficient color-coding to control variance and ensure mean separation.
The work delineates a computational-statistical gap, demonstrating that while information-theoretic recovery is possible at lower correlations, polynomial-time algorithms are limited by tree-based statistical thresholds.
High-Dimensional Procrustes Matching via Tree Counts
Problem Setting and Motivation
The Procrustes matching problem arises in high-dimensional inference tasks where two datasets, X and Y, each representing n points in Rd, are observed. These datasets correspond to correlated, unobserved objects but are obscured by an unknown permutation π (relabeling of indices) and an unknown orthogonal transformation Q (latent rotation). Mathematically, the generative model is:
Xi∼N(0,Id)Yi=ρQXπ(i)+1−ρ2Zi
with ρ∈[0,1] and Zi∼N(0,Id) independent. The objective is to recover the hidden permutation π given only Y0 and Y1.
While efficient algorithms exist in the low-dimensional regime (Y2), the high-dimensional regime (Y3) presents significant statistical and computational barriers, particularly when Y4 is a constant bounded away from Y5. Existing approaches require nearly-perfect correlation (Y6) for tractable matching when Y7 is large.
Tree Counting Framework and Algorithmic Innovation
The principal contribution is a polynomial-time algorithm for exact permutation recovery in the high-dimensional regime under constant correlation—specifically, when Y8 and Y9 for Otter’s tree-counting constant n0.
Key Algorithmic Features
Polynomial Features via Trees:
Each data point n1 is mapped to a vector of polynomial features, constructed as weighted counts of rooted, labeled bipartite trees embedded in n2. Each tree encodes a monomial in the entries of n3; the template family n4 comprises trees whose n5-nodes all have degree exactly n6 and whose root degree is large compared to tree size—so-called “wide” trees.
Permutation Recovery via Similarity Matching:
For each pair n7, a similarity score is computed as the inner product of these signature vectors over a suite of such trees, aggregated and weighted appropriately. Permutation recovery is then cast as solving for the bijection that maximizes aggregate similarity.
Efficient Counting by Color Coding:
Direct enumeration of tree counts is super-polynomial, but the algorithm leverages color-coding techniques to approximate such counts efficiently, yielding a total runtime of n8 for some constant n9.
Statistical Mechanisms
The key technical barrier with high-dimensional Procrustes matching is the low-rank structure and the absence of positive edge-wise correlation after the hidden rotation. However, matching pairs of edges that share Rd0-node endpoints between Rd1 and Rd2 restores correlation in the form of rank-2 moments, enabling the design of informative features.
Weighted counts of wide trees serve two purposes:
Variance Control: Tree counts indexed by high root degree and bounded local tree-width mitigate spurious correlations induced by low-rank effects.
Mean Separation: The aggregate statistics concentrate so that the true matching achieves significantly higher similarity than any false matching with high probability, provided Rd3.
Main Theoretical Results
Polynomial-Time Exact Recovery
Algorithmic Guarantee:
There exists a polynomial-time estimator recovering the permutation exactly with high probability whenever Rd4 and Rd5.
Tree Structure:
The required tree width (root degree Rd6) is Rd7; tree size is Rd8.
Mean-Variance Separation:
For true matches, expected similarity is asymptotically Rd9; for false matches, expectation is zero, and the variance is controlled to π0 for non-diagonal pairs.
Information-Theoretic Bound
An improved QAP-based information-theoretic analysis shows that (possibly intractable) estimators can succeed under
π1
matching known lower bounds for related linear assignment models. As π2 increases, the required correlation for possible recovery decreases, indicating a significant computational-statistical gap.
Computational-Statistical Gap
It is explicitly shown that, although information-theoretic recovery is feasible at vanishing correlation for sufficiently large π3, no known polynomial-time algorithm can match this bound. The tree-counting approach is barriered at π4, and low-degree polynomial testing shows this cannot be improved using only tree-based feature statistics.
Structural and Analytical Techniques
Combining Weingarten Calculus with Circuit Decomposition:
The analysis of joint moments under random π5 leverages Weingarten calculus for the orthogonal group, expressing moment calculations as sums over circuit decompositions of associated bipartite graphs. This enables explicit tracking of contributions from tree-based statistics and bounds on their concentration.
Wide Tree Construction:
The choice of “wide” trees suppresses the excess variance from overlapping substructures. The authors formalize this advantage with a detailed combinatorial and probabilistic analysis, linking the achievable correlation threshold to Otter’s constant—mirroring thresholds in random graph matching and matrix models.
Low-Degree-Polynomial Limitations:
The theoretical analysis demonstrates, via a low-degree framework, that counting trees with degree-π6 π7-nodes indeed achieves optimal separation for moment-based statistics: extending to more general trees does not improve the threshold unless the algorithm leaves the space of tree-count-based methods.
Implications and Future Directions
The demonstrated result represents the first polynomial-time method achieving exact Procrustes permutation recovery in the high-dimensional, constant-correlation regime. Key implications are:
Advancement in Geometric Matching:
The method bridges the algorithmic limitations in high-dimensional geometric and graph matching, indicating that fine control of feature-aggregation using combinatorial tree structures extends optimal recovery regimes well beyond classical spectral or low-rank approaches.
Barriers and Open Problems:
The robust computational-statistical gap suggests that tree-counting approaches capture the optimal threshold among low-degree algorithms, but further improvement (toward information-theoretic limits) may require fundamentally different algorithmic paradigms—possibly outside the class of polynomial-time algorithms, or using richer (non-tree) structural features.
Broader Applications and Generalizations:
This analytical template is generalizable to other matching and alignment problems in high-dimensions, including those with more complex latent transformations, heteroskedastic noise, or manifold structures.
Conclusion
This work rigorously characterizes the phase transition and algorithmic feasibility of high-dimensional Procrustes matching based on tree-count polynomial features. The connection to Otter's constant is established as a sharp computational threshold within tree-counting algorithms, and the authors robustly delineate the algorithmic and statistical frontiers for this family of high-dimensional inference problems. The analysis leverages fine combinatorial, probabilistic, and group-integral methods, and clarifies the limits of low-degree polynomial approaches for latent alignment in high dimensions. These results will inform both practical design and theoretical understanding for geometric matching and related inference problems in modern machine learning and statistics (2607.08538).