Papers
Topics
Authors
Recent
Search
2000 character limit reached

High-Dimensional Procrustes Matching via Tree Counts

Published 9 Jul 2026 in stat.ML, cs.IT, cs.LG, and math.ST | (2607.08538v1)

Abstract: Suppose we observe two sets of nn Gaussian vectors in R<sup>d\mathbb{R}<sup>d, with the promise that, after applying a permutation of [n][n] and a rotation of R<sup>d\mathbb{R}<sup>d, the two sets are ρρ-correlated. The Procrustes matching problem asks us to recover the unknown permutation of [n][n] that aligns the two sets. The problem is well-studied in the low-dimensional regime d=O(log⁡n)d=O(\log n), but the high-dimensional regime d≫log⁡nd\gg \log n has remained largely uncharted: prior matching guarantees require nearly perfect correlation ρ=1−o(1)ρ=1-o(1), even for information-theoretic recovery. Our main result is a polynomial-time algorithm for exact recovery at constant correlation. The algorithm works by computing and comparing weighted counts of a specially chosen family of ``wide'' trees. So long as d≥polylog(n)d\ge \mathrm{polylog}(n), the algorithm succeeds with high probability for any $ρ<sup>2&gt;\sqrtα$, where α≈0.338α\approx 0.338 is Otter's tree-counting constant. We complement this algorithmic result with an improved information-theoretic guarantee, showing that exact recovery is possible when ρ<sup>2</sup>≳max⁡log⁡n/d,log⁡n/nρ<sup>2</sup> \gtrsim \max{\log n/d,\sqrt{\log n/n}}. We also carry out a low-degree advantage calculation, which suggests that the condition $ρ<sup>2</sup> &gt; \sqrtα$ is necessary for any tree-counting algorithm.

Summary

  • The paper introduces a polynomial-time algorithm that recovers the hidden permutation exactly using weighted tree-count features in a high-dimensional setting with constant correlation.
  • It employs a novel mapping of data points to polynomial feature vectors using wide trees and efficient color-coding to control variance and ensure mean separation.
  • The work delineates a computational-statistical gap, demonstrating that while information-theoretic recovery is possible at lower correlations, polynomial-time algorithms are limited by tree-based statistical thresholds.

High-Dimensional Procrustes Matching via Tree Counts

Problem Setting and Motivation

The Procrustes matching problem arises in high-dimensional inference tasks where two datasets, XX and YY, each representing nn points in Rd\mathbb{R}^d, are observed. These datasets correspond to correlated, unobserved objects but are obscured by an unknown permutation π\pi (relabeling of indices) and an unknown orthogonal transformation QQ (latent rotation). Mathematically, the generative model is:

Xi∼N(0,Id)Yi=ρQXπ(i)+1−ρ2ZiX_i \sim \mathcal{N}(0, I_d) \qquad Y_i = \rho Q X_{\pi(i)} + \sqrt{1-\rho^2}Z_i

with ρ∈[0,1]\rho \in [0,1] and Zi∼N(0,Id)Z_i \sim \mathcal{N}(0, I_d) independent. The objective is to recover the hidden permutation π\pi given only YY0 and YY1.

While efficient algorithms exist in the low-dimensional regime (YY2), the high-dimensional regime (YY3) presents significant statistical and computational barriers, particularly when YY4 is a constant bounded away from YY5. Existing approaches require nearly-perfect correlation (YY6) for tractable matching when YY7 is large.

Tree Counting Framework and Algorithmic Innovation

The principal contribution is a polynomial-time algorithm for exact permutation recovery in the high-dimensional regime under constant correlation—specifically, when YY8 and YY9 for Otter’s tree-counting constant nn0.

Key Algorithmic Features

  • Polynomial Features via Trees:

Each data point nn1 is mapped to a vector of polynomial features, constructed as weighted counts of rooted, labeled bipartite trees embedded in nn2. Each tree encodes a monomial in the entries of nn3; the template family nn4 comprises trees whose nn5-nodes all have degree exactly nn6 and whose root degree is large compared to tree size—so-called “wide” trees.

  • Permutation Recovery via Similarity Matching:

For each pair nn7, a similarity score is computed as the inner product of these signature vectors over a suite of such trees, aggregated and weighted appropriately. Permutation recovery is then cast as solving for the bijection that maximizes aggregate similarity.

  • Efficient Counting by Color Coding:

Direct enumeration of tree counts is super-polynomial, but the algorithm leverages color-coding techniques to approximate such counts efficiently, yielding a total runtime of nn8 for some constant nn9.

Statistical Mechanisms

The key technical barrier with high-dimensional Procrustes matching is the low-rank structure and the absence of positive edge-wise correlation after the hidden rotation. However, matching pairs of edges that share Rd\mathbb{R}^d0-node endpoints between Rd\mathbb{R}^d1 and Rd\mathbb{R}^d2 restores correlation in the form of rank-2 moments, enabling the design of informative features.

Weighted counts of wide trees serve two purposes:

  1. Variance Control: Tree counts indexed by high root degree and bounded local tree-width mitigate spurious correlations induced by low-rank effects.
  2. Mean Separation: The aggregate statistics concentrate so that the true matching achieves significantly higher similarity than any false matching with high probability, provided Rd\mathbb{R}^d3.

Main Theoretical Results

Polynomial-Time Exact Recovery

  • Algorithmic Guarantee:

There exists a polynomial-time estimator recovering the permutation exactly with high probability whenever Rd\mathbb{R}^d4 and Rd\mathbb{R}^d5.

  • Tree Structure:

The required tree width (root degree Rd\mathbb{R}^d6) is Rd\mathbb{R}^d7; tree size is Rd\mathbb{R}^d8.

  • Mean-Variance Separation:

For true matches, expected similarity is asymptotically Rd\mathbb{R}^d9; for false matches, expectation is zero, and the variance is controlled to π\pi0 for non-diagonal pairs.

Information-Theoretic Bound

An improved QAP-based information-theoretic analysis shows that (possibly intractable) estimators can succeed under

π\pi1

matching known lower bounds for related linear assignment models. As π\pi2 increases, the required correlation for possible recovery decreases, indicating a significant computational-statistical gap.

Computational-Statistical Gap

It is explicitly shown that, although information-theoretic recovery is feasible at vanishing correlation for sufficiently large π\pi3, no known polynomial-time algorithm can match this bound. The tree-counting approach is barriered at π\pi4, and low-degree polynomial testing shows this cannot be improved using only tree-based feature statistics.

Structural and Analytical Techniques

  • Combining Weingarten Calculus with Circuit Decomposition:

The analysis of joint moments under random π\pi5 leverages Weingarten calculus for the orthogonal group, expressing moment calculations as sums over circuit decompositions of associated bipartite graphs. This enables explicit tracking of contributions from tree-based statistics and bounds on their concentration.

  • Wide Tree Construction:

The choice of “wide” trees suppresses the excess variance from overlapping substructures. The authors formalize this advantage with a detailed combinatorial and probabilistic analysis, linking the achievable correlation threshold to Otter’s constant—mirroring thresholds in random graph matching and matrix models.

  • Low-Degree-Polynomial Limitations:

The theoretical analysis demonstrates, via a low-degree framework, that counting trees with degree-π\pi6 π\pi7-nodes indeed achieves optimal separation for moment-based statistics: extending to more general trees does not improve the threshold unless the algorithm leaves the space of tree-count-based methods.

Implications and Future Directions

The demonstrated result represents the first polynomial-time method achieving exact Procrustes permutation recovery in the high-dimensional, constant-correlation regime. Key implications are:

  • Advancement in Geometric Matching:

The method bridges the algorithmic limitations in high-dimensional geometric and graph matching, indicating that fine control of feature-aggregation using combinatorial tree structures extends optimal recovery regimes well beyond classical spectral or low-rank approaches.

  • Barriers and Open Problems:

The robust computational-statistical gap suggests that tree-counting approaches capture the optimal threshold among low-degree algorithms, but further improvement (toward information-theoretic limits) may require fundamentally different algorithmic paradigms—possibly outside the class of polynomial-time algorithms, or using richer (non-tree) structural features.

  • Broader Applications and Generalizations:

This analytical template is generalizable to other matching and alignment problems in high-dimensions, including those with more complex latent transformations, heteroskedastic noise, or manifold structures.

Conclusion

This work rigorously characterizes the phase transition and algorithmic feasibility of high-dimensional Procrustes matching based on tree-count polynomial features. The connection to Otter's constant is established as a sharp computational threshold within tree-counting algorithms, and the authors robustly delineate the algorithmic and statistical frontiers for this family of high-dimensional inference problems. The analysis leverages fine combinatorial, probabilistic, and group-integral methods, and clarifies the limits of low-degree polynomial approaches for latent alignment in high dimensions. These results will inform both practical design and theoretical understanding for geometric matching and related inference problems in modern machine learning and statistics (2607.08538).

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 1 tweet with 4 likes about this paper.