- The paper demonstrates that in the regime d > n, detection is impossible when d >> (n h(p))^3, thereby resolving the longstanding detection threshold conjecture.
- It employs a rigorous KL divergence chain-rule, precise polynomial expansions of interaction functions, and posterior log-concavity to derive sharp statistical bounds.
- The findings imply that latent geometric information is lost in high dimensions, guiding practical network inference in latent space models.
Detection Thresholds for Random Geometric Graphs: Resolution in the High-Dimensional Regime
Introduction and Problem Overview
This paper addresses the intrinsic question of when random geometric graphs (RGGs), constructed via a hard-thresholding scheme on latent points sampled uniformly from the unit sphere in Rd, can be statistically distinguished from Erdős–Rényi (ER) random graphs with the same edge density. Specifically, the focus is on the "detection threshold": the critical point in the parameter space (dimension d, number of nodes n, and edge probability p) where geometry becomes information-theoretically undetectable from pure random connectivity.
Previous works have established sharp detection thresholds in the dense regime (constant p) and, up to polylogarithmic factors, for extremely sparse graphs (p=Θ(1/n)). However, a conjecture had remained open for the intermediate sparse regime: distinguishing between RGG and ER is conjectured to be possible if and only if d≪(nh(p))3, where h(p)=plog(1/p)+(1−p)log(1/(1−p)) is the binary entropy function. This paper resolves the conjecture for d>n and p≳n−2/3/logn, and improves bounds for the regime d0.
Model Setup and Key Results
An d1-node RGG is formed as follows:
- Latent vectors d2 are independently sampled from the unit sphere d3.
- For a threshold d4 defined such that d5, an edge exists between node d6 and d7 if and only if d8.
The detection problem is formulated as a binary hypothesis test: given a sample graph, decide between
- d9: the sample is RGG (n0)
- n1: the sample is ER graph n2 (n3)
with indices n4 and n5 for the corresponding distributions.
The central result is:
Detection between RGG and ER graphs is impossible in total variation when n6 and n7 for any fixed n8. This resolves the detection threshold conjecture in the regime n9, establishing information-theoretic impossibility at the conjectured scaling; the bound is tight up to constant factors.
Numerical and Technical Claims
- For the regime p0 and p1, the Kullback-Leibler divergence satisfies p2.
- Pinsker’s inequality then yields p3, showing indistinguishability in total variation.
- In the very sparse regime p4, the bound p5 still improves on previous contours, but a complete sharp threshold for p6 is left open.
Methods and Technical Innovations
Chain Rule Expansion and Posterior Analysis
The proof leverages a chain-rule expansion of the KL divergence, incrementally revealing vertices and their incident edges, framing the problem in terms of posterior moments of local functions of the latent configuration conditioned on the observed graph.
A key technical advance is the direct analysis of the posterior distribution p7 of the latent positions p8 given the RGG realization p9. This sidesteps the overly revealing approach—used in prior works—of conditioning on all but a fixed subset of latent positions, which loosens information-theoretic bounds.
Three-Stage Proof Sketch
- KL Expansion: Decompose the KL divergence into a sum over vertices, expressing each term as an expectation of squared local "interaction functions" p0 (functions of up to p1 latent points), where p2 ranges over small subgraphs.
- Approximation of Interaction Functions: Obtain precise polynomial expansions of p3 in terms of Gram matrix entries p4 for p5, bounding the remainder with strong (log-concavity-based) concentration results.
- Posterior Moment Control: Use transportation inequalities and strong log-concavity to bound the mean and high moments of the posterior overlap p6 for independent samples p7 from p8. This step is valid in the high-dimensional regime p9, critically exploiting structural log-concavity inherited from the prior over Gram matrices.
The key information-theoretic step is to compare the conditional mean squared overlap to the mutual information between the Gram matrix and the observed graph, leveraging generalization bounds from learning theory and strong subgaussianity properties.
Explicit Leading-Order Expansions
For p=Θ(1/n)0, the paper provides explicit expansions:
- The leading contribution to p=Θ(1/n)1 is proportional to p=Θ(1/n)2, with p=Θ(1/n)3.
- For higher p=Θ(1/n)4, leading contributions correspond to sums over pairwise products (matchings) or degree-structured subgraphs (e.g., triangles, quadruples) weighted by powers of p=Θ(1/n)5 and p=Θ(1/n)6 (the threshold scaling).
The sharp expansion of p=Θ(1/n)7 is pivotal for matching the lower bound to the upper bound, and controlling the remainder is essential for quantitative sharpness.
Strong Log-Concavity and Concentration
- The posterior law of p=Θ(1/n)8 (the Gram matrix), after truncation on operator norm, is shown to be p=Θ(1/n)9-strongly log-concave for d≪(nh(p))30.
- Gaussian concentration applied to Lipschitz functions of d≪(nh(p))31 under such log-concave laws ensures that high moments of posterior overlaps are negligible under the scaling d≪(nh(p))32.
Implications and Open Problems
Theoretical Implications
This work establishes the full sharp detection threshold for RGGs with binary kernel in the d≪(nh(p))33 regime for almost all nontrivial edge probabilities. In the context of random graph hypothesis testing, this:
- Solidifies the tight interplay between geometry, dimension, and statistical identifiability in high dimensions.
- Validates the signed triangle count as being optimal among polynomial tests, up to this regime, by matching the information-theoretic threshold.
- Provides a versatile framework that can be generalized to broader classes of latent space models, including those with isotropic Gaussian priors and random geometric graphs with smooth kernels.
Practical Significance
The result delineates the high-dimensional limits of network inference for latent space models, giving practitioners a precise prescription for when latent geometry is recoverable from network data and when the observed structure is indistinguishable from randomness.
Future Directions
- Extension to d≪(nh(p))34 Regime: The current method fundamentally relies on strong log-concavity, which does not extend to the underdetermined case d≪(nh(p))35 due to the degeneracy of the Gram matrix. New techniques are required to close this gap.
- Computational-to-Statistical Gaps: While the signed triangle count achieves the detection threshold in this regime, it is an open question whether computationally efficient tests remain optimal for all parameter ranges, or if algorithmic lower bounds (Low-Degree Likelihood/SoS hierarchies) reveal statistically hard but computationally easy phases.
- Universality and Generalization: Applying this posterior-based log-concavity framework to more general network models, including non-spherical latent spaces or more complex edge functions, remains a promising direction.
- Robustness to Model Misspecification: Extending these results to latent space models with noise, missing data, or adversarial perturbations is crucial for real-world applicability.
Conclusion
This paper conclusively establishes the sharp detection threshold for random geometric graphs in the high-dimensional regime d≪(nh(p))36, providing a definitive characterization of the phase transition for geometric detectability as a function of the intrinsic entropy of the edge distribution and the ambient dimension. The techniques, grounded in information theory, posterior overlap concentration, and high-dimensional probability, yield a blueprint for analyzing a wide class of relational models and latent space inference problems (2607.02013).