Papers
Topics
Authors
Recent
Search
2000 character limit reached

Resolution of the Detection Threshold Conjecture for Random Geometric Graphs in the $d>n$ Regime

Published 2 Jul 2026 in math.PR and math.ST | (2607.02013v1)

Abstract: A random geometric graph (RGG) is generated by first sampling latent points x1,,xnx_1,\ldots,x_n independently and uniformly from the unit sphere in R<sup>d\mathbb{R}<sup>d, and then connecting each pair (i,j)(i,j) if xi,xj\langle x_i,x_j\rangle exceeds some threshold ττ. We study the sharp detection threshold -- the largest dimension at which the RGG can be statistically distinguished from the Erdős--Rényi graph with the same edge density pp. This threshold is conjectured to be d(nh(p))<sup>3d \asymp (nh(p))<sup>3, where h(p)=plog1p+(1p)log11ph(p)=p \log \frac{1}{p} + (1-p) \log \frac{1}{1-p} is the binary entropy function. Previous works proved this conjecture for dense graphs with constant pp and, up to polylogarithmic factors, very sparse graphs with p=Θ(1/n)p=Θ(1/n). In this paper, we prove that detection is impossible when d(nh(p))<sup>3d\gg (nh(p))<sup>3 and d(1+ε)nd\ge (1+ε) n for any constant $ε&gt;0$, thereby resolving the conjecture in the regime pn<sup>2/3/log</sup>np\gtrsim n<sup>{-2/3}/\log</sup> n and improving upon the state of the art in the regime 1/npn<sup>2/3/log</sup>n1/n \ll p \ll n<sup>{-2/3}/\log</sup> n. The key to our proof is a sharp analysis of the posterior distribution of the latent points given the observed graph, obtained through an information-theoretic comparison argument combined with strong log-concavity.

Summary

  • The paper demonstrates that in the regime d > n, detection is impossible when d >> (n h(p))^3, thereby resolving the longstanding detection threshold conjecture.
  • It employs a rigorous KL divergence chain-rule, precise polynomial expansions of interaction functions, and posterior log-concavity to derive sharp statistical bounds.
  • The findings imply that latent geometric information is lost in high dimensions, guiding practical network inference in latent space models.

Detection Thresholds for Random Geometric Graphs: Resolution in the High-Dimensional Regime

Introduction and Problem Overview

This paper addresses the intrinsic question of when random geometric graphs (RGGs), constructed via a hard-thresholding scheme on latent points sampled uniformly from the unit sphere in Rd\mathbb{R}^d, can be statistically distinguished from Erdős–Rényi (ER) random graphs with the same edge density. Specifically, the focus is on the "detection threshold": the critical point in the parameter space (dimension dd, number of nodes nn, and edge probability pp) where geometry becomes information-theoretically undetectable from pure random connectivity.

Previous works have established sharp detection thresholds in the dense regime (constant pp) and, up to polylogarithmic factors, for extremely sparse graphs (p=Θ(1/n)p = \Theta(1/n)). However, a conjecture had remained open for the intermediate sparse regime: distinguishing between RGG and ER is conjectured to be possible if and only if d(nh(p))3d \ll (nh(p))^3, where h(p)=plog(1/p)+(1p)log(1/(1p))h(p) = p\log(1/p) + (1-p)\log(1/(1-p)) is the binary entropy function. This paper resolves the conjecture for d>nd > n and pn2/3/lognp \gtrsim n^{-2/3}/\log n, and improves bounds for the regime dd0.

Model Setup and Key Results

An dd1-node RGG is formed as follows:

  • Latent vectors dd2 are independently sampled from the unit sphere dd3.
  • For a threshold dd4 defined such that dd5, an edge exists between node dd6 and dd7 if and only if dd8.

The detection problem is formulated as a binary hypothesis test: given a sample graph, decide between

  • dd9: the sample is RGG (nn0)
  • nn1: the sample is ER graph nn2 (nn3) with indices nn4 and nn5 for the corresponding distributions.

The central result is: Detection between RGG and ER graphs is impossible in total variation when nn6 and nn7 for any fixed nn8. This resolves the detection threshold conjecture in the regime nn9, establishing information-theoretic impossibility at the conjectured scaling; the bound is tight up to constant factors.

Numerical and Technical Claims

  • For the regime pp0 and pp1, the Kullback-Leibler divergence satisfies pp2.
  • Pinsker’s inequality then yields pp3, showing indistinguishability in total variation.
  • In the very sparse regime pp4, the bound pp5 still improves on previous contours, but a complete sharp threshold for pp6 is left open.

Methods and Technical Innovations

Chain Rule Expansion and Posterior Analysis

The proof leverages a chain-rule expansion of the KL divergence, incrementally revealing vertices and their incident edges, framing the problem in terms of posterior moments of local functions of the latent configuration conditioned on the observed graph.

A key technical advance is the direct analysis of the posterior distribution pp7 of the latent positions pp8 given the RGG realization pp9. This sidesteps the overly revealing approach—used in prior works—of conditioning on all but a fixed subset of latent positions, which loosens information-theoretic bounds.

Three-Stage Proof Sketch

  1. KL Expansion: Decompose the KL divergence into a sum over vertices, expressing each term as an expectation of squared local "interaction functions" pp0 (functions of up to pp1 latent points), where pp2 ranges over small subgraphs.
  2. Approximation of Interaction Functions: Obtain precise polynomial expansions of pp3 in terms of Gram matrix entries pp4 for pp5, bounding the remainder with strong (log-concavity-based) concentration results.
  3. Posterior Moment Control: Use transportation inequalities and strong log-concavity to bound the mean and high moments of the posterior overlap pp6 for independent samples pp7 from pp8. This step is valid in the high-dimensional regime pp9, critically exploiting structural log-concavity inherited from the prior over Gram matrices.

The key information-theoretic step is to compare the conditional mean squared overlap to the mutual information between the Gram matrix and the observed graph, leveraging generalization bounds from learning theory and strong subgaussianity properties.

Explicit Leading-Order Expansions

For p=Θ(1/n)p = \Theta(1/n)0, the paper provides explicit expansions:

  • The leading contribution to p=Θ(1/n)p = \Theta(1/n)1 is proportional to p=Θ(1/n)p = \Theta(1/n)2, with p=Θ(1/n)p = \Theta(1/n)3.
  • For higher p=Θ(1/n)p = \Theta(1/n)4, leading contributions correspond to sums over pairwise products (matchings) or degree-structured subgraphs (e.g., triangles, quadruples) weighted by powers of p=Θ(1/n)p = \Theta(1/n)5 and p=Θ(1/n)p = \Theta(1/n)6 (the threshold scaling).

The sharp expansion of p=Θ(1/n)p = \Theta(1/n)7 is pivotal for matching the lower bound to the upper bound, and controlling the remainder is essential for quantitative sharpness.

Strong Log-Concavity and Concentration

  • The posterior law of p=Θ(1/n)p = \Theta(1/n)8 (the Gram matrix), after truncation on operator norm, is shown to be p=Θ(1/n)p = \Theta(1/n)9-strongly log-concave for d(nh(p))3d \ll (nh(p))^30.
  • Gaussian concentration applied to Lipschitz functions of d(nh(p))3d \ll (nh(p))^31 under such log-concave laws ensures that high moments of posterior overlaps are negligible under the scaling d(nh(p))3d \ll (nh(p))^32.

Implications and Open Problems

Theoretical Implications

This work establishes the full sharp detection threshold for RGGs with binary kernel in the d(nh(p))3d \ll (nh(p))^33 regime for almost all nontrivial edge probabilities. In the context of random graph hypothesis testing, this:

  • Solidifies the tight interplay between geometry, dimension, and statistical identifiability in high dimensions.
  • Validates the signed triangle count as being optimal among polynomial tests, up to this regime, by matching the information-theoretic threshold.
  • Provides a versatile framework that can be generalized to broader classes of latent space models, including those with isotropic Gaussian priors and random geometric graphs with smooth kernels.

Practical Significance

The result delineates the high-dimensional limits of network inference for latent space models, giving practitioners a precise prescription for when latent geometry is recoverable from network data and when the observed structure is indistinguishable from randomness.

Future Directions

  • Extension to d(nh(p))3d \ll (nh(p))^34 Regime: The current method fundamentally relies on strong log-concavity, which does not extend to the underdetermined case d(nh(p))3d \ll (nh(p))^35 due to the degeneracy of the Gram matrix. New techniques are required to close this gap.
  • Computational-to-Statistical Gaps: While the signed triangle count achieves the detection threshold in this regime, it is an open question whether computationally efficient tests remain optimal for all parameter ranges, or if algorithmic lower bounds (Low-Degree Likelihood/SoS hierarchies) reveal statistically hard but computationally easy phases.
  • Universality and Generalization: Applying this posterior-based log-concavity framework to more general network models, including non-spherical latent spaces or more complex edge functions, remains a promising direction.
  • Robustness to Model Misspecification: Extending these results to latent space models with noise, missing data, or adversarial perturbations is crucial for real-world applicability.

Conclusion

This paper conclusively establishes the sharp detection threshold for random geometric graphs in the high-dimensional regime d(nh(p))3d \ll (nh(p))^36, providing a definitive characterization of the phase transition for geometric detectability as a function of the intrinsic entropy of the edge distribution and the ambient dimension. The techniques, grounded in information theory, posterior overlap concentration, and high-dimensional probability, yield a blueprint for analyzing a wide class of relational models and latent space inference problems (2607.02013).

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 1 tweet with 4 likes about this paper.