---
title: Erdős–Rényi Comparison Graphs
url: https://www.emergentmind.com/topics/erdos-renyi-comparison-graphs
type: topic
---

# Erdős–Rényi Comparison Graphs

Erdős–Rényi Comparison Graphs provide a foundational setting for the analysis of statistical and computational questions involving two random graphs, particularly regarding the detection of structural correlation or recovery of latent alignments. These models form a precise mathematical framework for studying graph isomorphism, matching, and detection thresholds in settings where two graphs exhibit joint dependencies—most classically under edge-wise independence or mild correlation rooted in a hidden vertex correspondence. Broadly, the literature addresses both the information-theoretic and computational aspects of detecting and recovering such structure, elaborating tight thresholds for feasibility and algorithmic performance.

## 1. Model Definitions and Correlated Random Graph Ensembles

The canonical Erdős–Rényi comparison model involves two graphs $G_1, G_2$ on $n$ vertices, where each is a realization of $\mathbf{G}(n,p)$, and the joint law introduces correlation via a hidden permutation or parent graph structure. Two primary constructions are studied:

- **Edge-correlated model:** Given parameters $q \in (0,1)$ and correlation $\rho \in [0,1]$, a hidden permutation $\pi$ on $[n]$ induces dependency such that
  - For each edge $\{i, j\}$,
    $$
    \mathrm{P}[ \{i, j\} \in E(G_2) | \{\pi(i),\pi(j)\} \in E(G_1) ] = q + \rho(1-q),
    $$
    $$
    \mathrm{P}[ \{i, j\} \in E(G_2) | \{\pi(i),\pi(j)\} \notin E(G_1) ] = q - \rho q,
    $$
  ensuring $\mathrm{Cov}( 1_{\{\pi(i),\pi(j)\} \in G_1}, 1_{\{i,j\} \in G_2} ) = \rho$ while preserving the Erdős–Rényi marginals [2311.15931].
- **Subsampled-parent model ("parent graph" comparison):** A parent $G_0 \sim \mathbf{G}(n,p)$ yields $G_1$ and (after hidden permutation) $G_2$ via independent edge subsampling with probability $s$, imposing limited overlaps and thereby a parameterized degree of correlation. This setting naturally encodes recovery problems for latent alignments and detection of shared origin [2203.14573, 2502.12077].

These models underpin a host of comparison tasks, including hypothesis testing (independent vs. correlated), exact or partial permutation recovery, and algorithmic graph matching under null and alternative hypotheses.

## 2. Hypothesis Testing and Detection Thresholds

The archetypal *comparison* or *detection* problem is formulated as a two-sample hypothesis test between the null $H_0$ (independent marginals) and alternative $H_1$ (joint dependency through hidden correlation):

- $H_0$: $(G_1, G_2)$ are independent, each $\mathbf{G}(n, q)$.
- $H_1$: $(G_1, G_2)$ are edge-correlated as described in the previous section, or subsampled from a common parent.

**Sharp information-theoretic detection thresholds** in the sparse regime $p = n^{-\alpha + o(1)}$, $\alpha \in (0,1]$ are obtained via combinatorial and variational analysis of the structure in the *intersection graph* (i.e., the graph formed by common edges after optimal alignment). A critical result is that detection feasibility undergoes a phase transition characterized via:
$$
s_c(\alpha) = n^{-(1-\alpha)/2+o(1)}, \quad \text{or equivalently} \quad n p s^2 = A^*, \quad A^* = \varphi^{-1}(2)
$$
where $\varphi$ is a variationally defined "densest-subgraph rate" function [2203.14573]. If $nps^2 \gg A^*$, detection is possible with vanishing error; otherwise, no test significantly outperforms random guessing.

The optimal (information-theoretic) test relies on the **densest subgraph statistic** over all permutations of vertex alignments:
$$
T_{\mathrm{stat}}(G_1, G_2) = \max_{T \in S_n} \max_{U \subseteq V, |U| > n/\log n} \frac{|E_{T}(U)|}{|U|}
$$
where $E_T(U)$ are common edges in the induced subgraph under $T$. Exhaustive search is required for optimality, rendering the test infeasible for large $n$ [2203.14573].

## 3. Algorithmic Complexity: Limits and Low-Degree Barriers

Despite the existence of sharp information-theoretic boundaries for detection and recovery, significant computational barriers are present. The main lines of evidence are as follows:

- **Low-degree hardness:** For the edge-correlated model, even sophisticated polynomial-time algorithms (including subgraph count statistics, spectral methods, and message passing) can be expressed as low-degree polynomials in the adjacency entries. It is proved that no such method of degree $d = O(\rho^{-1})$ (dense regime) or $d = \exp\left\{ o( (\log n)/\log(nq) \wedge \sqrt{\log n} ) \right\}$ (sparse regime) suffices to reliably distinguish $H_0$ from $H_1$ when $\rho < \sqrt{\alpha}$, where $\alpha \approx 0.338$ is Otter's constant [2311.15931]. This suggests that the best-known polynomial-time methods achieve the computational frontier under standard conjectures, including for exact or partial permutation recovery.

- **Polynomial-time approximation schemes (PTAS):** In the case of *independent* comparison (i.e., no underlying correlation), the maximum edge overlap between $G_1$ and $G_2$ over all permutations can be tightly approximated in randomized polynomial time. The maximal achievable overlap is $\sim n/(2\alpha - 1)$ for $p = n^{-\alpha}$, $\alpha \in (1/2, 1)$, and a PTAS attains this up to any fixed $\varepsilon$ [2210.07823], reflecting a lack of information-computation gap in the purely random matching case.

## 4. Recovery of Latent Structure: Exact and Partial Matching

Recovery problems in the correlated comparison setting seek not just detection but estimation of the underlying permutation (or as many correct vertex pairs as possible):

- **Partial recovery in the sparse regime:** Given $p = n^{-\alpha + o(1)}$, $s$ such that $nps^2 = \lambda = O(1)$, the maximum fraction $\rho(\alpha, \lambda)$ of correctly matched pairs is asymptotically characterized by the limiting load distribution $\mu_\lambda$, specifically its upper tail $F_\lambda(t)$. Except for a countable collection of critical $\alpha$, one has
  $$
  \rho(\alpha, \lambda) = F_\lambda(\alpha^{-1})
  $$
  The "balanced load" function on the intersection graph, rooted in combinatorial optimization over densest subgraphs, governs achievable performance [2502.12077]. Computationally efficient procedures achieving this bound remain open, with low-degree lower bounds suggesting polynomial-time intractability at small $\lambda$.

## 5. Connections to Universality and Random Matrix Theory

The comparison paradigm also yields insights into the universality of large-scale random graph statistics. Classical results show that both the component-size scaling in critical random geometric graphs and the spectral statistics of sparse Erdős–Rényi graphs coincide with those of canonical random matrix ensembles:

- **Critical component scaling:** The 2D-torus random geometric graph exhibits the same $n^{2/3}$ scaling and parabolic-drift diffusion limit for largest component sizes as the Erdős–Rényi model in the critical window,
  $$
  \big(n^{-2/3} C_1, n^{-2/3} C_2, \ldots\big) \xrightarrow{\ell^2} (\xi_1, \xi_2, \ldots)
  $$
  with $(\xi_i)$ given by Brownian excursions with parabolic drift, and $C_{\max}$ converging to a Tracy–Widom-type law [2308.07696].

- **Spectral universality:** For $pN \gg N^{2/3}$, the bulk and edge eigenvalue statistics of Erdős–Rényi adjacency matrices align with those of the Gaussian Orthogonal Ensemble (GOE) after appropriate normalization. This includes convergence of correlation functions and Tracy–Widom fluctuations for the largest eigenvalues [1103.3869]. These universality results depend on local semicircle laws, Dyson Brownian motion, and Green-function comparison via moment-matching.

## 6. Open Questions and Future Directions

Several research directions remain at the frontier of Erdős–Rényi comparison theory:

- Can low-degree hardness be strengthened to full sum-of-squares lower bounds, making computational impossibility evidence more robust [2311.15931]?
- What new algorithmic strategies, if any, could breach the low-degree barrier, especially for detection thresholds below $\rho = \sqrt{\alpha}$?
- Under what regimes do information-theoretic and computational thresholds coincide or diverge, and how does this depend on specifics of the graph parameters and comparison model?
- What are the precise locations and structural mechanisms of the exceptional, atom-driven jumps in recoverability identified for the load distribution in partial recovery?

These questions represent ongoing efforts to delineate sharp boundaries in random graph inference and matching, both statistically and computationally.

## 7. References to Key Results

| Focus Area                        | Reference                                                        | arXiv ID      |
|------------------------------------|------------------------------------------------------------------|---------------|
| Detection/Recovery thresholds      | Ding & Du, "Detection threshold for correlated Erdős-Rényi..."   | 2203.14573    |
| Low-degree hardness                | Ding, Du, & Li, "Low-Degree Hardness of Detection..."            | 2311.15931    |
| Critical scaling, universality     | "Scaling of Components in Critical Geometric Random Graphs..."   | 2308.07696    |
| Spectral statistics, universality  | Erdős, Yau et al., "Spectral Statistics of Erdős–Rényi Graphs II"| 1103.3869     |
| Partial recovery, balanced load    | Du, "Optimal recovery of correlated Erdős–Rényi graphs"          | 2502.12077    |
| PTAS for maximal overlap           | "A polynomial-time approximation scheme for the maximal overlap..."| 2210.07823  |

Source: https://www.emergentmind.com/topics/erdos-renyi-comparison-graphs