- The paper proves a local central limit theorem for the number of copies of any fixed clique K_r in G_{n,p} throughout the essentially optimal range n^{-1/m(K_r)} ≪ p ≤ 1/2.
- The authors combine Fourier inversion, Stein’s method, quantitative central limit bounds, and a refined decoupling argument to control both low- and high-frequency characteristic functions across four sparsity regimes.
- The result settles sparse-regime anti-concentration for clique counts and extends local limit theory beyond triangles, while leaving analogous results for general connected graphs and non-complete strictly balanced graphs open.
The paper establishes a local central limit theorem (LCLT) for the number of copies of a fixed clique Kr in the Erdős–Rényi random graph Gn,p, in an essentially optimal range of edge probabilities (2608.16882). Written by Cohen Antonir, Hoshen, and Zhukovskii, it resolves the Gilmer–Kopparty conjecture for cliques in the sparse regime p=o(1), which had remained open after prior work settled the dense case.
Background and context
Let XH denote the number of labelled copies of a fixed graph H in Gn,p, with mean μH and variance σH2. Ruciński proved that (XH−μH)/σH converges in distribution to a standard normal whenever p≫n−1/m(H) and Gn,p0, where Gn,p1 is the maximum density. Gilmer and Kopparty conjectured that this integral central limit theorem can be strengthened to a local one: pointwise,
Gn,p2
Progress on this conjecture proceeded regime by regime: Gilmer and Kopparty handled triangle counts at constant Gn,p3; Röllin and Ross covered triangles above the appearance threshold; Berkowitz introduced the decoupling method and obtained quantitative LCLTs for triangles and cliques at constant Gn,p4; Sah and Sawhney extended to all connected graphs at constant Gn,p5; Sah, Sawhney, and Zhu obtained multivariate versions; and Araújo and Mattos recently closed the gap for the triangle, confirming the conjecture fully for Gn,p6. The sparse regime for general cliques was untouched.
Main result
For fixed integer Gn,p7, writing Gn,p8 for the number of Gn,p9-copies, p=o(1)0 and p=o(1)1 its mean and standard deviation, and p=o(1)2, the main theorem states that for all p=o(1)3,
p=o(1)4
where p=o(1)5 is the support lattice of p=o(1)6. Since p=o(1)7, this covers the full range where a CLT is expected, up to the upper endpoint p=o(1)8 (the complement regime follows by symmetry considerations not addressed here). The theorem comes with explicit error rates across four sub-regimes of p=o(1)9, e.g. XH0 when XH1 with XH2, and XH3 just above the appearance threshold. The authors note these rates are not optimised — polynomial factors XH4 could likely be replaced by polylogarithms — but even unoptimised, the constant-XH5 error improves on Berkowitz's and Sah–Sawhney's bounds.
A corollary worth emphasising: because an LCLT determines XH6 asymptotically, the result settles the anti-concentration question for clique counts in the sparse regime, where previously only the triangle case was understood.
Proof architecture
The argument follows the now-standard Fourier route: by Fourier inversion on the lattice, the LCLT reduces to showing
XH7
Low frequencies are handled by combining Stein's method with the quantitative CLT of Barbour, Karoński, and Ruciński, yielding an integral bound of order XH8 where XH9. The substantive contribution lies in three lemmas controlling high frequencies, each covering a different density regime:
- Above the 2-density (H0): H1 for H2.
- Dense-ish sparse regime (H3): the same decay for H4.
- Near the appearance threshold: H5, where H6 is chosen so that H7.
The regime split reflects a structural dichotomy governed by the 2-density: below it, a typical edge lies in no H8 and copies are typically edge-disjoint, so H9 behaves like a sum of nearly independent variables; above it, an edge lies in Gn,p0 cliques and dependencies become essential.
The decoupling method and its extension
The core technique, inherited from Berkowitz and refined by Sah–Sawhney, is a separation-of-variables argument: viewing Gn,p1 as a degree-Gn,p2 polynomial in independent edge indicators, repeated Cauchy–Schwarz ("decoupling") bounds Gn,p3 by a power of the characteristic function of a signed sum over "rainbow" cliques with respect to a partition Gn,p4 and colouring scheme Gn,p5. The signed sum splits into Gn,p6 (cliques using exactly one vertex in each relevant part) and an error term Gn,p7 capturing the remaining dependencies.
Two reductions then apply. First, Gn,p8 equals a weighted sum Gn,p9 of independent Bernoulli indicators over edges μH0, whose characteristic function decays via the elementary bound μH1. A delicate balancing act is required: weights must be small enough that μH2 (so the Bernoulli bound is effective), yet large enough that sufficiently many edges carry typical weight — achieved via second- and fourth-moment estimates plus Paley–Zygmund, and a matching decomposition into perfect matchings to handle dependence between weights.
Second, the error term is controlled by bounding μH3 for large μH4. Non-vanishing contributions come only from tuples of rainbow cliques in which every vertex of μH5 appears in zero or at least two members of the tuple. The paper's most substantial technical novelty is the combinatorial analysis of such tuples: tight lower bounds on the number of edges in their union as a function of the vertex distribution across partition parts, followed by an optimisation showing
μH6
At constant μH7 this term is degenerate and easy; in the sparse regime its sharp control is what makes the theorem possible.
Two further adaptations deserve mention. For μH8, weights of vertex-disjoint edges are no longer independent (unlike the triangle case exploited by Araújo and Mattos); the authors circumvent this by first revealing μH9 and proving concentration of clique-extension counts before exposing edge ownership. Near the appearance threshold, dependencies vanish entirely — most edges lie in at most one rainbow copy — so the error-term analysis becomes unnecessary and the proof simplifies considerably, relying instead on Kim–Vu polynomial concentration to show that σH20 edges carry weight σH21.
Finally, the intervals of frequencies covered by varying the parameters σH22 of the partition must be stitched together: the authors define σH23 via a dominance condition among the quantities σH24 and verify that the union of covered intervals spans σH25, with the extreme frequencies near σH26 handled by a separate argument choosing σH27 so that expected rainbow-copy counts per edge are σH28.
Limitations and open questions
Several caveats are stated plainly. The result covers only cliques; extending the LCLT to arbitrary connected graphs σH29 in the sparse regime remains open, and the paper's techniques lean heavily on the strict balancedness of (XH−μH)/σH0 and (XH−μH)/σH1 minus an edge. The upper endpoint is (XH−μH)/σH2, with the complementary range presumably following by symmetry but not treated here. The quantitative error terms are admittedly unoptimised — the polylogarithmic exponents (e.g. (XH−μH)/σH3) and the polynomial losses (XH−μH)/σH4 are not tight, and tightening them would require refining both the moment computations and the low-frequency Stein bound. The conjecture for disconnected graphs fails in general, so any generalisation must retain connectivity. Whether the Gilmer–Kopparty conjecture holds for strictly balanced but non-complete graphs below the 2-density threshold, where the independence structure differs qualitatively, is left unanswered.
Conclusion
This paper completes the verification of the Gilmer–Kopparty local central limit conjecture for clique counts across essentially the entire admissible range of (XH−μH)/σH5, providing the first LCLT for subgraph counts in the sparse regime beyond triangles. Its principal technical contributions — the high-moment analysis of the decoupling error under sparsity, and the handling of weight dependencies for (XH−μH)/σH6 — extend the decoupling framework to regimes where dependencies between subgraph copies are no longer negligible, and yield as a byproduct an asymptotic answer to the anti-concentration problem for sparse clique counts.