Papers
Topics
Authors
Recent
Search
2000 character limit reached

A Local Central Limit Theorem for Clique Counts in Sparse Random Graphs

Published 17 Aug 2026 in math.CO and math.PR | (2608.16882v1)

Abstract: Let XHX_H denote the number of copies of a fixed graph HH in Gn,pG_{n, p}. Gilmer and Kopparty conjectured that XHX_H satisfies a local central limit theorem (LCLT) provided that HH is connected, pn<sup>1/m(H)p \gg n<sup>{-1/m(H)}, and n<sup>2</sup>(1p)1n<sup>2</sup> (1-p) \gg 1, where m(H)m(H) is the maximum density. Following the work of Berkowitz, Sah and Sawhney confirmed this conjecture for every constant pp, leaving the regime where p=o(1)p=o(1) open. In this regime, the only case addressed in the literature is when H=K3H=K_3, where, in a paper, Araújo and Mattos confirmed the conjecture for p(4n<sup>1/2,</sup>1/2)p \in (4n<sup>{-1/2},</sup> 1/2). This, together with a general result of Röllin and Ross, essentially settles the conjecture for the triangle. We generalise these results by showing that an LCLT holds for H=KrH = K_r (for any fixed r3r \ge 3) in the regime n<sup>1/m(H)</sup>p1/2n<sup>{-1/m(H)}\ll</sup> p\leq 1/2, essentially settling the conjecture for cliques.

Summary

  • The paper proves a local central limit theorem for the number of copies of any fixed clique K_r in G_{n,p} throughout the essentially optimal range n^{-1/m(K_r)} ≪ p ≤ 1/2.
  • The authors combine Fourier inversion, Stein’s method, quantitative central limit bounds, and a refined decoupling argument to control both low- and high-frequency characteristic functions across four sparsity regimes.
  • The result settles sparse-regime anti-concentration for clique counts and extends local limit theory beyond triangles, while leaving analogous results for general connected graphs and non-complete strictly balanced graphs open.

The paper establishes a local central limit theorem (LCLT) for the number of copies of a fixed clique KrK_r in the Erdős–Rényi random graph Gn,pG_{n,p}, in an essentially optimal range of edge probabilities (2608.16882). Written by Cohen Antonir, Hoshen, and Zhukovskii, it resolves the Gilmer–Kopparty conjecture for cliques in the sparse regime p=o(1)p = o(1), which had remained open after prior work settled the dense case.

Background and context

Let XHX_H denote the number of labelled copies of a fixed graph HH in Gn,pG_{n,p}, with mean μH\mu_H and variance σH2\sigma_H^2. Ruciński proved that (XHμH)/σH(X_H - \mu_H)/\sigma_H converges in distribution to a standard normal whenever pn1/m(H)p \gg n^{-1/m(H)} and Gn,pG_{n,p}0, where Gn,pG_{n,p}1 is the maximum density. Gilmer and Kopparty conjectured that this integral central limit theorem can be strengthened to a local one: pointwise,

Gn,pG_{n,p}2

Progress on this conjecture proceeded regime by regime: Gilmer and Kopparty handled triangle counts at constant Gn,pG_{n,p}3; Röllin and Ross covered triangles above the appearance threshold; Berkowitz introduced the decoupling method and obtained quantitative LCLTs for triangles and cliques at constant Gn,pG_{n,p}4; Sah and Sawhney extended to all connected graphs at constant Gn,pG_{n,p}5; Sah, Sawhney, and Zhu obtained multivariate versions; and Araújo and Mattos recently closed the gap for the triangle, confirming the conjecture fully for Gn,pG_{n,p}6. The sparse regime for general cliques was untouched.

Main result

For fixed integer Gn,pG_{n,p}7, writing Gn,pG_{n,p}8 for the number of Gn,pG_{n,p}9-copies, p=o(1)p = o(1)0 and p=o(1)p = o(1)1 its mean and standard deviation, and p=o(1)p = o(1)2, the main theorem states that for all p=o(1)p = o(1)3,

p=o(1)p = o(1)4

where p=o(1)p = o(1)5 is the support lattice of p=o(1)p = o(1)6. Since p=o(1)p = o(1)7, this covers the full range where a CLT is expected, up to the upper endpoint p=o(1)p = o(1)8 (the complement regime follows by symmetry considerations not addressed here). The theorem comes with explicit error rates across four sub-regimes of p=o(1)p = o(1)9, e.g. XHX_H0 when XHX_H1 with XHX_H2, and XHX_H3 just above the appearance threshold. The authors note these rates are not optimised — polynomial factors XHX_H4 could likely be replaced by polylogarithms — but even unoptimised, the constant-XHX_H5 error improves on Berkowitz's and Sah–Sawhney's bounds.

A corollary worth emphasising: because an LCLT determines XHX_H6 asymptotically, the result settles the anti-concentration question for clique counts in the sparse regime, where previously only the triangle case was understood.

Proof architecture

The argument follows the now-standard Fourier route: by Fourier inversion on the lattice, the LCLT reduces to showing

XHX_H7

Low frequencies are handled by combining Stein's method with the quantitative CLT of Barbour, Karoński, and Ruciński, yielding an integral bound of order XHX_H8 where XHX_H9. The substantive contribution lies in three lemmas controlling high frequencies, each covering a different density regime:

  • Above the 2-density (HH0): HH1 for HH2.
  • Dense-ish sparse regime (HH3): the same decay for HH4.
  • Near the appearance threshold: HH5, where HH6 is chosen so that HH7.

The regime split reflects a structural dichotomy governed by the 2-density: below it, a typical edge lies in no HH8 and copies are typically edge-disjoint, so HH9 behaves like a sum of nearly independent variables; above it, an edge lies in Gn,pG_{n,p}0 cliques and dependencies become essential.

The decoupling method and its extension

The core technique, inherited from Berkowitz and refined by Sah–Sawhney, is a separation-of-variables argument: viewing Gn,pG_{n,p}1 as a degree-Gn,pG_{n,p}2 polynomial in independent edge indicators, repeated Cauchy–Schwarz ("decoupling") bounds Gn,pG_{n,p}3 by a power of the characteristic function of a signed sum over "rainbow" cliques with respect to a partition Gn,pG_{n,p}4 and colouring scheme Gn,pG_{n,p}5. The signed sum splits into Gn,pG_{n,p}6 (cliques using exactly one vertex in each relevant part) and an error term Gn,pG_{n,p}7 capturing the remaining dependencies.

Two reductions then apply. First, Gn,pG_{n,p}8 equals a weighted sum Gn,pG_{n,p}9 of independent Bernoulli indicators over edges μH\mu_H0, whose characteristic function decays via the elementary bound μH\mu_H1. A delicate balancing act is required: weights must be small enough that μH\mu_H2 (so the Bernoulli bound is effective), yet large enough that sufficiently many edges carry typical weight — achieved via second- and fourth-moment estimates plus Paley–Zygmund, and a matching decomposition into perfect matchings to handle dependence between weights.

Second, the error term is controlled by bounding μH\mu_H3 for large μH\mu_H4. Non-vanishing contributions come only from tuples of rainbow cliques in which every vertex of μH\mu_H5 appears in zero or at least two members of the tuple. The paper's most substantial technical novelty is the combinatorial analysis of such tuples: tight lower bounds on the number of edges in their union as a function of the vertex distribution across partition parts, followed by an optimisation showing

μH\mu_H6

At constant μH\mu_H7 this term is degenerate and easy; in the sparse regime its sharp control is what makes the theorem possible.

Two further adaptations deserve mention. For μH\mu_H8, weights of vertex-disjoint edges are no longer independent (unlike the triangle case exploited by Araújo and Mattos); the authors circumvent this by first revealing μH\mu_H9 and proving concentration of clique-extension counts before exposing edge ownership. Near the appearance threshold, dependencies vanish entirely — most edges lie in at most one rainbow copy — so the error-term analysis becomes unnecessary and the proof simplifies considerably, relying instead on Kim–Vu polynomial concentration to show that σH2\sigma_H^20 edges carry weight σH2\sigma_H^21.

Finally, the intervals of frequencies covered by varying the parameters σH2\sigma_H^22 of the partition must be stitched together: the authors define σH2\sigma_H^23 via a dominance condition among the quantities σH2\sigma_H^24 and verify that the union of covered intervals spans σH2\sigma_H^25, with the extreme frequencies near σH2\sigma_H^26 handled by a separate argument choosing σH2\sigma_H^27 so that expected rainbow-copy counts per edge are σH2\sigma_H^28.

Limitations and open questions

Several caveats are stated plainly. The result covers only cliques; extending the LCLT to arbitrary connected graphs σH2\sigma_H^29 in the sparse regime remains open, and the paper's techniques lean heavily on the strict balancedness of (XHμH)/σH(X_H - \mu_H)/\sigma_H0 and (XHμH)/σH(X_H - \mu_H)/\sigma_H1 minus an edge. The upper endpoint is (XHμH)/σH(X_H - \mu_H)/\sigma_H2, with the complementary range presumably following by symmetry but not treated here. The quantitative error terms are admittedly unoptimised — the polylogarithmic exponents (e.g. (XHμH)/σH(X_H - \mu_H)/\sigma_H3) and the polynomial losses (XHμH)/σH(X_H - \mu_H)/\sigma_H4 are not tight, and tightening them would require refining both the moment computations and the low-frequency Stein bound. The conjecture for disconnected graphs fails in general, so any generalisation must retain connectivity. Whether the Gilmer–Kopparty conjecture holds for strictly balanced but non-complete graphs below the 2-density threshold, where the independence structure differs qualitatively, is left unanswered.

Conclusion

This paper completes the verification of the Gilmer–Kopparty local central limit conjecture for clique counts across essentially the entire admissible range of (XHμH)/σH(X_H - \mu_H)/\sigma_H5, providing the first LCLT for subgraph counts in the sparse regime beyond triangles. Its principal technical contributions — the high-moment analysis of the decoupling error under sparsity, and the handling of weight dependencies for (XHμH)/σH(X_H - \mu_H)/\sigma_H6 — extend the decoupling framework to regimes where dependencies between subgraph copies are no longer negligible, and yield as a byproduct an asymptotic answer to the anti-concentration problem for sparse clique counts.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Tweets

Sign up for free to view the 1 tweet with 0 likes about this paper.