Upper Tail Problem in Random Graphs
- Upper Tail Problem is defined as the asymptotic study of probabilities that a random variable significantly exceeds its typical value, formulated via entropy minimization and variational principles.
- It employs methodologies including optimization over weighted graphs, high-moments techniques, and entropic-stability to reveal the competition between diffuse and localized strategies in sparse regimes.
- The approach extends to diverse settings such as directed graphs, hypergraphs, and spectral observables, providing insights into rare-event geometry and optimal witness classification.
In probabilistic combinatorics, the upper tail problem asks for asymptotics of probabilities that a random quantity substantially exceeds its typical value. In its classical form, for a fixed graph and , one studies
where is the number of copies of in , together with the logarithmic rate and the structure of the random object conditioned on this rare event. Across sparse random graphs, directed graphs, hypergraphs, and spectral observables, the subject has developed around a recurrent competition between diffuse Poisson-type fluctuations and localized planting mechanisms such as cliques, hubs, stars, or mixed hubs (Bhattacharya et al., 2015, Ganguly et al., 2022, Harel et al., 2019).
1. Formal definition and variational formulation
For a fixed graph , one standard normalization is
In the sparse regime, a central objective is to identify the first-order asymptotics of . A major structural reduction expresses the problem through an entropy minimization over weighted graphs. Writing 0 for the family of symmetric 1 matrices 2 with 3 and 4, one defines the relative-entropy functional 5 and the weighted homomorphism density 6, and then considers
7
For 8, the nonlinear large-deviation principle reduces the upper tail to this discrete variational problem: 9 Thus, solving 0 to leading order yields 1 (Bhattacharya et al., 2015).
This variational perspective has two consequences. First, it shifts the problem from direct probability estimates to an optimization problem with entropic cost 2. Second, it makes the upper tail problem intrinsically structural: the optimizer describes the weighted graph most likely to realize the excess count. In later developments, analogous formulations reappear for directed graphs, hypergraphs, constrained sparse graph models, and spectral statistics (Park, 2024, Liu et al., 2019, Bhattacharya et al., 2019, Bhattacharya et al., 2018).
2. Dominant mechanisms: clique, anti-clique, hub, and core
In the sparse graph case, the leading constant in the exponent is often governed by the independence polynomial. If 3 denotes the induced subgraph of 4 on the vertices of degree 5, and
6
then the sparse-regime solution identifies a parameter 7 through
8
If 9 is irregular, then
0
If 1 is 2-regular on 3 vertices, then two competing constructions coexist: an anti-clique of cost 4 and a planted clique of cost 5, giving
6
For triangles,
7
with a clique-versus-anti-clique transition at 8 (Bhattacharya et al., 2015).
A complementary framework is the high-moments and entropic-stability method. For a polynomial 9 with nonnegative coefficients on the 0-biased discrete hypercube, one defines a rate function 1 by minimizing the cost of conditioning a small coordinate set 2 to be present while forcing
3
If the family of such cores is entropically stable and 4, then
5
and, with high probability conditioned on the upper-tail event, one of the near-optimal cores appears in the sample (Harel et al., 2019).
These two viewpoints are closely aligned. The variational formulation identifies the entropic minimizer at the level of weighted graphs, while the entropic-stability framework identifies the discrete witnesses that realize the same rare event. This suggests that the upper tail problem is fundamentally about classifying optimal witnesses, not merely estimating probabilities.
3. Canonical sparse-regime behaviors and solved examples
Several solved models exhibit a trichotomy or two-term competition between diffuse and localized strategies. For 6-stars 7, letting
8
and
9
one has
0
where
1
and
2
The same work interprets this as a “Poisson 3 localization” transition and gives a first complete sparse solution for an irregular graph (Akhmejanova et al., 6 Jan 2025).
For 4-term arithmetic progressions in a 5-biased subset of 6, the 7-plane is partitioned into three disjoint regimes. In the Gaussian regime,
8
In the Poisson regime,
9
where 0. In the localised regime,
1
and whenever 2,
3
The proof combines exponential tilting, martingale concentration, and an entropic-stability lemma for small dense seeds (Harel et al., 2024).
Induced subgraph counts introduce additional phenomena. For induced 4, the sparse theory has a complete first-order asymptotic, but in the range
5
the upper tail does not admit a naive mean-field approximation. In that regime, no single product-tilt can match the true lower bound, and the entropy minimizer ceases to be a two-point tilt in 6; instead one must condition on the appearance of “one of many” moderately sized bipartite subgraphs (Antonir, 2022).
For ordinary cycle counts, the classical 7-gap can be closed up to constants. If 8 is the 9-cycle and
0
then
1
Thus the exponent is of order 2, eliminating the extra 3 factor present in general Janson–Oleszkiewicz–Ruciński bounds (Raz, 2019).
4. The constant-average-degree triangle problem
A particularly sharp form of the upper tail problem arises when 4 with fixed 5. Let 6 be the number of triangles in 7. Then 8 weakly converges to the Poisson distribution with mean 9, and for fixed 0,
1
The upper-tail problem asks how fast 2 may grow before the Poisson-tail approximation breaks down, and what replaces it beyond that threshold (Ganguly et al., 2022).
The transition is sharp. With
3
the change of regime occurs when
4
In the subcritical regime,
5
with 6 and 7,
8
Thus, to first order, the Poisson-tail size remains correct. In the supercritical regime,
9
one has
0
and the dominant strategy becomes planting an almost-clique of size 1 (Ganguly et al., 2022).
The conditioned structure is equally explicit. For fixed 2, let
3
and
4
Whenever 5,
6
Moreover,
7
in the subcritical regime, while
8
in the supercritical regime. The rare event is therefore realized either by almost 9 vertex-disjoint triangles or by a medium-sized clique-like structure (Ganguly et al., 2022).
The proof combines several graph-theoretic ingredients. A connected triangle-induced graph spanned by 00 triangles has tree-excess 01, where 02 is concave and satisfies 03. If 04 is the event that there is a triangle-induced subgraph with 05 triangles, then
06
for 07. The event 08 can be certified by disjoint triangle-induced subgraphs, so the van den Berg–Kesten inequality reduces the probability estimate to an optimization over partitions of 09. A stability version of the Kruskal–Katona bound then explains why the supercritical optimizer is almost-clique-like. The result settles the constant-average-degree triangle upper tail problem and answers a question of Aldous (Ganguly et al., 2022).
5. Spectral, directed, hypergraph, and constrained variants
The upper tail problem extends beyond copy counts. For the largest eigenvalue 10 of the adjacency matrix of 11, if 12 and 13, then
14
In the same regime,
15
The exponent for 16 again encodes a clique-regime and an anti-clique-regime. In a more localized sparse regime, when
17
the joint upper tail of the top 18 eigenvalues becomes polynomial in 19: 20 and, conditioned on the event, the graph admits a decomposition into a disjoint union of stars and a spectrally negligible remainder (Bhattacharya et al., 2018, Bhattacharya et al., 2020).
For directed random graphs, the same general strategy leads to a directed variational problem over weighted digraphs. In the sparse regime 21, the upper-tail rate is asymptotically given by the corresponding entropy minimization. One obtains upper and lower bounds that differ by a constant factor of at most 22, and these coincide for several classes of digraphs, including triangles, stars, directed 23-cycles, and balanced digraphs (Park, 2024).
For 24-uniform hypergraphs, the sparse theory is more intricate than in graphs. A conjectural rate function based on compatible collections of mixed hubs and stable labelings was formulated for 25. This framework predicts that the optimizer may require genuinely mixed-hub structures rather than only cliques or 1-hubs, and a 26-uniform counterexample shows that naïve hub-only conjectures fail. Subsequent work confirmed the conjecture for complete 27-partite 28-graphs, tight cycles, and the Fano plane, and also established a general lower bound under explicit edge-covering conditions (Liu et al., 2019, Cook et al., 30 Sep 2025).
The same upper-tail logic persists under structural constraints on the ambient random graph. For sparse 29-regular graphs, sparse uniform 30-edge graphs, and inhomogeneous block models, the logarithmic decay rate is again described by an entropy minimization under homomorphism-count constraints, and joint upper tails for multiple graphs admit analogous variational formulas (Bhattacharya et al., 2019).
6. Methods, counterexamples, and open directions
A large part of the subject is method-driven. One important development is a BK-inequality based combinatorial sparsification method that recovers the missing 31 factor in settings where classical Kim–Vu and Janson–Ruciński arguments only yielded exponents of order 32. In a prototype weighted hypergraph theorem, the method gives
33
thereby matching the expected logarithmic correction in many sparse problems. The mechanism is deletion-based sparsification combined with the van den Berg–Kesten inequality and refined Chernoff bounds with minor dependencies (Warnke, 2016).
At the same time, the subject contains sharp negative results against overly simple universality claims. The DeMarco–Kahn upper tail conjecture predicted that two mechanisms—clustered planting and disjoint planting—should always determine the exponent. This is false. An infinite family of counterexamples is given by graphs 34 obtained from an 35-cycle by attaching 36 pendant edges at a single cycle-vertex. For these graphs, in a range
37
one has
38
revealing a third, locally-disjoint mechanism beyond the conjectured clustered and disjoint ones (Šileikis et al., 2018).
Current open directions are highly model-specific. For constant-average-degree triangles, proposed extensions include other fixed subgraphs, inhomogeneous random graphs with bounded average degree, and dynamic random graph processes (Ganguly et al., 2022). For spectral tails, open problems include the case 39 for 40, a joint large-deviation principle for 41, higher edge-eigenvalues, and lower tails in the sparse regime (Bhattacharya et al., 2018). For arithmetic progressions, the boundary cases 42, 43, and 44 remain delicate (Harel et al., 2024). For hypergraphs, the full mixed-hub variational theory is still incomplete outside the presently verified classes (Liu et al., 2019, Cook et al., 30 Sep 2025).
A plausible implication is that the modern upper tail problem is best understood not as a single theorem, but as a classification program for rare-event geometry. The probability exponent, the optimizer of the entropy functional, and the conditioned structure are increasingly treated as inseparable parts of the same object.