Papers
Topics
Authors
Recent
Search
2000 character limit reached

Upper Tail Problem in Random Graphs

Updated 14 July 2026
  • Upper Tail Problem is defined as the asymptotic study of probabilities that a random variable significantly exceeds its typical value, formulated via entropy minimization and variational principles.
  • It employs methodologies including optimization over weighted graphs, high-moments techniques, and entropic-stability to reveal the competition between diffuse and localized strategies in sparse regimes.
  • The approach extends to diverse settings such as directed graphs, hypergraphs, and spectral observables, providing insights into rare-event geometry and optimal witness classification.

In probabilistic combinatorics, the upper tail problem asks for asymptotics of probabilities that a random quantity substantially exceeds its typical value. In its classical form, for a fixed graph HH and GGn,pG\sim\mathcal G_{n,p}, one studies

Pr(XH(1+δ)E[XH]),\Pr\bigl(X_H\ge (1+\delta)\,\mathbb E[X_H]\bigr),

where XHX_H is the number of copies of HH in GG, together with the logarithmic rate logPr()-\log \Pr(\cdot) and the structure of the random object conditioned on this rare event. Across sparse random graphs, directed graphs, hypergraphs, and spectral observables, the subject has developed around a recurrent competition between diffuse Poisson-type fluctuations and localized planting mechanisms such as cliques, hubs, stars, or mixed hubs (Bhattacharya et al., 2015, Ganguly et al., 2022, Harel et al., 2019).

1. Formal definition and variational formulation

For a fixed graph HH, one standard normalization is

RH(n,p,δ):=logPr(XH(1+δ)E[XH]).R_H(n,p,\delta):=-\log\Pr\bigl(X_H\ge (1+\delta)\,\mathbb E[X_H]\bigr).

In the sparse regime, a central objective is to identify the first-order asymptotics of RH(n,p,δ)R_H(n,p,\delta). A major structural reduction expresses the problem through an entropy minimization over weighted graphs. Writing GGn,pG\sim\mathcal G_{n,p}0 for the family of symmetric GGn,pG\sim\mathcal G_{n,p}1 matrices GGn,pG\sim\mathcal G_{n,p}2 with GGn,pG\sim\mathcal G_{n,p}3 and GGn,pG\sim\mathcal G_{n,p}4, one defines the relative-entropy functional GGn,pG\sim\mathcal G_{n,p}5 and the weighted homomorphism density GGn,pG\sim\mathcal G_{n,p}6, and then considers

GGn,pG\sim\mathcal G_{n,p}7

For GGn,pG\sim\mathcal G_{n,p}8, the nonlinear large-deviation principle reduces the upper tail to this discrete variational problem: GGn,pG\sim\mathcal G_{n,p}9 Thus, solving Pr(XH(1+δ)E[XH]),\Pr\bigl(X_H\ge (1+\delta)\,\mathbb E[X_H]\bigr),0 to leading order yields Pr(XH(1+δ)E[XH]),\Pr\bigl(X_H\ge (1+\delta)\,\mathbb E[X_H]\bigr),1 (Bhattacharya et al., 2015).

This variational perspective has two consequences. First, it shifts the problem from direct probability estimates to an optimization problem with entropic cost Pr(XH(1+δ)E[XH]),\Pr\bigl(X_H\ge (1+\delta)\,\mathbb E[X_H]\bigr),2. Second, it makes the upper tail problem intrinsically structural: the optimizer describes the weighted graph most likely to realize the excess count. In later developments, analogous formulations reappear for directed graphs, hypergraphs, constrained sparse graph models, and spectral statistics (Park, 2024, Liu et al., 2019, Bhattacharya et al., 2019, Bhattacharya et al., 2018).

2. Dominant mechanisms: clique, anti-clique, hub, and core

In the sparse graph case, the leading constant in the exponent is often governed by the independence polynomial. If Pr(XH(1+δ)E[XH]),\Pr\bigl(X_H\ge (1+\delta)\,\mathbb E[X_H]\bigr),3 denotes the induced subgraph of Pr(XH(1+δ)E[XH]),\Pr\bigl(X_H\ge (1+\delta)\,\mathbb E[X_H]\bigr),4 on the vertices of degree Pr(XH(1+δ)E[XH]),\Pr\bigl(X_H\ge (1+\delta)\,\mathbb E[X_H]\bigr),5, and

Pr(XH(1+δ)E[XH]),\Pr\bigl(X_H\ge (1+\delta)\,\mathbb E[X_H]\bigr),6

then the sparse-regime solution identifies a parameter Pr(XH(1+δ)E[XH]),\Pr\bigl(X_H\ge (1+\delta)\,\mathbb E[X_H]\bigr),7 through

Pr(XH(1+δ)E[XH]),\Pr\bigl(X_H\ge (1+\delta)\,\mathbb E[X_H]\bigr),8

If Pr(XH(1+δ)E[XH]),\Pr\bigl(X_H\ge (1+\delta)\,\mathbb E[X_H]\bigr),9 is irregular, then

XHX_H0

If XHX_H1 is XHX_H2-regular on XHX_H3 vertices, then two competing constructions coexist: an anti-clique of cost XHX_H4 and a planted clique of cost XHX_H5, giving

XHX_H6

For triangles,

XHX_H7

with a clique-versus-anti-clique transition at XHX_H8 (Bhattacharya et al., 2015).

A complementary framework is the high-moments and entropic-stability method. For a polynomial XHX_H9 with nonnegative coefficients on the HH0-biased discrete hypercube, one defines a rate function HH1 by minimizing the cost of conditioning a small coordinate set HH2 to be present while forcing

HH3

If the family of such cores is entropically stable and HH4, then

HH5

and, with high probability conditioned on the upper-tail event, one of the near-optimal cores appears in the sample (Harel et al., 2019).

These two viewpoints are closely aligned. The variational formulation identifies the entropic minimizer at the level of weighted graphs, while the entropic-stability framework identifies the discrete witnesses that realize the same rare event. This suggests that the upper tail problem is fundamentally about classifying optimal witnesses, not merely estimating probabilities.

3. Canonical sparse-regime behaviors and solved examples

Several solved models exhibit a trichotomy or two-term competition between diffuse and localized strategies. For HH6-stars HH7, letting

HH8

and

HH9

one has

GG0

where

GG1

and

GG2

The same work interprets this as a “Poisson GG3 localization” transition and gives a first complete sparse solution for an irregular graph (Akhmejanova et al., 6 Jan 2025).

For GG4-term arithmetic progressions in a GG5-biased subset of GG6, the GG7-plane is partitioned into three disjoint regimes. In the Gaussian regime,

GG8

In the Poisson regime,

GG9

where logPr()-\log \Pr(\cdot)0. In the localised regime,

logPr()-\log \Pr(\cdot)1

and whenever logPr()-\log \Pr(\cdot)2,

logPr()-\log \Pr(\cdot)3

The proof combines exponential tilting, martingale concentration, and an entropic-stability lemma for small dense seeds (Harel et al., 2024).

Induced subgraph counts introduce additional phenomena. For induced logPr()-\log \Pr(\cdot)4, the sparse theory has a complete first-order asymptotic, but in the range

logPr()-\log \Pr(\cdot)5

the upper tail does not admit a naive mean-field approximation. In that regime, no single product-tilt can match the true lower bound, and the entropy minimizer ceases to be a two-point tilt in logPr()-\log \Pr(\cdot)6; instead one must condition on the appearance of “one of many” moderately sized bipartite subgraphs (Antonir, 2022).

For ordinary cycle counts, the classical logPr()-\log \Pr(\cdot)7-gap can be closed up to constants. If logPr()-\log \Pr(\cdot)8 is the logPr()-\log \Pr(\cdot)9-cycle and

HH0

then

HH1

Thus the exponent is of order HH2, eliminating the extra HH3 factor present in general Janson–Oleszkiewicz–Ruciński bounds (Raz, 2019).

4. The constant-average-degree triangle problem

A particularly sharp form of the upper tail problem arises when HH4 with fixed HH5. Let HH6 be the number of triangles in HH7. Then HH8 weakly converges to the Poisson distribution with mean HH9, and for fixed RH(n,p,δ):=logPr(XH(1+δ)E[XH]).R_H(n,p,\delta):=-\log\Pr\bigl(X_H\ge (1+\delta)\,\mathbb E[X_H]\bigr).0,

RH(n,p,δ):=logPr(XH(1+δ)E[XH]).R_H(n,p,\delta):=-\log\Pr\bigl(X_H\ge (1+\delta)\,\mathbb E[X_H]\bigr).1

The upper-tail problem asks how fast RH(n,p,δ):=logPr(XH(1+δ)E[XH]).R_H(n,p,\delta):=-\log\Pr\bigl(X_H\ge (1+\delta)\,\mathbb E[X_H]\bigr).2 may grow before the Poisson-tail approximation breaks down, and what replaces it beyond that threshold (Ganguly et al., 2022).

The transition is sharp. With

RH(n,p,δ):=logPr(XH(1+δ)E[XH]).R_H(n,p,\delta):=-\log\Pr\bigl(X_H\ge (1+\delta)\,\mathbb E[X_H]\bigr).3

the change of regime occurs when

RH(n,p,δ):=logPr(XH(1+δ)E[XH]).R_H(n,p,\delta):=-\log\Pr\bigl(X_H\ge (1+\delta)\,\mathbb E[X_H]\bigr).4

In the subcritical regime,

RH(n,p,δ):=logPr(XH(1+δ)E[XH]).R_H(n,p,\delta):=-\log\Pr\bigl(X_H\ge (1+\delta)\,\mathbb E[X_H]\bigr).5

with RH(n,p,δ):=logPr(XH(1+δ)E[XH]).R_H(n,p,\delta):=-\log\Pr\bigl(X_H\ge (1+\delta)\,\mathbb E[X_H]\bigr).6 and RH(n,p,δ):=logPr(XH(1+δ)E[XH]).R_H(n,p,\delta):=-\log\Pr\bigl(X_H\ge (1+\delta)\,\mathbb E[X_H]\bigr).7,

RH(n,p,δ):=logPr(XH(1+δ)E[XH]).R_H(n,p,\delta):=-\log\Pr\bigl(X_H\ge (1+\delta)\,\mathbb E[X_H]\bigr).8

Thus, to first order, the Poisson-tail size remains correct. In the supercritical regime,

RH(n,p,δ):=logPr(XH(1+δ)E[XH]).R_H(n,p,\delta):=-\log\Pr\bigl(X_H\ge (1+\delta)\,\mathbb E[X_H]\bigr).9

one has

RH(n,p,δ)R_H(n,p,\delta)0

and the dominant strategy becomes planting an almost-clique of size RH(n,p,δ)R_H(n,p,\delta)1 (Ganguly et al., 2022).

The conditioned structure is equally explicit. For fixed RH(n,p,δ)R_H(n,p,\delta)2, let

RH(n,p,δ)R_H(n,p,\delta)3

and

RH(n,p,δ)R_H(n,p,\delta)4

Whenever RH(n,p,δ)R_H(n,p,\delta)5,

RH(n,p,δ)R_H(n,p,\delta)6

Moreover,

RH(n,p,δ)R_H(n,p,\delta)7

in the subcritical regime, while

RH(n,p,δ)R_H(n,p,\delta)8

in the supercritical regime. The rare event is therefore realized either by almost RH(n,p,δ)R_H(n,p,\delta)9 vertex-disjoint triangles or by a medium-sized clique-like structure (Ganguly et al., 2022).

The proof combines several graph-theoretic ingredients. A connected triangle-induced graph spanned by GGn,pG\sim\mathcal G_{n,p}00 triangles has tree-excess GGn,pG\sim\mathcal G_{n,p}01, where GGn,pG\sim\mathcal G_{n,p}02 is concave and satisfies GGn,pG\sim\mathcal G_{n,p}03. If GGn,pG\sim\mathcal G_{n,p}04 is the event that there is a triangle-induced subgraph with GGn,pG\sim\mathcal G_{n,p}05 triangles, then

GGn,pG\sim\mathcal G_{n,p}06

for GGn,pG\sim\mathcal G_{n,p}07. The event GGn,pG\sim\mathcal G_{n,p}08 can be certified by disjoint triangle-induced subgraphs, so the van den Berg–Kesten inequality reduces the probability estimate to an optimization over partitions of GGn,pG\sim\mathcal G_{n,p}09. A stability version of the Kruskal–Katona bound then explains why the supercritical optimizer is almost-clique-like. The result settles the constant-average-degree triangle upper tail problem and answers a question of Aldous (Ganguly et al., 2022).

5. Spectral, directed, hypergraph, and constrained variants

The upper tail problem extends beyond copy counts. For the largest eigenvalue GGn,pG\sim\mathcal G_{n,p}10 of the adjacency matrix of GGn,pG\sim\mathcal G_{n,p}11, if GGn,pG\sim\mathcal G_{n,p}12 and GGn,pG\sim\mathcal G_{n,p}13, then

GGn,pG\sim\mathcal G_{n,p}14

In the same regime,

GGn,pG\sim\mathcal G_{n,p}15

The exponent for GGn,pG\sim\mathcal G_{n,p}16 again encodes a clique-regime and an anti-clique-regime. In a more localized sparse regime, when

GGn,pG\sim\mathcal G_{n,p}17

the joint upper tail of the top GGn,pG\sim\mathcal G_{n,p}18 eigenvalues becomes polynomial in GGn,pG\sim\mathcal G_{n,p}19: GGn,pG\sim\mathcal G_{n,p}20 and, conditioned on the event, the graph admits a decomposition into a disjoint union of stars and a spectrally negligible remainder (Bhattacharya et al., 2018, Bhattacharya et al., 2020).

For directed random graphs, the same general strategy leads to a directed variational problem over weighted digraphs. In the sparse regime GGn,pG\sim\mathcal G_{n,p}21, the upper-tail rate is asymptotically given by the corresponding entropy minimization. One obtains upper and lower bounds that differ by a constant factor of at most GGn,pG\sim\mathcal G_{n,p}22, and these coincide for several classes of digraphs, including triangles, stars, directed GGn,pG\sim\mathcal G_{n,p}23-cycles, and balanced digraphs (Park, 2024).

For GGn,pG\sim\mathcal G_{n,p}24-uniform hypergraphs, the sparse theory is more intricate than in graphs. A conjectural rate function based on compatible collections of mixed hubs and stable labelings was formulated for GGn,pG\sim\mathcal G_{n,p}25. This framework predicts that the optimizer may require genuinely mixed-hub structures rather than only cliques or 1-hubs, and a GGn,pG\sim\mathcal G_{n,p}26-uniform counterexample shows that naïve hub-only conjectures fail. Subsequent work confirmed the conjecture for complete GGn,pG\sim\mathcal G_{n,p}27-partite GGn,pG\sim\mathcal G_{n,p}28-graphs, tight cycles, and the Fano plane, and also established a general lower bound under explicit edge-covering conditions (Liu et al., 2019, Cook et al., 30 Sep 2025).

The same upper-tail logic persists under structural constraints on the ambient random graph. For sparse GGn,pG\sim\mathcal G_{n,p}29-regular graphs, sparse uniform GGn,pG\sim\mathcal G_{n,p}30-edge graphs, and inhomogeneous block models, the logarithmic decay rate is again described by an entropy minimization under homomorphism-count constraints, and joint upper tails for multiple graphs admit analogous variational formulas (Bhattacharya et al., 2019).

6. Methods, counterexamples, and open directions

A large part of the subject is method-driven. One important development is a BK-inequality based combinatorial sparsification method that recovers the missing GGn,pG\sim\mathcal G_{n,p}31 factor in settings where classical Kim–Vu and Janson–Ruciński arguments only yielded exponents of order GGn,pG\sim\mathcal G_{n,p}32. In a prototype weighted hypergraph theorem, the method gives

GGn,pG\sim\mathcal G_{n,p}33

thereby matching the expected logarithmic correction in many sparse problems. The mechanism is deletion-based sparsification combined with the van den Berg–Kesten inequality and refined Chernoff bounds with minor dependencies (Warnke, 2016).

At the same time, the subject contains sharp negative results against overly simple universality claims. The DeMarco–Kahn upper tail conjecture predicted that two mechanisms—clustered planting and disjoint planting—should always determine the exponent. This is false. An infinite family of counterexamples is given by graphs GGn,pG\sim\mathcal G_{n,p}34 obtained from an GGn,pG\sim\mathcal G_{n,p}35-cycle by attaching GGn,pG\sim\mathcal G_{n,p}36 pendant edges at a single cycle-vertex. For these graphs, in a range

GGn,pG\sim\mathcal G_{n,p}37

one has

GGn,pG\sim\mathcal G_{n,p}38

revealing a third, locally-disjoint mechanism beyond the conjectured clustered and disjoint ones (Šileikis et al., 2018).

Current open directions are highly model-specific. For constant-average-degree triangles, proposed extensions include other fixed subgraphs, inhomogeneous random graphs with bounded average degree, and dynamic random graph processes (Ganguly et al., 2022). For spectral tails, open problems include the case GGn,pG\sim\mathcal G_{n,p}39 for GGn,pG\sim\mathcal G_{n,p}40, a joint large-deviation principle for GGn,pG\sim\mathcal G_{n,p}41, higher edge-eigenvalues, and lower tails in the sparse regime (Bhattacharya et al., 2018). For arithmetic progressions, the boundary cases GGn,pG\sim\mathcal G_{n,p}42, GGn,pG\sim\mathcal G_{n,p}43, and GGn,pG\sim\mathcal G_{n,p}44 remain delicate (Harel et al., 2024). For hypergraphs, the full mixed-hub variational theory is still incomplete outside the presently verified classes (Liu et al., 2019, Cook et al., 30 Sep 2025).

A plausible implication is that the modern upper tail problem is best understood not as a single theorem, but as a classification program for rare-event geometry. The probability exponent, the optimizer of the entropy functional, and the conditioned structure are increasingly treated as inseparable parts of the same object.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Upper Tail Problem.