Papers
Topics
Authors
Recent
Search
2000 character limit reached

Planted Anomalous Community

Updated 11 July 2026
  • Planted anomalous community is a hidden subset of vertices exhibiting systematic deviation from a random graph model, defined by elevated within-community edge probabilities.
  • The model leverages parameters such as sparsity (α) and community size (β) to determine detectability, highlighting a sharp computational-statistical gap in sparse settings.
  • Efficient detection techniques include linear-time edge counting and exhaustive scan statistics, while recovery methods use probabilistic, spectral, and Bayesian approaches for community localization.

A planted anomalous community is a hidden subset of vertices whose induced interactions deviate systematically from a background random graph or hypergraph model. In the canonical sparse-graph formulation, one observes GG(N,q)G \sim \mathcal{G}(N,q) under the null and GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q) under the alternative, where a hidden subset S[N]S \subset [N] of mean size KK has within-community edge probability p=qin=cqp=q_{\mathrm{in}}=cq with c>1c>1, while all other edges occur with probability q=Nαq=N^{-\alpha}; with K=Θ(Nβ)K=\Theta(N^\beta), the problem exhibits a sharp computational-statistical theory governed by α\alpha and β\beta (Hajek et al., 2014). Closely related formulations replace Bernoulli edges by Gaussian weights, hyperedges, directed comparisons, or circular phases, but retain the same organizing idea: a small latent subset induces a coherent deviation relative to a null ensemble (Corinzia et al., 2020, Kunisky et al., 2024, Ameen et al., 9 Jan 2026).

1. Formal definition and canonical model

In the Erdős–Rényi background model, GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)0 has independent edges, each present with probability GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)1. The planted anomalous community model introduces a hidden subset GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)2, GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)3, such that edges within GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)4 appear with elevated probability GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)5, while all other edges remain at probability GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)6. The null and alternative are

GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)7

and the paper’s main regime is GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)8, GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)9, with S[N]S \subset [N]0 fixed, S[N]S \subset [N]1, and S[N]S \subset [N]2. The planted subgraph size is random in the main formulation: each vertex is included in S[N]S \subset [N]3 independently with probability S[N]S \subset [N]4, so S[N]S \subset [N]5 with mean S[N]S \subset [N]6; this random-size model is statistically equivalent, up to constants, to the fixed-size model in the regimes considered. The basic signal quantities are

S[N]S \subset [N]7

Increasing S[N]S \subset [N]8 or decreasing S[N]S \subset [N]9 makes detection harder (Hajek et al., 2014).

Within this literature, “planted anomalous community” and “planted dense subgraph” are effectively synonymous in the graph case. The same structural role appears in the weighted KK0-uniform planted KK1-densest sub-hypergraph model, where the planted set induces a coherent mean shift KK2 on the KK3 hyperedges entirely within the community, and the maximum-likelihood estimator is the KK4-densest sub-hypergraph objective KK5 (Corinzia et al., 2020). In directed graphs, the anomaly can instead be a ranked community: the planted subset is not distinguished by higher edge density, but by unusual consistency of pairwise orderings (Kunisky et al., 2024). In circular-data models, the planted community is defined by phase coherence on the internal edges of a complete graph, with signal edges drawn either from a common arc of length KK6 or from a von Mises distribution with common location parameter KK7 (Ameen et al., 9 Jan 2026).

2. Statistical detectability in sparse Erdős–Rényi graphs

For the Bernoulli planted dense subgraph model with KK8, the sharp information-theoretic detection threshold is

KK9

If p=qin=cqp=q_{\mathrm{in}}=cq0, no test of any computational complexity can reliably distinguish p=qin=cqp=q_{\mathrm{in}}=cq1 from p=qin=cqp=q_{\mathrm{in}}=cq2; the total variation distance tends to zero. If p=qin=cqp=q_{\mathrm{in}}=cq3, there exists a generally computationally intensive test achieving vanishing error probability. A non-asymptotic impossibility statement is given by Proposition 1: if

p=qin=cqp=q_{\mathrm{in}}=cq4

then

p=qin=cqp=q_{\mathrm{in}}=cq5

for some function p=qin=cqp=q_{\mathrm{in}}=cq6 with p=qin=cqp=q_{\mathrm{in}}=cq7, so detection is impossible in that regime (Hajek et al., 2014).

The same result is often summarized through three regions in the p=qin=cqp=q_{\mathrm{in}}=cq8-plane. The “simple regime” is

p=qin=cqp=q_{\mathrm{in}}=cq9

where a linear-time test based on total edge count is statistically optimal. The “hard regime” is

c>1c>10

where detection is statistically possible via scan statistics or similar exhaustive procedures but is conjectured to be computationally intractable. The “impossible regime” is

c>1c>11

where detection is information-theoretically impossible. The critical sparsity

c>1c>12

marks the point where the information-theoretic boundary and the efficient frontier coincide: when c>1c>13, c>1c>14, and no computational-statistical gap remains; when c>1c>15, the statistical boundary is c>1c>16, which is strictly below the efficient boundary (Hajek et al., 2014).

This phase diagram is specific to the constant-factor anomaly c>1c>17, but the organizing principle is broader. A plausible implication is that sparsity does not simply make the problem uniformly harder; rather, it changes which statistic is optimal and whether exhaustive combinatorial structure is necessary.

3. Efficient procedures, scan methods, and computational hardness

Two explicit test statistics organize the graph case. The linear-time procedure uses the total number of edges,

c>1c>18

with threshold

c>1c>19

It rejects q=Nαq=N^{-\alpha}0 if q=Nαq=N^{-\alpha}1, runs in q=Nαq=N^{-\alpha}2, and succeeds when

q=Nαq=N^{-\alpha}3

The computationally intensive scan statistic is

q=Nαq=N^{-\alpha}4

with threshold

q=Nαq=N^{-\alpha}5

It rejects q=Nαq=N^{-\alpha}6 if q=Nαq=N^{-\alpha}7, has runtime q=Nαq=N^{-\alpha}8 in naive form, and succeeds when

q=Nαq=N^{-\alpha}9

or, ignoring logarithms,

K=Θ(Nβ)K=\Theta(N^\beta)0

Thus the scan achieves the statistical threshold up to log factors, while the linear statistic matches the best possible polynomial-time boundary in the simple regime (Hajek et al., 2014).

The computational lower bound is obtained through a randomized reduction from planted clique. Under the planted clique hypothesis—namely that planted clique detection is intractable in polynomial time for K=Θ(Nβ)K=\Theta(N^\beta)1 at any constant K=Θ(Nβ)K=\Theta(N^\beta)2—the paper constructs a randomized mapping from K=Θ(Nβ)K=\Theta(N^\beta)3 to K=Θ(Nβ)K=\Theta(N^\beta)4 with K=Θ(Nβ)K=\Theta(N^\beta)5, K=Θ(Nβ)K=\Theta(N^\beta)6, and K=Θ(Nβ)K=\Theta(N^\beta)7. The reduction preserves the null exactly and approximates the alternative in total variation, so any polynomial-time solver for the planted dense subgraph instance would transfer to a polynomial-time solver for planted clique. If the hypothesis holds for all K=Θ(Nβ)K=\Theta(N^\beta)8, the efficient boundary simplifies to

K=Θ(Nβ)K=\Theta(N^\beta)9

and no polynomial-time algorithm achieves reliable detection when

α\alpha0

This is the formal source of the computational-statistical gap for α\alpha1 (Hajek et al., 2014).

Later low-degree work on planted-vs-planted testing in planted dense subgraph models identifies a sharp strong-testing threshold at

α\alpha2

achieved by counting balanced unicyclic graphs, and shows that trees are uninformative while cyclic structures carry the signal. That threshold coincides, down to the sharp constant, with the known low-degree recovery threshold (Skeja et al., 3 Jun 2026). This suggests that the combinatorial role of cycle-like substructures remains central even when the hypotheses are both planted rather than planted-versus-null.

4. Recovery, densest subgraph, and exact localization

Detection asks only whether a planted anomaly is present; recovery asks for the community itself. In the sparse Bernoulli model α\alpha3, α\alpha4, α\alpha5, exact or near-exact recovery is known to be possible information-theoretically if and only if α\alpha6 and α\alpha7. Efficient recovery is known in the region

α\alpha8

via convex relaxations or spectral or iterative methods. The same framework also yields hardness for average-case approximation of densest α\alpha9-subgraph: under planted clique hardness, any constant-factor approximation is hard on average in the hard regime

β\beta0

whereas in the simple region β\beta1, recovering the planted community and thus obtaining a β\beta2-approximation is possible in polynomial time. The reduction extends to deterministic-size planted subgraphs for monotone tests and gives average-case computational hardness of recovery in the same hard regime (Hajek et al., 2014).

Weighted models sharpen the distinction between community recovery and anomalous-subgraph recovery. In the Gaussian weighted planted dense subgraph model with planted set of size β\beta3, exact recovery is impossible when the same β\beta4, even statistically, whereas the maximum-likelihood estimator succeeds when β\beta5, and the semidefinite relaxation succeeds down to the threshold value of β\beta6. By contrast, in the Gaussian weighted stochastic block model with two symmetric communities, exact recovery is possible, both statistically and algorithmically, down to β\beta7. The paper therefore shows that exact recovery of two symmetric communities is a strictly easier problem than recovering a planted dense subgraph of size half the total number of nodes (Pandey et al., 2024).

For linear-size planted dense subgraphs β\beta8, however, simple spectral procedures can be optimal. In the sparse Bernoulli planted dense subgraph model with

β\beta9

exact recovery is achieved by a linear-combination-of-eigenvectors spectral algorithm whenever

GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)00

matching the information-theoretic threshold. In submatrix localization with

GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)01

thresholding the top eigenvector achieves exact recovery whenever

GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)02

The same paper proves optimal exact recovery for a censored planted dense subgraph model via a signed-adjacency spectral algorithm under

GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)03

(Dhara et al., 2022).

5. Heterogeneous, semi-random, and unbalanced graph models

The homogeneous Erdős–Rényi background is mathematically clean but restrictive. In an inhomogeneous random graph GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)04, edges are independent conditional on a nonnegative weight vector GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)05, with

GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)06

The planted model GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)07 introduces a latent Bernoulli label vector GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)08 and within-community edge probability GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)09 against baseline GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)10. For testing

GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)11

the proposed polynomial-time statistic is based on triangle and 6-cycle densities,

GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)12

Under mild moment bounds on GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)13, GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)14 under GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)15. Under GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)16,

GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)17

and if GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)18, power tends to one exactly when

GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)19

This supplies a parameter-free polynomial-time test for planted communities in heterogeneous networks (Yuan et al., 2021).

A more general inhomogeneous framework allows arbitrary baseline probabilities GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)20. Under the null, GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)21; under the alternative, a hidden set GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)22 of size GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)23 has GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)24 for GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)25. The scan statistic

GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)26

is maximized over GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)27, GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)28. The information-theoretic lower bound and the upper bound of the scan test are both driven by

GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)29

showing that the “most informative subgraph” inside the planted community, rather than the full community, can determine detectability in inhomogeneous graphs (Bogerd et al., 2019).

Other graph models change the geometry of the anomaly rather than the noise law. In the semi-random planted sparse vertex cut model, the anomaly is a pair of well-connected groups GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)30 and GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)31 linked mainly by few connector or ambassador vertices GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)32 and GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)33; the relevant complexity measure is balanced vertex expansion, and semidefinite programming yields exact recovery or constant-factor bi-criteria approximation under spectral-gap and randomness conditions (Louis et al., 2018). In the planted partition model with arbitrarily many and highly unbalanced communities, Diamond Percolation retains an edge GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)34 when the number of common neighbors

GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)35

satisfies GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)36, then returns connected components of the retained graph. Under the size-sparsity assumption

GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)37

the method achieves exact, almost exact, or weak recovery, including power-law community-size regimes (Gösgens et al., 2 Apr 2025).

6. Higher-order, directed, and circular generalizations

In GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)38-uniform hypergraphs with Gaussian weights, the planted GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)39-densest sub-hypergraph model selects a planted community GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)40 of size GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)41 and observes

GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)42

with GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)43. The maximum-likelihood estimator is

GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)44

The normalized signal-to-noise ratio

GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)45

governs exact recovery. The paper provides upper and lower information-theoretic thresholds for exact recovery and an approximate message passing threshold

GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)46

exhibiting a statistical-computational gap that widens with sparsity (Corinzia et al., 2020). A complementary GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)47-uniform sub-hypergraph stochastic block model gives exact-recovery limits in terms of GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)48 times a divergence between within- and outside-hyperedge laws: exact recovery is impossible when

GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)49

and achievable by maximum likelihood when

GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)50

(Liang et al., 2021).

Directed formulations replace density by order consistency. In the ranked-community model, a hidden set GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)51 of size GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)52 carries a latent ranking GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)53, and observed pairwise orderings inside GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)54 agree with GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)55 with probability GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)56, while all other observed directions are uniform. In the log-density regime

GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)57

strong detection is statistically possible if

GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)58

and polynomial-time strong detection is achievable if

GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)59

Strong recovery is statistically possible if GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)60, while polynomial-time strong recovery is achievable if

GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)61

The anomaly is thus defined by unusual consistency of orientations rather than elevated edge density (Kunisky et al., 2024).

Circular-data models provide another non-density generalization. In the community setting, one observes phases GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)62 on the edges of a complete graph. Under the null, all GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)63 are i.i.d. uniform on GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)64. Under the alternative, there is a community GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)65 of size GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)66 and an unknown phase GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)67 such that internal edges follow either a hard-arc law GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)68 or a von Mises law

GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)69

For the hard-arc model, weak detection is impossible if

GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)70

For the von Mises model, weak detection is impossible if

GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)71

The main achievability tools are interval scans and the coherence statistic

GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)72

which is a phase-invariant analog of a scan over coherent edge orientations (Ameen et al., 9 Jan 2026). This suggests that the notion of “anomalous community” is best understood as a latent subset inducing a structured deviation—density, weight, rank consistency, or phase coherence—rather than as a purely topological dense block.

7. Operational anomaly scoring and uncertainty quantification

Some recent work adopts an explicitly operational rather than minimax definition. In co-membership-based generic anomalous communities detection, an anomalous community is a community whose member set contains many “unexpected” vertices when considered against the broader network’s co-membership structure. The method constructs a bipartite utility graph GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)73, trains an XGBoost link-prediction classifier on community-membership edges GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)74, and interprets the predicted membership probability GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)75 as the probability that vertex GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)76 belongs to community GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)77. Community anomaly scores are then aggregated as

GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)78

with analogous label-based scores. On the Reddit-based anomaly-infused dataset, the best meta-feature achieved GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)79; on the fully simulated dataset, GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)80 was obtained by GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)81 and GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)82. The methodology is domain-free and relies on co-membership rather than internal density (Lapid et al., 2022).

A probabilistic generative approach models anomalous edges directly. In the mixed-membership Poisson model, each edge has a binary latent anomaly indicator GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)83, with

GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)84

The posterior anomaly probability for an undirected graph is

GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)85

and a planted anomalous community is recovered when the posterior anomaly probabilities concentrate on a coherent edge set. The paper emphasizes that anomalies are defined relative to the learned community-based null model, not relative to a fixed density threshold (Safdari et al., 2022).

Bayesian uncertainty quantification for sparse community models clarifies how confident one can be in a detected anomaly. In the sparse planted bi-section model, when the posterior recovers the true community assignment exactly, any sequence of credible sets of levels bounded away from zero is also a consistent sequence of confidence sets. In the almost-exact regime, if

GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)86

then the GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)87-enlargements of credible sets achieve asymptotic frequentist coverage, and minimal-diameter credible sets satisfy

GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)88

with high probability. In regimes where GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)89 and GG(N,K,p,q)G \sim \mathcal{G}(N,K,p,q)90 are very close, enlarged credible sets can still deliver asymptotic coverage via remote contiguity arguments (Kleijn et al., 2018). A plausible implication is that, even when exact localization of a planted anomalous community is statistically or computationally delicate, neighborhood-based uncertainty sets can remain interpretable and valid.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Planted Anomalous Community.