Papers
Topics
Authors
Recent
Search
2000 character limit reached

Likelihood Ratio Scoring via Block Rejection

Updated 21 December 2025
  • The paper introduces ε-regularization to stabilize the likelihood ratio statistic in sparse stochastic block models, ensuring valid hypothesis testing.
  • It applies a Monte Carlo approximation to efficiently manage the sum over 2^n labelings, yielding asymptotic power-Poisson distributions with rigorous error control.
  • The method effectively distinguishes Erdős–Rényi graphs from balanced community structures, outperforming traditional spectral and subgraph-count tests in high-SNR regimes.

Likelihood ratio scoring with block rejection refers to a rigorous statistical methodology for hypothesis testing in network data, specifically for distinguishing between an Erdős–Rényi (ER) random graph and a balanced two-community stochastic block model (SBM) in the bounded degree regime. The standard likelihood ratio (LR) approach degenerates in sparse settings due to the asymptotic orthogonality of the probability measures when the signal-to-noise ratio exceeds a certain threshold. To address this, an ε\varepsilon–regularization is introduced to stabilize the LR statistic, allowing for valid inference through block rejection rules. The resulting test yields asymptotic distributions characterized by power-Poisson laws, and achieves robust performance via Monte Carlo approximation, with strong theoretical and empirical guarantees in the high-SNR regime (Yuan et al., 2018).

1. Problem Formulation and Model Definitions

Consider the testing problem where an observed undirected graph GG on nn vertices is to be distinguished between:

  • H0H_0: GG(n,p0)G \sim G(n, p_0) (Erdős–Rényi, with p0=(a+b)/(2n)p_0=(a+b)/(2n))
  • H1H_1: GG(n,a,b)G \sim G(n, a, b) (balanced bisection SBM), where each vertex is labeled ou{±1}o_u \in \{\pm 1\} independently, and

Pr(Auv=1o)={a/n,ou=ov b/n,ouova>b>0\Pr(A_{uv}=1 \mid o) = \begin{cases} a/n, & o_u = o_v \ b/n, & o_u \neq o_v \end{cases} \quad a > b > 0

The “signal‐to‐noise ratio” (SNR) is defined as GG0. When GG1 and GG2 are fixed (bounded degree regime), the classical LR statistic fails due to the lack of contiguity between the distributions. The problem forms the foundation for hypothesis testing in community detection and the determination of the number of communities.

2. Regularized Likelihood Ratio Statistic

An GG3–regularization is introduced, modifying the intra-group and inter-group edge probabilities:

GG4

The regularized likelihood under the SBM model for a label assignment GG5 is

GG6

where GG7 if GG8, GG9 if nn0. The ER log-likelihood is nn1, and the nn2–regularized LR statistic is the average over all labelings:

nn3

Regularization ensures the ratio nn4 is strictly less than one, suppressing the explosive variance observed with the standard LR in the high-SNR bounded-degree setting.

3. Asymptotic Power-Poisson Laws

As nn5 and nn6, the asymptotic distributions of nn7 are infinite “power-Poisson” products. For each cycle length nn8:

nn9

  • Under H0H_00:

H0H_01

  • Under H0H_02 (with block-signal parameter H0H_03):

H0H_04

These results rely on a Janson-type contiguity criterion and sufficient control over mixed moments of cycle counts and H0H_05.

4. Rejection Rule and Error Rates

Given a desired significance level H0H_06, compute the H0H_07-quantile H0H_08 of H0H_09 under GG(n,p0)G \sim G(n, p_0)0 (GG(n,p0)G \sim G(n, p_0)1). The block rejection rule:

  • Reject GG(n,p0)G \sim G(n, p_0)2 if GG(n,p0)G \sim G(n, p_0)3 achieves

GG(n,p0)G \sim G(n, p_0)4

GG(n,p0)G \sim G(n, p_0)5

Power analysis reveals that for suitable regularization parameters and growing average degree, the test is asymptotically powerful whenever GG(n,p0)G \sim G(n, p_0)6.

5. Monte Carlo Approximation and Computational Considerations

Direct evaluation of GG(n,p0)G \sim G(n, p_0)7 is infeasible since it requires summing over GG(n,p0)G \sim G(n, p_0)8 labelings. A Monte Carlo (MC) estimator using GG(n,p0)G \sim G(n, p_0)9 i.i.d. labelings achieves

p0=(a+b)/(2n)p_0=(a+b)/(2n)0

with computational cost p0=(a+b)/(2n)p_0=(a+b)/(2n)1, where p0=(a+b)/(2n)p_0=(a+b)/(2n)2 and p0=(a+b)/(2n)p_0=(a+b)/(2n)3 is required for negligible MC error. This allows practical application for moderate system sizes, underlining the method's potential in sparse network regimes.

6. Empirical Studies and Performance Benchmarking

Simulations for values p0=(a+b)/(2n)p_0=(a+b)/(2n)4, p0=(a+b)/(2n)p_0=(a+b)/(2n)5, and p0=(a+b)/(2n)p_0=(a+b)/(2n)6 correspond to SNR p0=(a+b)/(2n)p_0=(a+b)/(2n)7 with p0=(a+b)/(2n)p_0=(a+b)/(2n)8. With appropriately chosen p0=(a+b)/(2n)p_0=(a+b)/(2n)9, the empirical size at level H1H_10 matches the theoretical prediction, and power increases with H1H_11 and H1H_12. The regularized LR procedure outperforms the spectral test of Bickel–Sarkar and the subgraph-count test of Gao–Lafferty in these sparse regimes.

On real data, e.g., the “political books” co-purchase network (H1H_13), the test shows high success rates (83%–100%) in correctly rejecting the null when distinct communities are merged, again exceeding the performance of competing methods in sparse graphs.

7. Extensions, Limitations, and Future Directions

The H1H_14–regularization is critical for restoring statistical contiguity and mitigating the erratic behavior of the standard LR statistic in SBMs with bounded average degree. No test is possible for H1H_15 due to information-theoretic lower bounds. The MC approach's computational cost remains exponential in the worst case; further analysis of deterministic or mean-field approximations is needed. Extensions to multi-community SBMs, degree-corrected models, and connections with semidefinite and spectral relaxations are posited as promising avenues for future work. The likelihood ratio scoring with block rejection framework establishes a principled, theoretically robust paradigm for community hypothesis testing in sparse networks (Yuan et al., 2018).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Likelihood Ratio Scoring with Block Rejection.