Papers
Topics
Authors
Recent
Search
2000 character limit reached

Clauset–Shalizi–Newman Protocol (Binned Data)

Updated 4 July 2026
  • The protocol is a statistically principled procedure for fitting power-law tails in binned data by integrating estimation, cutoff selection, and model testing.
  • It employs maximum-likelihood estimation, a KS goodness-of-fit test, and likelihood-ratio tests to determine the best heavy-tailed model.
  • The method automatically selects the tail onset by minimizing the KS distance, balancing bias and variance in binned observations.

Searching arXiv for the original Clauset–Shalizi–Newman framework and the binned-data adaptation. The Clauset–Shalizi–Newman protocol, in the form adapted to binned observations, is a statistically principled procedure for fitting and testing power-law tails when empirical data are available only as counts in known bins. In this formulation, the protocol combines maximum-likelihood estimation of the scaling exponent, data-driven selection of the lower cutoff, a goodness-of-fit test based on the Kolmogorov–Smirnov statistic, and likelihood-ratio comparisons against alternative heavy-tailed models. The adaptation is designed for settings in which fluctuations in the empirical tail are already substantial and are further worsened when information is lost from binning; its purpose is to recover, as far as possible, the inferential structure of the Clauset–Shalizi–Newman framework under this constraint (Virkar et al., 2012).

1. Problem setting and inferential objective

The protocol addresses binned data of the form {(bi,bi+1),hi}\{(b_i,b_{i+1}),h_i\}, where B=(b1,b2,,bk)B=(b_1,b_2,\dots,b_k) are bin boundaries and hih_i is the count observed in bin ii. The central inferential question is whether the upper tail of such data is plausibly consistent with a power-law distribution, and, if so, over which range and with what exponent (Virkar et al., 2012).

The adaptation is motivated by the fact that accurate identification of power-law patterns has significant consequences for correctly understanding and modeling complex systems, while statistical evidence for or against the power-law hypothesis is complicated by large fluctuations in the empirical distribution's tail. In the binned setting, this difficulty is compounded by the loss of information induced by coarsening. The protocol therefore does not treat estimation, goodness-of-fit testing, and model comparison as separable tasks; rather, it organizes them into a single workflow in which the fitted tail region, the exponent estimate, and the plausibility of the model are assessed jointly.

A defining feature of the protocol is that the lower boundary of the power-law regime is not assumed known. Instead, the retained tail begins at a bin boundary bminb_{\min}, chosen from the observed bin structure. If the retained tail contains counts

n=i:bibminhi,n=\sum_{i:\,b_i\ge b_{\min}} h_i,

then all subsequent fitting and testing are performed on that tail alone. This structure formalizes the common situation in which only the upper part of an empirical distribution may exhibit power-law behavior, while the body does not.

2. Maximum-likelihood estimation of the exponent

For a fixed lower cutoff bmb_m, the protocol estimates the scaling exponent α>1\alpha>1 by maximum likelihood using only bins with bibmb_i\ge b_m. The log-likelihood is

L(α)=n(α1)lnbm+i:bibmhiln[bi1αbi+11α].\mathcal L(\alpha) =n(\alpha-1)\ln b_{m} +\sum_{i:\,b_i\ge b_{m}} h_i\,\ln\Bigl[b_i^{\,1-\alpha}-b_{i+1}^{\,1-\alpha}\Bigr]\,.

The maximum-likelihood estimate B=(b1,b2,,bk)B=(b_1,b_2,\dots,b_k)0 is the value of B=(b1,b2,,bk)B=(b_1,b_2,\dots,b_k)1 that maximizes this expression. In general, the optimization must be solved numerically (Virkar et al., 2012).

In the special case of logarithmic binning,

B=(b1,b2,,bk)B=(b_1,b_2,\dots,b_k)2

the estimator has a closed form: B=(b1,b2,,bk)B=(b_1,b_2,\dots,b_k)3 For the same logarithmic-binning case, an approximate standard error is

B=(b1,b2,,bk)B=(b_1,b_2,\dots,b_k)4

These formulas provide the parametric core of the protocol. They convert the binned counts directly into a tail-exponent estimate without reconstructing unbinned observations. A plausible implication is that the procedure preserves the inferential target of the unbinned Clauset–Shalizi–Newman framework while explicitly accounting for the discretization imposed by bins.

3. Selection of the lower cutoff

The lower cutoff is treated as an unknown to be selected from among the bin boundaries. For each candidate B=(b1,b2,,bk)B=(b_1,b_2,\dots,b_k)5, the protocol performs two operations: it fits B=(b1,b2,,bk)B=(b_1,b_2,\dots,b_k)6 by maximum likelihood using only bins with B=(b1,b2,,bk)B=(b_1,b_2,\dots,b_k)7, and it computes the Kolmogorov–Smirnov distance B=(b1,b2,,bk)B=(b_1,b_2,\dots,b_k)8 between the empirical counts in those bins and the cdf of the fitted power law. The selected cutoff is

B=(b1,b2,,bk)B=(b_1,b_2,\dots,b_k)9

This criterion is described as automatically balancing bias against variance: if hih_i0 is too small, non–power-law data distort the fit; if hih_i1 is too large, too little data remain for reliable estimation (Virkar et al., 2012).

This cutoff-selection step is central to the protocol’s logic. It operationalizes the idea that a power law, if present, may govern only a tail subset of the observations. The procedure therefore avoids imposing a global power-law model on the entire binned sample and instead estimates the onset of scaling behavior from the data themselves.

A common misunderstanding is to treat hih_i2 as a cosmetic reporting choice. Within this protocol it is not cosmetic: it determines the sample size hih_i3 used for estimation, the cdf against which goodness of fit is measured, and the set of observations on which model-comparison tests are conducted.

4. Kolmogorov–Smirnov goodness-of-fit test

Once hih_i4 and hih_i5 have been chosen, the model cdf in the tail is defined by

hih_i6

and hih_i7 denotes the corresponding empirical cdf, defined as normalized cumulative counts in the same bins. The Kolmogorov–Smirnov distance is then

hih_i8

The associated hih_i9-value is obtained by a semi-parametric bootstrap. Synthetic data sets of size ii0 are generated by drawing with probability ii1 from the fitted power-law tail and with probability ii2 from the empirical body below ii3; each simulated data set is then re-binned, re-fitted, and assigned its own KS distance ii4. The ii5-value is the fraction of simulated ii6 values that exceed the observed ii7. If ii8, the power-law hypothesis is rejected as an implausible fit (Virkar et al., 2012).

The inferential role of this test is narrower than simple visual agreement. The protocol does not regard a fitted exponent alone as evidence for a power law. Instead, the power-law hypothesis must survive an explicit goodness-of-fit test after accounting for the same fitting and cutoff-selection procedure used on the observed data. This suggests that apparent linearity in log-log summaries is insufficient within the protocol’s evidentiary standard.

5. Likelihood-ratio comparison with alternatives

If the power law is not rejected, the protocol compares it against alternative explanations using likelihood-ratio tests on the same tail data, namely bins with ii9. For a power-law model bminb_{\min}0 and an alternative bminb_{\min}1, such as exponential, log-normal, stretched exponential, or power-law with cutoff, the test statistic is

bminb_{\min}2

A positive bminb_{\min}3 favors the power law, whereas a negative value favors the alternative (Virkar et al., 2012).

To assess significance, the protocol uses Vuong’s test. One computes the per-bin log-likelihood differences

bminb_{\min}4

their mean bminb_{\min}5, and their standard deviation bminb_{\min}6 over the bminb_{\min}7 tail observations, and then forms

bminb_{\min}8

Under the null of equal model performance, bminb_{\min}9 is asymptotically standard normal, and the two-tailed n=i:bibminhi,n=\sum_{i:\,b_i\ge b_{\min}} h_i,0-value is

n=i:bibminhi,n=\sum_{i:\,b_i\ge b_{\min}} h_i,1

If n=i:bibminhi,n=\sum_{i:\,b_i\ge b_{\min}} h_i,2, the sign of n=i:bibminhi,n=\sum_{i:\,b_i\ge b_{\min}} h_i,3, equivalently of n=i:bibminhi,n=\sum_{i:\,b_i\ge b_{\min}} h_i,4, is taken as a reliable indication of the better model; otherwise no decision is possible. In the nested case of power law versus power law with cutoff, a n=i:bibminhi,n=\sum_{i:\,b_i\ge b_{\min}} h_i,5 test on n=i:bibminhi,n=\sum_{i:\,b_i\ge b_{\min}} h_i,6 is used instead.

This component of the protocol is designed to prevent a false dichotomy between “not rejected” and “established.” A non-rejected power law may still be outperformed by another heavy-tailed model on the same data. Conversely, if the likelihood-ratio test is not significant, the prescribed conclusion is not that the models are equivalent in a substantive sense, but that the comparison is inconclusive.

6. Effect of binning and algorithmic workflow

The protocol explicitly states that any coarsening of bins reduces statistical power. For logarithmic binning by ratio n=i:bibminhi,n=\sum_{i:\,b_i\ge b_{\min}} h_i,7, an asymptotic relation is given for the sample size needed to achieve the same variance in n=i:bibminhi,n=\sum_{i:\,b_i\ge b_{\min}} h_i,8 under two different coarseness levels: n=i:bibminhi,n=\sum_{i:\,b_i\ge b_{\min}} h_i,9 As bmb_m0 increases, or as bmb_m1 increases, many bins in the tail become empty, and both the exponent estimate and the goodness-of-fit and model-comparison tests lose power. The stated recommendation is therefore to choose the finest bins feasible in data collection or reporting to maximize statistical accuracy (Virkar et al., 2012).

The complete workflow is given as a seven-step algorithm. For each candidate tail boundary bmb_m2 with at least two nonempty bins above it, one fits bmb_m3 by maximizing the binned log-likelihood and computes the KS distance bmb_m4. One then selects bmb_m5, adopts the corresponding bmb_m6, computes the observed KS statistic bmb_m7, and generates bmb_m8 synthetic data sets by semi-parametric bootstrap. The bootstrap procedure re-draws from the fitted tail and the empirical body, re-bins the data, re-selects bmb_m9, re-fits α>1\alpha>10, and recomputes α>1\alpha>11. The resulting bootstrap α>1\alpha>12-value is

α>1\alpha>13

If α>1\alpha>14, the power law is rejected; otherwise alternative models are fit by MLE on the same tail, compared by α>1\alpha>15, and adjudicated by the Vuong statistic or reported as inconclusive when significance is absent.

This workflow integrates estimation uncertainty, threshold selection, and model comparison into a single procedure. A plausible implication is that the protocol is best understood not as a formula for estimating α>1\alpha>16, but as a full inferential pipeline for binned heavy-tailed data.

7. Interpretation, scope, and common points of confusion

The protocol is an adaptation of the Clauset–Shalizi–Newman framework to the binned context, not a separate theory of heavy tails. Its contribution lies in preserving the framework’s statistically principled structure under loss of resolution: maximum-likelihood fitting, Kolmogorov–Smirnov goodness-of-fit testing, and likelihood-ratio comparisons against alternative explanations are all retained, but reformulated in terms of counts over bins rather than individual observations (Virkar et al., 2012).

Several recurrent confusions are clarified by the protocol itself. First, a fitted power-law exponent does not by itself validate the power-law hypothesis; the goodness-of-fit test can still reject the model. Second, a non-rejected power law is not automatically preferred to alternatives; likelihood-ratio tests may favor another model or remain inconclusive. Third, the effects of binning are not merely computational; they directly reduce statistical power, especially when the bins are coarse or the exponent is large. Fourth, the lower cutoff is not fixed in advance by assumption but chosen by minimizing the KS distance over candidate bin boundaries.

The protocol’s scope is specifically the tail of binned empirical data with heavy-tailed patterns. It yields three outputs: a statistically principled fit of a power-law tail to binned data, a goodness-of-fit α>1\alpha>17-value, and model-comparison tests against alternative heavy-tailed distributions. Within that scope, its significance lies in replacing informal visual diagnosis with an explicit inferential procedure adapted directly from the Clauset–Shalizi–Newman framework to the binned setting.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Clauset-Shalizi-Newman Protocol.