Adjusted Wilson Confidence Interval
- Adjusted Wilson confidence intervals are a family of modified Wilson score intervals that improve on the classical Wald intervals for binomial proportions.
- They incorporate techniques such as rare-event relative margin adjustments, finite-population corrections, test inversion for exact coverage, and pseudo-observation smoothing.
- These adjustments enhance coverage accuracy, reduce interval width, and offer practical solutions for applications including privacy-preserving data analysis and selective classification.
The adjusted Wilson confidence interval is not a single canonical construction but a family of modifications, recalibrations, and domain-specific generalizations built around the Wilson score interval for a binomial proportion. Across the recent literature, “adjustment” may refer to changing the criterion by which Wilson intervals are judged in rare-event settings, replacing the nominal sample size by a finite-population effective sample size, transforming an approximate Wilson interval into an exact interval by test inversion, or adding pseudo-observations to obtain a smoothed Wilson-like procedure (McGrath et al., 2021, O'Neill, 2021, Wang, 2021, Kahouadji, 13 Aug 2025). The unifying theme is dissatisfaction with the classical Wald interval and the search for interval estimators that retain closed-form convenience while improving coverage, width, or coherence under nonstandard sampling regimes.
1. Wilson score interval as the reference construction
For a binomial sample with observed proportion and confidence level , the standard Wilson score interval is
It arises by inverting a score-test or normal-approximation inequality in which the null standard error is used rather than the plug-in standard error of the Wald interval (Dalitz, 2018).
The Wilson interval is the baseline against which adjusted variants are defined because it remedies several structural defects of Wald. It remains inside the parameter space , behaves sensibly at small , and has monotonicity and consistency properties that can be stated explicitly. In the formulation using , its width satisfies and ; the lower and upper bounds are increasing in the observed sample proportion ; and (O'Neill, 2021).
These properties explain why Wilson is repeatedly treated as the default closed-form alternative to Wald. In broad comparisons, Wald has too low coverage, especially near 0 or 1, whereas Wilson fluctuates around the nominal level and is a good practical compromise, although its coverage can still deteriorate near the boundaries and may be less favorable than HPD intervals there (Dalitz, 2018).
2. Rare-event adjustment by relative margin of error
In rare-event inference, the principal “adjustment” is often not a new Wilson formula but a new precision criterion. For event probabilities as small as 2 or 3, a fixed absolute margin of error can be practically meaningless. The relevant quantity is the relative margin of error
4
together with the expected width
5
the expected margin of error
6
and the expected relative margin of error
7
Performance is then judged jointly with the expected coverage probability
8
This framework was developed specifically for rare-event binomial proportions and applied to Wald, Clopper–Pearson, Wilson, and Agresti–Coull intervals (McGrath et al., 2021).
Within that framework, the main empirical conclusion is that Wilson is usually the most attractive default method. It is typically close to nominal coverage, less conservative than Clopper–Pearson, better behaved than Wald at small or moderate 9, and often requires the smallest sample size to satisfy both coverage and relative-width criteria. The same study recommends planning with 0 as a practical range, regards 1 as a loose upper bound compatible with order-of-magnitude precision, and treats 2 as generally unacceptable (McGrath et al., 2021).
A central point is terminological. In this rare-event setting, the literature does not introduce a special modified Wilson interval formula. The “adjusted Wilson” idea is conceptual and operational: retain the standard Wilson score interval, but evaluate and plan it using 3 rather than absolute width. This distinction matters because it separates interval algebra from inferential adequacy in extreme-probability regimes (McGrath et al., 2021).
3. Finite-population-corrected Wilson intervals
A mathematically distinct adjustment arises when the target of inference is not an infinite-population parameter but a finite population of size 4. In that setting, the Wilson interval is generalized by replacing the nominal sample size 5 with an effective sample size. For inference about the finite-population proportion
6
the effective sample size is
7
For inference about the unsampled proportion
8
the effective sample size is
9
The adjusted intervals are then 0 and 1 respectively (O'Neill, 2021).
This correction preserves the Wilson form while changing its operating characteristics in a way that tracks the inferential target. For the population proportion, 2 yields 3, so the interval reduces to the ordinary Wilson interval; if 4, then 5 and the interval collapses to the point estimate 6. For the unsampled proportion, 7 again gives the ordinary Wilson form, but 8 yields 9 and the interval 0, reflecting the absence of an unsampled remainder (O'Neill, 2021).
The finite-population generalization also inherits monotonicity and boundedness properties through the effective sample sizes. For the population-proportion interval, width decreases as confidence decreases or as 1 increases, but increases with 2, because a fixed sample represents a smaller fraction of a larger population. The unsampled-proportion interval is structurally different: it obeys the lower bound
3
so for fixed finite 4 it has a nonzero minimum width and is not consistent in the same way as the population-proportion interval (O'Neill, 2021).
4. Exact-coverage modification through the h-function
Another use of “adjusted Wilson” refers to exactification by test inversion. The h-function method starts from any confidence interval
5
and defines the distance-to-interval statistic
6
From this, one constructs
7
and then the modified interval
8
Theorem 1 states that 9 is a valid p-value and that 0 is a 1 exact confidence interval (Wang, 2021).
Applied to the Wilson interval for a binomial proportion, this produces the modified Wilson interval denoted 2. The point is not to alter Wilson’s algebraic score formula but to wrap it in an exact inversion procedure. In the binomial example with 3, the original Wilson interval 4 has infimum coverage probability 5, whereas the modified interval 6 has 7 coverage. As expected, the modified interval is wider than the original Wilson interval (Wang, 2021).
This line of work also gives a refinement theorem. If the starting interval is already exact, then the same modification yields a subset interval, and repeated modification produces a nonincreasing sequence of exact intervals converging to a fixed point. In that sense, the adjusted Wilson interval is part of a broader exact-CI calculus rather than a single named closed form (Wang, 2021).
5. Pseudo-observation adjusted Wilson intervals
A different literature uses “adjusted Wilson” for pseudo-observation smoothing. The adjusted Wilson of type 8 is defined by
9
Operationally, one adds 0 pseudo-observations, half successes and half failures, and then computes a Wald interval using the adjusted estimate 1 and the enlarged sample size 2 (Kahouadji, 13 Aug 2025).
Under this nomenclature, the adjusted Wilson interval is a smoothed version of Wald and is compared directly with standard Wald and standard Wilson across a large simulation grid: 3, 4, and confidence levels 5, 6, and 7. The summary diagnostic is the satisfactory pixel percentage, defined through rainbow color-coded coverage plots in which “pink” pixels indicate coverage at least as large as the nominal level (Kahouadji, 13 Aug 2025).
The reported best pseudo-observation counts are 8 for 9, 0 for 1, and 2 for 3. On the 4 grid, the satisfactory pixel percentages are 5 for adjusted Wilson 6 at 7, 8 for adjusted Wilson 9 at 0, and 1 for adjusted Wilson 2 at 3, each exceeding the corresponding Wilson values 4, 5, and 6, as well as the corresponding Wald values 7, 8, and 9 (Kahouadji, 13 Aug 2025). This pseudo-count interpretation is close in spirit to the Agresti–Coull adjustment, and it is one source of the term’s terminological ambiguity.
6. Comparative interpretation and application-specific generalizations
A common misconception is that the adjusted Wilson confidence interval denotes a single universally accepted formula. The literature does not support that interpretation. Depending on context, the “adjustment” may mean a relative-margin-of-error criterion for rare events, a finite-population correction through 0 or 1, exactification through the h-function, or pseudo-observation smoothing by 2 (McGrath et al., 2021, O'Neill, 2021, Wang, 2021, Kahouadji, 13 Aug 2025).
Comparative performance is correspondingly context-dependent. In general-purpose binomial inference, Wilson is a preferred closed-form method relative to Wald, but not an unqualified optimum: HPD intervals can have better coverage near 3 close to 4 or 5, albeit at the cost of numerical computation (Dalitz, 2018). Coverage-adjusted Clopper–Pearson intervals, built by tuning the nominal level to match mean coverage under a prior or posterior weighting distribution, are often superior to Wilson and Jeffreys near the boundaries, whereas Wilson seems preferable around 6 (Thulin, 2012). These results indicate that “adjusted Wilson” should not be read as a universal dominance claim over all competitors.
The phrase also appears in domain-specific extensions where the Wilson logic is embedded into a larger inferential device. Under differential privacy, a principled Wilson-based two-step interval propagates uncertainty from the privatized proportion by drawing plausible latent sample proportions and aggregating the corresponding Wilson intervals; it tends to have high coverage but is conservative, and Bayesian credible intervals are often preferred overall (Kao et al., 4 Nov 2025). In selective binary classification, Wilson is generalized to local weighted Bernoulli experiments through Wilson Score Kernel Density Estimation, producing per-instance confidence bounds used for abstention decisions (Iversen et al., 24 Feb 2026). Outside classical interval estimation, an “improved Wilson score interval method” for answer ranking combines the Wilson lower bound with a Spotlight Index, and a score/Wilson-style adaptation to the trinomial Net Promoter Score shrinks the estimator toward 7, although in that NPS study the best practical 8 performer is the Adjusted Wald 9 rather than the score/Wilson family (Cao, 2018, Rocks, 2016).
Taken together, these strands establish the adjusted Wilson confidence interval as a family resemblance concept rather than a single object. What remains stable is the Wilson score interval’s role as a structurally well-behaved alternative to Wald. What changes from paper to paper is the target of correction: rare-event precision, finite-population coherence, exact coverage, pseudo-count smoothing, privacy-noise propagation, local nonparametric inference, or application-specific ranking and multinomial adaptation.