Papers
Topics
Authors
Recent
Search
2000 character limit reached

Serfling’s Inequality: Finite-Sampling Tail Bounds

Updated 18 February 2026
  • Serfling’s inequality is a finite-sampling exponential concentration inequality that extends Hoeffding’s bounds to sampling without replacement.
  • It utilizes a telescoping sum and conditional moment generating functions to derive sharp tail bounds incorporating finite-sample corrections.
  • The method is crucial in hypergeometric settings and two-sample empirical processes, enhancing statistical inference in nonparametric tests.

Serfling’s inequality is a finite-sampling exponential concentration inequality that extends Hoeffding’s classical bounds for sums of independent bounded random variables to the case of sampling without replacement from a finite population. It provides sharp exponential tail bounds for deviations of sample means from the population mean under sampling fraction corrections, with particular relevance for hypergeometric and empirical process contexts (Greene et al., 2015).

1. Formal Statement and Interpretation

Let {c1,…,cN}\{c_1,\ldots,c_N\} be a finite population (“urn”) with ci∈Rc_i \in \mathbb{R}, population mean μN=1N∑i=1Nci\mu_N = \frac{1}{N}\sum_{i=1}^N c_i, variance σN2=1N∑i=1N(ci−μN)2\sigma_N^2 = \frac{1}{N}\sum_{i=1}^N (c_i - \mu_N)^2, minimum aN=min⁡icia_N = \min_i c_i, and maximum bN=max⁡icib_N = \max_i c_i. Consider sampling without replacement n≤Nn \le N elements, yielding Y1,…,YnY_1,\ldots,Y_n, with sample mean Y‾n=1n∑i=1nYi\overline Y_n = \frac1n\sum_{i=1}^n Y_i. Define the finite-sampling fractions fn∗=n−1Nf_n^* = \frac{n-1}{N} and ci∈Rc_i \in \mathbb{R}0.

Serfling’s inequality states that for any ci∈Rc_i \in \mathbb{R}1,

ci∈Rc_i \in \mathbb{R}2

A frequently employed specialization is when ci∈Rc_i \in \mathbb{R}3: ci∈Rc_i \in \mathbb{R}4 These forms quantify the upper tail probabilities of the sample mean deviating from the population mean under sampling without replacement, with explicit finite-sample corrections.

2. Analytical Strategy and Key Proof Elements

The proof adapts Hoeffding’s martingale-based approach for independent variables to the without-replacement regime. The essential steps are:

  • Expressing ci∈Rc_i \in \mathbb{R}5 as a telescoping sum of conditional expectations.
  • Stepwise control of the conditional moment-generating function ci∈Rc_i \in \mathbb{R}6, leveraging that at each stage, the remaining population values remain bounded in ci∈Rc_i \in \mathbb{R}7.
  • Demonstrating, by induction,

ci∈Rc_i \in \mathbb{R}8

  • Invoking Markov’s inequality ci∈Rc_i \in \mathbb{R}9 and optimizing μN=1N∑i=1Nci\mu_N = \frac{1}{N}\sum_{i=1}^N c_i0 to obtain the exponential rate. The correction factor μN=1N∑i=1Nci\mu_N = \frac{1}{N}\sum_{i=1}^N c_i1 directly arises from the diminishing uncertainty after each observed draw, distinguishing the without-replacement scenario from the i.i.d. case (Greene et al., 2015).

3. Relationship to Hoeffding’s Inequality and Refinements

Hoeffding’s classical bound for independent random variables μN=1N∑i=1Nci\mu_N = \frac{1}{N}\sum_{i=1}^N c_i2 with μN=1N∑i=1Nci\mu_N = \frac{1}{N}\sum_{i=1}^N c_i3 is

μN=1N∑i=1Nci\mu_N = \frac{1}{N}\sum_{i=1}^N c_i4

For sampling with replacement from μN=1N∑i=1Nci\mu_N = \frac{1}{N}\sum_{i=1}^N c_i5, substituting μN=1N∑i=1Nci\mu_N = \frac{1}{N}\sum_{i=1}^N c_i6 yields

μN=1N∑i=1Nci\mu_N = \frac{1}{N}\sum_{i=1}^N c_i7

Serfling’s bound, with its μN=1N∑i=1Nci\mu_N = \frac{1}{N}\sum_{i=1}^N c_i8 exponent augmentation, tightens the concentration in the without-replacement regime. Since μN=1N∑i=1Nci\mu_N = \frac{1}{N}\sum_{i=1}^N c_i9, the bound always strictly improves upon the naive i.i.d. Hoeffding bound for sampling without replacement.

A conceivable enhancement is to replace σN2=1N∑i=1N(ci−μN)2\sigma_N^2 = \frac{1}{N}\sum_{i=1}^N (c_i - \mu_N)^20 with the more accurate σN2=1N∑i=1N(ci−μN)2\sigma_N^2 = \frac{1}{N}\sum_{i=1}^N (c_i - \mu_N)^21, which has been achieved in special cases but remains open in general. Bennett-type refinements, applying Ehm’s representation of the hypergeometric as sums of independent Bernoullis, yield

σN2=1N∑i=1N(ci−μN)2\sigma_N^2 = \frac{1}{N}\sum_{i=1}^N (c_i - \mu_N)^22

where σN2=1N∑i=1N(ci−μN)2\sigma_N^2 = \frac{1}{N}\sum_{i=1}^N (c_i - \mu_N)^23, σN2=1N∑i=1N(ci−μN)2\sigma_N^2 = \frac{1}{N}\sum_{i=1}^N (c_i - \mu_N)^24. For binary populations (σN2=1N∑i=1N(ci−μN)2\sigma_N^2 = \frac{1}{N}\sum_{i=1}^N (c_i - \mu_N)^25), an explicit Hoeffding-style bound follows (Greene et al., 2015).

4. Hypergeometric Specialization

Consider the population made up of σN2=1N∑i=1N(ci−μN)2\sigma_N^2 = \frac{1}{N}\sum_{i=1}^N (c_i - \mu_N)^26 ones and σN2=1N∑i=1N(ci−μN)2\sigma_N^2 = \frac{1}{N}\sum_{i=1}^N (c_i - \mu_N)^27 zeros. Here, σN2=1N∑i=1N(ci−μN)2\sigma_N^2 = \frac{1}{N}\sum_{i=1}^N (c_i - \mu_N)^28 is the sample proportion of ones; σN2=1N∑i=1N(ci−μN)2\sigma_N^2 = \frac{1}{N}\sum_{i=1}^N (c_i - \mu_N)^29; aN=min⁡icia_N = \min_i c_i0. Serfling’s bound yields

aN=min⁡icia_N = \min_i c_i1

Specifically, for aN=min⁡icia_N = \min_i c_i2,

aN=min⁡icia_N = \min_i c_i3

Such hypergeometric tail bounds are central in settings where binary attributes are counted under finite sampling, for example, in quality control and resampling inference (Greene et al., 2015).

5. Finite-Sampling Correction Terms and Open Questions

The “primitive” finite-sampling correction in Serfling’s exponent is aN=min⁡icia_N = \min_i c_i4. The “true” correction aN=min⁡icia_N = \min_i c_i5, corresponding to the classical finite-population variance reduction, appears in recent Bennett-type and Hoeffding-type inequalities but has not been universally established for Serfling’s bound in general settings. The proximity of these correction factors is essential for maximal sharpness in empirical process and finite-population inferential theory. Existing results and evidence suggest such refinement is plausible and likely achievable in further generalizations, especially under independence approximations (Greene et al., 2015).

6. Applications to Two-Sample Empirical Process Statistics

Serfling’s inequality forms the backbone of corrected exponential tail bounds for two-sample Kolmogorov–Smirnov (K–S) statistics. For independent samples aN=min⁡icia_N = \min_i c_i6 and aN=min⁡icia_N = \min_i c_i7 from the same continuous distribution aN=min⁡icia_N = \min_i c_i8, with empirical CDFs aN=min⁡icia_N = \min_i c_i9 and bN=max⁡icib_N = \max_i c_i0, the two-sample one-sided K–S statistic is

bN=max⁡icib_N = \max_i c_i1

Viewing the pooled empirical CDF bN=max⁡icib_N = \max_i c_i2 as the population, each empirical CDF is a sample without replacement, and Serfling’s bound applies. In the balanced case bN=max⁡icib_N = \max_i c_i3, with bN=max⁡icib_N = \max_i c_i4,

bN=max⁡icib_N = \max_i c_i5

with the two-sided statistic

bN=max⁡icib_N = \max_i c_i6

The finite-sampling correction bN=max⁡icib_N = \max_i c_i7 adjusts the exponent of the classical Dvoretzky–Kiefer–Wolfowitz–Massart inequality bN=max⁡icib_N = \max_i c_i8. For unbalanced samples (bN=max⁡icib_N = \max_i c_i9), conjecturally, similar exponential bounds hold: n≤Nn \le N0 These corrections have significance for the tightness and calibration of empirical process-based inference, notably in nonparametric hypothesis testing and distributional comparison (Greene et al., 2015).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Serfling’s Inequality.