Papers
Topics
Authors
Recent
Search
2000 character limit reached

SDR Variance Estimates in Small Domains

Published 18 Aug 2026 in stat.ME | (2608.17353v1)

Abstract: Successive Difference Replication (SDR) is a replication based method of variance estimation introduced by Fay and Train (1995) for estimators based on complex multistage surveys, especially those including a final systematic sampling stage. The method has been used for many years as the primary variance-estimation methodology in large national household surveys administered by the Census Bureau, including the American Community Survey and also the Current Population Survey's monthly estimates based on self-representing strata. In settings where it is applied, generally no second method of variance estimation has been available, so the performance of SDR has been studied via simulation by various authors, for variances of survey totals and of nonlinear survey estimators. This paper begins with a thorough exposition of the SDR method and review of previously published results on the small-domain biases of SDR variance estimation. It is shown that the number D of cycles used in implementing SDR should be 3 or larger, in order to control the variability of SDR estimates, but need not be larger than 5. Beyond that, the value of D is virtually irrelevant to the occurrence of small-domain bias in SDR. The SDR method is shown via theoretical formulas and simulation to inflate average estimated variances in small domains by amounts that vary systematically with the patterns of attribute means and variances and survey weights in consecutively enumerated strata. The degree of average variance inflation is generally moderate, no more than 15 percent in domains with sample size 20, but can be larger in special settings. Moreover, SDR estimates are extremely variable in small domains, with standard deviations often far larger than any biases.

Authors (2)

Summary

  • The paper provides a in-depth analysis of Successive Difference Replication (SDR) variance estimation in small domains, exploring both theoretical and empirical aspects.
  • A detailed simulation shows that the variability in SDR estimation is high, not significantly reduced by tuning parameters, making variability dominate any potential biases.

Background and motivation

Successive Difference Replication (SDR), introduced by Fay and Train, is the primary variance-estimation method for large Census Bureau household surveys including the American Community Survey (ACS) and the Current Population Survey (CPS) self-representing strata. It was designed for complex multi-stage designs with a final systematic-sampling stage, where no design-unbiased variance formula exists. Despite its long operational use — including as an input to Small Area Estimation (SAE) programs such as SAIPE and Voting Rights Act Section 203 determinations — its behavior for domain estimates from non-SRS samples had not been systematically studied. Prior simulations (Huang and Bell; Sukasih and Jang; Ash; Opsomer et al.) largely considered whole-sample totals from SRS samples of iid superpopulations. This paper fills that gap with a theoretical treatment of SDR bias in stratified SRS settings, simulation evidence on small-domain variance inflation, and practical recommendations.

The method and a quadratic-form representation

SDR defines replicate weight factors from adjacent row pairs of an R×RR \times R Hadamard matrix via fj,r=1+2−3/2(hγj,r−hγj+1,r)f_{j,r} = 1 + 2^{-3/2}(h_{\gamma_j,r} - h_{\gamma_{j+1},r}), with replicate totals Y^r\hat{Y}_r combined as V^SDR(Y^)=4R∑r(Y^r−Y^)2\hat{V}^{\tt SDR}(\hat{Y}) = \frac{4}{R}\sum_r (\hat{Y}_r - \hat{Y})^2. The paper's central structural contribution is to rewrite this estimator, for arbitrary sample size nn and number of permutation cycles DD, as a quadratic form V^SDR=∑g,g′Qg,g′Y(g)Y(g′)\hat{V}^{\tt SDR} = \sum_{g,g'} Q_{g,g'} Y^{(g)} Y^{(g')} in distinct weighting-group totals Y(g)Y^{(g)}, with QQ depending only on combinatorial properties of the index sequence (γj)(\gamma_j) and not on the particular Hadamard matrix. For fj,r=1+2−3/2(hγj,r−hγj+1,r)f_{j,r} = 1 + 2^{-3/2}(h_{\gamma_j,r} - h_{\gamma_{j+1},r})0 this recovers Wolter's SD2 estimator as half a sum of squared differences of consecutive group totals; for fj,r=1+2−3/2(hγj,r−hγj+1,r)f_{j,r} = 1 + 2^{-3/2}(h_{\gamma_j,r} - h_{\gamma_{j+1},r})1 the matrix fj,r=1+2−3/2(hγj,r−hγj+1,r)f_{j,r} = 1 + 2^{-3/2}(h_{\gamma_j,r} - h_{\gamma_{j+1},r})2 is no longer tridiagonal but is computable in closed form. These formulas allow exact computation of SDR estimates without retaining replicates, and are implemented in an accompanying R package.

The paper also documents that all known implementations with fj,r=1+2−3/2(hγj,r−hγj+1,r)f_{j,r} = 1 + 2^{-3/2}(h_{\gamma_j,r} - h_{\gamma_{j+1},r})3, fj,r=1+2−3/2(hγj,r−hγj+1,r)f_{j,r} = 1 + 2^{-3/2}(h_{\gamma_j,r} - h_{\gamma_{j+1},r})4 (including ACS's) necessarily produce duplicated row-index pairs — 22 duplicates in the actual ACS indexing — and shows that a duplicate-free set of at most 7 cycles exists within the standard construction, while arguing statistically that this is immaterial.

Choice of fj,r=1+2−3/2(hγj,r−hγj+1,r)f_{j,r} = 1 + 2^{-3/2}(h_{\gamma_j,r} - h_{\gamma_{j+1},r})5 and fj,r=1+2−3/2(hγj,r−hγj+1,r)f_{j,r} = 1 + 2^{-3/2}(h_{\gamma_j,r} - h_{\gamma_{j+1},r})6

Under iid superpopulations with SRS sampling, the paper derives fj,r=1+2−3/2(hγj,r−hγj+1,r)f_{j,r} = 1 + 2^{-3/2}(h_{\gamma_j,r} - h_{\gamma_{j+1},r})7 exactly in terms of the first four moments of fj,r=1+2−3/2(hγj,r−hγj+1,r)f_{j,r} = 1 + 2^{-3/2}(h_{\gamma_j,r} - h_{\gamma_{j+1},r})8 and the sparse matrix representation. The resulting coefficients of variation of fj,r=1+2−3/2(hγj,r−hγj+1,r)f_{j,r} = 1 + 2^{-3/2}(h_{\gamma_j,r} - h_{\gamma_{j+1},r})9 relative to Y^r\hat{Y}_r0 hover near 0.17–0.20 regardless of Y^r\hat{Y}_r1, confirming Krewski–Rao/Wolter intuition that precision depends on the number of replicates (Y^r\hat{Y}_r2), not the sample size. Two actionable conclusions follow:

  • Y^r\hat{Y}_r3 is sufficient: beyond three or four cycles, further cycling yields essentially no reduction in variability of the variance estimate.
  • Y^r\hat{Y}_r4 is nearly irrelevant to small-domain bias: simulations across many superpopulations show strikingly equivalent distributions of Y^r\hat{Y}_r5 for Y^r\hat{Y}_r6 and Y^r\hat{Y}_r7 across expected domain sample sizes up to 75 and beyond.

This directly contradicts Ash's earlier suggestion that more cycles improve bias control under deterministic attribute patterns.

Bias theory for stratified SRS designs

For stratified SRS samples from superpopulations with stratumwise constant means and variances, and Y^r\hat{Y}_r8, the paper proves an interpretable decomposition of Y^r\hat{Y}_r9: the target variance plus a leading term

V^SDR(Y^)=4R∑r(Y^r−Y^)2\hat{V}^{\tt SDR}(\hat{Y}) = \frac{4}{R}\sum_r (\hat{Y}_r - \hat{Y})^20

plus a lower-order term. The resulting approximate relative bias is governed by how sharply these quantities — effectively the parameters V^SDR(Y^)=4R∑r(Y^r−Y^)2\hat{V}^{\tt SDR}(\hat{Y}) = \frac{4}{R}\sum_r (\hat{Y}_r - \hat{Y})^21 — change between consecutively ordered strata. Special cases make the dependence concrete:

  • Single-stratum domains: relative bias is approximately proportional to V^SDR(Y^)=4R∑r(Y^r−Y^)2\hat{V}^{\tt SDR}(\hat{Y}) = \frac{4}{R}\sum_r (\hat{Y}_r - \hat{Y})^22 and decreases with the squared coefficient of variation of attributes.
  • Alternating strata (alternating sample sizes or alternating means): relative bias scales linearly with the number of alternating strata V^SDR(Y^)=4R∑r(Y^r−Y^)2\hat{V}^{\tt SDR}(\hat{Y}) = \frac{4}{R}\sum_r (\hat{Y}_r - \hat{Y})^23 and with V^SDR(Y^)=4R∑r(Y^r−Y^)2\hat{V}^{\tt SDR}(\hat{Y}) = \frac{4}{R}\sum_r (\hat{Y}_r - \hat{Y})^24, so severe alternation can inflate bias multiplicatively even for moderate domain sizes.
  • Uniformly small domain fractions V^SDR(Y^)=4R∑r(Y^r−Y^)2\hat{V}^{\tt SDR}(\hat{Y}) = \frac{4}{R}\sum_r (\hat{Y}_r - \hat{Y})^25: the numerator scales as the square of average V^SDR(Y^)=4R∑r(Y^r−Y^)2\hat{V}^{\tt SDR}(\hat{Y}) = \frac{4}{R}\sum_r (\hat{Y}_r - \hat{Y})^26 while the denominator scales linearly, so bias is negligible for very small domains embedded in large surveys.
  • Isolated strata: the one setting found where relative bias need not vanish as domain sample size grows, though it remains modest when the average ratio V^SDR(Y^)=4R∑r(Y^r−Y^)2\hat{V}^{\tt SDR}(\hat{Y}) = \frac{4}{R}\sum_r (\hat{Y}_r - \hat{Y})^27 stays below roughly 0.10–0.15.

An important caveat acknowledged by the authors is that SDR has no built-in finite population correction, unlike the true stratified-SRS variance, and the theory offers no clear guidance on when or how to apply such a correction when sampling fractions and domain proportions vary across strata.

Simulation findings

Factorial simulations (10 strata, Gamma-distributed stratum means, Beta-distributed domain fractions with varying correlation and dispersion, 10,000 samples per scenario) confirm the theory:

  • When expected domain sample size exceeds 25, relative bias stays below 20% even in the most extreme scenarios; for domains of size around 20, average inflation is generally no more than about 15%, though larger biases occur in special configurations.
  • Bias is largest when overall domain proportion is large, cross-stratum correlation of domain indicators is zero (maximizing V^SDR(Y^)=4R∑r(Y^r−Y^)2\hat{V}^{\tt SDR}(\hat{Y}) = \frac{4}{R}\sum_r (\hat{Y}_r - \hat{Y})^28-alternation), and the CV of V^SDR(Y^)=4R∑r(Y^r−Y^)2\hat{V}^{\tt SDR}(\hat{Y}) = \frac{4}{R}\sum_r (\hat{Y}_r - \hat{Y})^29 is high.
  • Variability dominates bias: density plots of single-sample ratios nn0 show extremely wide spreads in small domains, often far exceeding any systematic bias. This is arguably the most consequential finding for practice: the noise of SDR variance estimates in small domains cannot be reduced by tuning the method.

The authors also test a remediation strategy — replacing nn1 with the median over random permutations of stratum order — and find it reduces variability but does not reliably reduce bias; it can introduce negative bias when the original ordering is already favorable. They explicitly decline to recommend it.

Relevance to small area estimation

Since SDR variances feed directly into SAE linking models (often after smoothing through Generalized Variance Functions), the results bear on programs such as SAIPE and VRA Section 203 determinations. The paper's premise — following Huang and Bell — that bias should be small when domain proportions are uniformly tiny holds, but the alternating and isolated-stratum examples imply that among the many moderate-sized domains typical of agency applications, a few may carry SDR variance biases in the 20–50% range, and conditional predictions for those areas can be seriously affected (consistent with Bell's sensitivity analysis). Practical suggestions include treating structurally empty strata as such (omitting sampled units from strata containing no domain elements) and, where outcome-variable structure permits, ordering geographic units to respect estimated nn2 patterns — while cautioning against blind permutation.

Limitations and open questions

The theoretical formulas assume stratified SRS from stratumwise iid superpopulations with complete response; nonresponse mechanisms are explicitly ignored throughout, and clustered population structures are not addressed — nor were they in any prior published study of SDR bias. Real surveys sort respondents in complicated ways (ACS uses a re-sort reflecting both sampling stages; CPS modifies geographic sorting by month), so the mapping from the idealized nn3 analysis to operational sort orderings is only partially established; the CPS anomaly documented by Trudell et al., arising from baseweights defined at the sampling stage rather than on respondent files, indicates that implementation details matter in ways this framework does not yet capture. Whether the theory extends to surveys with iid clusters as units — as CPS's 4-household cluster replicate weighting implicitly assumes — is asserted as plausible but not proven here. Finally, the question of how best to reorder respondents when outcome-specific parameters nn4 are unknown remains unresolved, since the optimal ordering depends on quantities the analyst cannot observe.

Conclusion

This paper provides the first systematic theoretical and simulation-based account of SDR variance estimation for small-domain totals, delivering closed-form quadratic-form representations valid for all nn5 and nn6, exact bias formulas under stratified SRS, and clear guidance: use nn7 (with nn8, or 160 where feasible), expect modest upward bias — generally below ~15% for domains of size 20, occasionally larger under alternating or isolated stratum structure — and recognize that the dominant problem in small domains is the high variability of the SDR variance estimate itself, which no choice of replication parameters can remedy.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.