- The paper introduces the Post-Hoc Inference Engine, which calibrates survey weights to hierarchical Bayes posterior draws so researchers can estimate uncertainty for cross-classified statistics.
- Uncorrected posterior-calibrated intervals can severely undercover non-calibration cells, reaching 13.5%, while design-based compositional variance exceeds model variance by 8–40 times.
- The proposed Calibrated Bayes Interval combines sampling and posterior uncertainty, raising Monte Carlo coverage to 92–99% while keeping coefficients of variation between 0.8% and 4.9%.
Problem and contribution
This paper addresses the inference problem that arises after a survey has been designed and fitted under the Hierarchical Bayes (HB) small-area framework of Tam (Tam, 18 Mar 2026). That earlier work showed that combining Bethel multivariate allocation with HB modelling can reduce a Labour Force Survey sample by 80% (from n∗=91,308 to nHB=18,262) while maintaining domain-level precision and near-nominal coverage. The present paper tackles the operational sequel: once such a reduced survey is fielded, a National Statistical Office (NSO) will publish cross-classified tables — employment by occupation, income by employment status, health status by occupation — many involving variables outside the original HB calibration system. The question is how to attach valid interval estimates to these post-hoc tabulations without either understating uncertainty or propagating posterior draws mechanically in a way that yields poor repeated-sampling coverage.
The central methodological device is the Post-Hoc Inference Engine (PHIE). For each MCMC draw T^(b) of the HB posterior domain totals, PHIE performs a chi-square calibration of the Horvitz–Thompson weights to that draw, producing replicate weights; applying them to any cross-tabulated statistic yields an empirical posterior distribution whose quantiles form a credible interval (CrI). The paper's key structural insight is a three-tier taxonomy showing that not all cells share the same inferential status, and that uncorrected PHIE intervals can undercover severely for non-calibration-supported cells.
The PHIE construction
The setting is a stratified simple random sample with HT weights, partitioned into D domains, with a calibration design vector stacking V×D domain-variable blocks (p=VD constraints). Lemma 1 gives the closed-form solution to the chi-square distance calibration problem: for any target vector t,
wi′(t)=wi(1+(t−T^)⊤G−1yi),
where G=Y⊤WY is assumed full rank. This is exactly the GREG weight of Deville and Särndal [deville1992] applied with HB posterior draws rather than fixed population totals as targets. A single unified weight set simultaneously satisfies all p constraints. Lemma 2 then establishes, via the continuous mapping theorem, that the empirical quantiles of the replicate statistics nHB=18,2620 converge to the true posterior quantiles of the induced functional nHB=18,2621 as nHB=18,2622.
The paper is careful about scope: the resulting interval is a credible interval for the induced posterior only, treating the sample, design vectors, and design weights as fixed. Whether this coincides with the full posterior depends on the cell's relationship to the calibration system — the motivation for the taxonomy.
Three-tier taxonomy and the CBI correction
Tier 1-E cells reproduce a calibration constraint exactly. By the calibration property, the PHIE CrI is identical to reading the interval directly from MCMC output; coverage is exact in a single run.
Tier 2-CA / Tier 2-NCA cells sum a calibration variable over an attribute filter (calibration-derived or not). Because the filter cuts across domains, the cell total decomposes via fixed estimated within-domain shares nHB=18,2623, so PHIE propagates only Component 2 (domain-total posterior variance) and ignores Component 1 (design-based sampling variance of the shares). By the law of total variance, this produces a "quasi-posterior" interval that undercovers whenever compositional variability is non-negligible. The remedy is a Calibrated Bayes Interval (CBI) in the sense of Little [little2012], combining Taylor-linearised design-based variance of the share (with stratum design effects) with the MCMC variance of domain totals:
nHB=18,2624
Two pre-publication diagnostics are introduced: nHB=18,2625, detecting cells with no distinctive calibration footprint, and nHB=18,2626, detecting near-orthogonality between the cell's calibration profile and the posterior-mean residual. Both explain when Component 2 collapses toward zero.
Tier 3-NCV cells involve a non-calibration outcome variable. The extension treats the outcome as a ratio-estimator numerator with denominator given by the most correlated calibration variable nHB=18,2627 ("linking variable"). When nHB=18,2628 is large the ratio estimator is efficient; when nHB=18,2629, Component 2 vanishes and the CBI reduces approximately to a design-based confidence interval, which the paper recommends as the primary published interval in that case.
Empirical results
The evaluation uses 5% Australian 2021 Census microdata (T^(b)0, T^(b)1, T^(b)2 strata, T^(b)3, T^(b)4, T^(b)5), where population truth permits direct coverage assessment across 200 Monte Carlo replications. Three findings stand out.
First, Tier 1-E coverage is near-nominal: single-run coverage is 8/8 and MC coverage averages 99.2%, confirming the calibration property empirically.
Second, uncorrected PHIE intervals undercover substantially for Tier 2 and Tier 3 cells — as low as 13.5% for health condition by occupation and 18.5% for the nil-income band. This is the paper's strongest quantitative warning: propagating only domain-total posterior uncertainty is structurally inadequate because compositional sampling variability dominates. Indeed, for every occupation group, Component 1 exceeds Component 2 by a factor of 8–40×, establishing that cross-tabulation uncertainty is driven primarily by design-based sampling variability rather than HB model uncertainty.
Third, the CBI correction restores coverage robustly: MC coverage rises to 92–99% across all tiers and both Tier 3 examples (strong linking correlation T^(b)6 for income; negligible T^(b)7 for health condition). Residual deviations from 95% are within two Monte Carlo standard errors (T^(b)8 pp at 200 replications).
A practically important caveat concerns single-run versus Monte Carlo coverage. For income bands, single-run CBI coverage is 11/15; for health condition by occupation it is only 3/8. The paper attributes these misses to realisation-to-realisation variability of the plug-in Component 1 estimate — concentrated in cells with small T^(b)9 — rather than systematic failure, but concedes that an NSO conducting one survey cannot rely on single-run coverage alone and should validate against Census benchmarks or administrative proxies before release.
Publishability
CBI-based coefficients of variation remain within standard publication thresholds: all CVs lie below the 5% publishable limit, ranging from 0.8–4.9% across tiers, with the maximum (4.9%, health condition among Machinery Operators) at the boundary. This indicates the corrected intervals are wide enough to be honest yet narrow enough for routine NSO dissemination.
Limitations and open questions
The paper states several limitations explicitly. Validity is assessed empirically rather than asymptotically: the Bernstein–von Mises conditions are not satisfied in this hierarchical, finite-sample setting, so no general mathematical guarantee of nominal coverage exists for arbitrary cross-classified statistics. The CBI is a plug-in estimator — uncertainty in the estimated design effects D0 is not propagated (potentially inflating intervals in thin strata), and replacing true shares D1 introduces D2 plug-in error. The framework also inherits the adequacy of the underlying HB linking model: if covariates are weak or the model misspecified, posterior totals may borrow strength in the wrong direction and posterior uncertainty may be too small, overstating the benefit of sample reduction. Finally, the reduced MCMC settings used in the Monte Carlo study (burnin 200, 500 iterations) were a computational necessity, though sensitivity checks on 20 replications with full settings differed by less than 1.5 percentage points. Open questions include whether the diagnostics D3 and D4 can be converted into formal pre-publication acceptance rules, and how single-run coverage shortfalls in low-D5 cells should be handled operationally beyond the recommended supplementary design-based intervals.
Conclusion
This paper supplies the missing inferential machinery for publishing cross-classified statistics from an HB-calibrated unit record file. Its contributions are the posterior-calibrated replicate-weight engine, the demonstration that Tier 1-E cells inherit exact posterior intervals while filtered and non-calibration cells do not, the CBI correction restoring near-nominal repeated-sampling coverage, and the ratio-estimator extension for non-calibration variables. The empirical message is direct: cross-tabulations from a mini-max HB survey must not be treated as ordinary weighted tabulations, since their dominant uncertainty source is compositional sampling variability, not model uncertainty — but a two-component calibrated Bayes interval keeps both coverage and publishable precision intact for the cases examined.