DPComb: Count Modeling & p-value Adjustment
- DPComb is a label for two unrelated methodologies: one for Bayesian adaptive COM-Poisson regression to model count data and one for optimal discrete p-value combination in multiple testing.
- In count-data modeling, DPComb employs a Dirichlet process mixture of COM-Poisson regressions to flexibly capture conditional distributions and extract noncrossing quantiles without relying on jittering.
- For discrete p-value combination, DPComb uses an optimal-transport framework to adjust independent p-values under classical additive rules, ensuring accurate Type I error calibration.
DPComb is used in the cited literature for two distinct statistical constructions. In Bayesian count-data modeling, it denotes a Dirichlet process-based COM-Poisson density regression approach for estimating conditional distributions and deriving noncrossing conditional quantiles without jittering (Chanialidis et al., 2014). In multiple-testing methodology, it denotes an optimal-transport framework, and an R package of the same name, for adjusting and combining independent discrete -values under classical additive combination rules such as Fisher, Pearson, George, Stouffer, and Edgington (Contador et al., 4 Aug 2025). A separate diffusion-model paper explicitly notes that divide-and-conquer posterior sampling is canonically named DCPS rather than DPComb, even if the latter is used informally as an alias (Janati et al., 2024).
1. Terminological scope and disambiguation
The shared label covers unrelated methodological objects in count-data Bayesian nonparametrics and in discrete -value combination. The overlap is nominal rather than conceptual.
| Source | Meaning of DPComb | Domain |
|---|---|---|
| (Chanialidis et al., 2014) | Adaptive DP mixture of COM-Poisson regressions for density regression and count quantiles | Bayesian count-data modeling |
| (Contador et al., 4 Aug 2025) | Optimal adjustment and combination of independent discrete -values; also the R package DPComb | Multiple testing and meta-analysis |
| (Janati et al., 2024) | Not the canonical name; the diffusion method is called DCPS | Bayesian inverse problems with diffusion priors |
This naming collision matters because the two principal usages solve different problems. The count-data DPComb models an entire conditional pmf and obtains quantiles by cdf inversion. The discrete-testing DPComb starts from already computed discrete -values and adjusts transformed test statistics so that classical global null calibrations remain accurate despite discreteness. A plausible implication is that database and software searches for “DPComb” require domain-specific disambiguation.
2. DPComb as Bayesian density regression for count data
In the count-data setting, DPComb was proposed to address the fact that for the conditional cdf is a step function and the conditional quantile
is piecewise-constant and not a continuous function of the model parameters (Chanialidis et al., 2014). The paper identifies three difficulties in direct count quantile regression: discreteness, instability induced by the discontinuity of the quantile map in the parameters, and the limitations of jittering, where is added to form . For small counts, especially 0 or 1, the added noise can dominate the estimated quantiles, and fitting each 2 separately can produce crossing quantiles.
The proposed solution is density regression with an adaptive DP mixture of COM-Poisson regressions. The COM-Poisson pmf is written as
3
with 4 reducing to Poisson, 5 corresponding to under-dispersion, and 6 to over-dispersion. The centered reparameterization 7 yields
8
Under the approximation reported in the paper,
9
and 0 is the mode. This reparameterization improves interpretability relative to 1 when 2.
Regression enters through log links for both centering and dispersion:
3
Hence, under the same approximation,
4
The hierarchical structure places a Dirichlet process prior on component-specific regression parameters 5, so that
6
with a convenient base measure given as 7. The paper further uses a predictor-dependent construction following Dunson, Pillai, and Park (2007), with kernel-based proximity weights between observed predictors. The resulting conditional pmf at a fixed 8 is
9
which allows flexible, covariate-dependent multimodality and dispersion patterns. The paper states that mixtures of COM-Poisson components can approximate any discrete distribution on 0 arbitrarily well.
3. Inference, quantile extraction, and behavior of the count-data method
Quantiles in the count-data DPComb are not modeled directly; they are extracted from the fitted cdf (Chanialidis et al., 2014). For each component,
1
and the mixture cdf is
2
The conditional 3-quantile is then
4
Because 5 is monotone nondecreasing in 6, the map 7 is nondecreasing for each fixed 8, so crossing quantiles are prevented by construction.
A central computational issue is that the COM-Poisson likelihood is doubly intractable because of the normalizing constant 9. The paper therefore combines three ingredients: Neal’s MCMC for nonconjugate DP mixtures, adaptive predictor-dependent mixture ideas of Dunson et al. (2007), and the exchange algorithm. For a generic move 0 with proposal 1 and auxiliary draw 2, the acceptance ratio is
3
so the normalizing constants cancel. For COM-Poisson with 4, the unnormalized kernel is
5
The allocation step uses predictor-dependent prior weights
6
where 7 may be chosen as a Gaussian kernel. Auxiliary samples from COM-Poisson are obtained by rejection sampling, and posterior predictive evaluation of 8 uses truncated series, asymptotics or bounds, and log-scale numerical stabilization.
The paper’s practical guidance includes weakly informative Gaussian priors for 9 and 0, a Gamma prior for the concentration parameter, initialization from Poisson regression, and posterior checks based on trace plots, deviance, observed versus predicted histograms, and empirical versus fitted quantiles. Per sweep computational cost is reported as 1 for allocation updates plus the cost of exchange steps. The paper also notes extensions such as zero-inflated COM-Poisson mixtures, alternative links for 2, alternative predictor-dependent DPs, and alternative count kernels.
Empirically, the paper summarizes two simulation designs with 3: a binomial regression model 4 and a heterogeneous mixture 5. Using mean absolute error of estimated conditional quantiles averaged across 6 and 7, the discrete Bayesian density regression outperforms jittering, and jittering frequently produces crossing quantiles except, in that study, when 8. On housebreakings versus deprivation in Greater Glasgow, the method is reported to capture overdispersion and skewness that standard Poisson regression misses, with estimated 9, 0, and 1 quantiles that better reflect the empirical distribution of counts.
4. DPComb as optimal adjustment and combination of independent discrete 2-values
In the multiple-testing setting, DPComb addresses the problem of combining independent discrete 3-values 4 under the global null
5
when classical continuous calibrations become conservative or miscalibrated because the null support is a grid with jumps (Contador et al., 4 Aug 2025). The framework assumes that each discrete 6-value satisfies
7
with cdf 8 for 9. It considers additive combination statistics of the form
0
where 1 is a strictly increasing continuous cdf with finite first two moments.
The method replaces the discrete transformed component 2 or 3 by a surrogate 4 that is closest in 5 to the continuous ideal 6, 7. Theorem 1 in the paper states that on the event 8 the 9-optimal modification is the conditional mean on the corresponding interval:
0
for the 1 form, and
2
for the 3 form. The adjusted global statistic is
4
For the five classical combiners treated in the paper, the continuous statistics and surrogate calibrations are as follows.
| Method | Continuous statistic | Surrogate calibration |
|---|---|---|
| Fisher | 5 | 6 |
| Pearson | 7 | 8 |
| George | 9 | 0 |
| Stouffer | 1 | 2 |
| Edgington | 3 | 4 |
The paper gives explicit adjusted values for each interval 5. Writing 6 and 7, the formulas are
8
9
00
01
and
02
The framework therefore gives a unified discrete adjustment for Fisher, Pearson, George, Stouffer, and Edgington without adding randomness.
5. Variance matching, asymptotics, and reported empirical behavior
A key result for the discrete-03-value DPComb is the variance decomposition
04
reported as Lemma 1 (Contador et al., 4 Aug 2025). This explains why direct calibration by the continuous null is conservative: the adjusted discrete statistic 05 preserves the mean of the continuous ideal 06 but has smaller variance. DPComb therefore matches both mean and variance through a surrogate family. For Fisher and Pearson this is a Gamma family containing the usual 07 null as a special case; for Stouffer the Normal family remains exact under continuity; for George and Edgington the sums are approximated by Normal laws.
The paper reports closed-form expressions for the per-component variances 08:
09
10
11
12
and
13
with 14. Lemmas 2 and 3 state asymptotic Type I validity in the i.i.d. and independent non-identical cases, respectively:
15
Finite-sample accuracy is linked to the scaled Wasserstein distance 16 and the variance ratio 17, and Lemma 4 gives the upper bound
18
The paper’s simulations evaluate four stylized null cdfs: PL, PR, PC, and PS. The reported findings are that Fisher performs best when mass is concentrated at the right tail (PR) and worst when mass is concentrated at the left tail (PL), Pearson behaves oppositely, all methods are similar and well-calibrated when mass is central (PC), and normal-based surrogates have advantages when mass is at both ends (PS). Empirical Type I error at 19 or 20 converges to nominal for all methods as 21 increases.
Two benchmark power studies connect DPComb to uniformly most powerful behavior when the likelihood ratio test is monotone in a classical combiner. In the geometric example, Fisher is UMP for the right-sided alternative and Pearson is UMP for the left-sided alternative; the paper reports that DPComb’s Fisher and Pearson procedures match the UMP LRT power curves, with small discrepancies attributable to discrete conservativeness when exact 22-quantiles are unattainable. In the circular example, the LRT is monotone in 23, so Edgington is UMP; the adjusted Edgington procedure exhibits the highest power in simulations. A rare-variant case-control study with 2000 subjects and 15 SNPs shows Gene 1 significant for right-sided and two-sided alternatives across multiple combiners, Gene 2 significant for the right-sided alternative only, and Edgington yielding the smallest gene-level 24-value among the five methods.
6. Software implementation, practical use, and the DCPS naming issue
The R package DPComb implements the discrete-25-value methodology end to end (Contador et al., 4 Aug 2025). The paper states that it computes the 26-optimal adjusted statistics for Fisher, Pearson, George, Stouffer, and Edgington; constructs matched surrogates; and returns global 27-values. It supports left-, right-, and two-sided 28-values, provides utilities to compute the null 29-value distribution for binomial, Poisson, negative binomial, (noncentral) hypergeometric, and geometric models as well as user-specified discrete cdfs, and reports 30, 31, and Wald-type diagnostics for method selection. The reported computational profile is 32 precomputation for 33 distinct support points, followed by 34 conversion of observed 35 to 36 and 37 calibration by Gamma or Normal quantiles. The paper recommends the method when 38-values are discrete, independent, and small-39 calibration matters, and it cautions that the theory assumes independence.
The count-data DPComb has a different practical niche. It is designed for conditional count distributions with complex covariate effects, over- or under-dispersion, and the need for noncrossing conditional quantiles. Its main limitations, as stated in the paper, are computational cost from exchange steps, potentially slow mixing for high-dimensional 40, and sensitivity to the kernel choice and bandwidth 41.
The diffusion-model paper “Divide-and-Conquer Posterior Sampling for Denoising Diffusion Priors” explicitly states that the method introduced there is called DCPS, not DPComb (Janati et al., 2024). It further notes that if a query uses “DPComb” for this divide-and-conquer posterior sampling approach, that should be understood only as an alias, with DCPS remaining the canonical name. This suggests that the statistically established meanings of DPComb are the count-data density-regression method and the discrete-42-value combination method, whereas diffusion-based posterior sampling belongs under the separate heading DCPS.