---
title: 'Compound p-values: Methods & Applications'
url: https://www.emergentmind.com/topics/compound-p-values
type: topic
---

# Compound p-values: Methods & Applications

Compound p-values are p-like inferential quantities formed by combining information beyond a single simple test statistic. In the literature, the term is used in several related but nonidentical senses: as p-values that depend on all available data rather than only the data for one hypothesis; as combined p-values obtained by merging several study-level or test-level p-values; and, in a recent multiple-testing formulation, as p-values that satisfy superuniformity only on average across the true nulls rather than coordinatewise. Related work also treats p-values for composite null models and partial-conjunction or replicability hypotheses as part of the same conceptual family [1108.4848] [2409.19812] [2507.21465] [2001.05126].

## 1. Terminological scope

The phrase has no single universally accepted definition. Some papers use it for cross-hypothesis borrowing in multiple testing, some for evidence combination across studies, and some for average-validity relaxations of ordinary p-values. A useful way to organize the literature is by the inferential object being preserved: individual null validity, combined-study evidence, or average null validity.

| Usage | Core condition or construction | Representative source |
|---|---|---|
| Cross-hypothesis borrowing | \(P_k=P_k(X)\) depends on all data but remains a bona fide p-value | [1108.4848], [2409.19812] |
| Combined evidence | Merge \(p_1,\dots,p_K\) into one valid p-value | [1212.4966], [1707.06897] |
| Average-validity multiple testing | \(\sum_{i\in H} P(p_i\le t)\le mt\) | [2507.21465] |
| Composite-null validity | \(P(p\le a)\le a\) uniformly over nuisance values | [2001.05126] |

In the cross-hypothesis sense, compound p-values are ordinary valid p-values that use information from other hypotheses. Habiger-style compound p-values are explicitly described as bona fide p-values that depend on the full data \(X\), rather than only on the \(k\)-th component \(X_k\); the same section of the later e-value literature suggests that “non-separable p-values” is a more precise label for this usage [2409.19812]. In the average-validity sense, by contrast, the defining condition is not per-null superuniformity but the weaker ensemble property
\[
\sum_{i\in H} P(p_i\le t)\le mt,\qquad t\in[0,1],
\]
which generalizes ordinary p-values and is the formal definition of compound p-values in that framework [2507.21465].

A broader interpretation includes p-values for composite nulls with nuisance parameters. In that literature, the relevant question is how to define a valid p-value for \(H_0:X\sim f_0(x;\theta)\) when \(\theta\) is unknown, for example via
\[
p_S=\sup_{\theta\in\Theta} p(\theta),\qquad
p_C=\sup_{\theta\in C_\beta}p(\theta)+\beta,
\]
with validity understood as \(P_{H_0}(p\le a)\le a\) for all \(a\in[0,1]\) [2001.05126].

## 2. Combination rules and merged p-values

A major branch of the subject treats compound p-values as single omnibus p-values obtained by combining several input p-values. Under independence, classical combination rules have distinct power profiles. Fisher’s statistic
\[
X_F=-2\sum_{i=1}^k \log p_i
\]
is sensitive to many moderately small p-values; Stouffer’s inverse-normal sum
\[
Z_S=\frac{1}{\sqrt{k}}\sum_{i=1}^k \Phi^{-1}(1-p_i)
\]
is natural when evidence aggregates linearly on a \(Z\)-scale; Tippett’s method uses \(\min_i p_i\) and is strongest for sparse alternatives; Edgington’s method uses the sum \(\sum_i p_i\) and is less dominated by a single extreme observation. The central comparative result is that no combiner is uniformly best: different rules are optimal against different alternative structures [1707.06897].

When dependence is unknown, valid merging requires explicit calibration. A general arbitrary-dependence theory based on generalized means shows that
\[
2\bar p,\qquad
e\Bigl(\prod_{i=1}^K p_i\Bigr)^{1/K},\qquad
(e\ln K)\Bigl(\frac1K\sum_{i=1}^K p_i^{-1}\Bigr)^{-1}
\]
are valid merged p-values built from the arithmetic, geometric, and harmonic means, respectively. The factor \(2\) for the arithmetic mean is exact and cannot be reduced in general; \(e\) is asymptotically precise for the geometric mean; and \(\ln K\) is the correct asymptotic scale for the harmonic mean, while \(e\ln K\) is a finite-sample safe rule [1212.4966].

The same combination logic appears in partial-conjunction and replicability testing. If \(p_{(1)}\le\cdots\le p_{(s)}\) are ordered p-values and the target null is that at most \(\gamma-1\) component nulls are false, then a valid partial-conjunction p-value is obtained by applying any valid global-null combiner \(g\) to the largest \(s-\gamma+1\) p-values:
\[
P^\star=g\bigl(p_{(\gamma)},\dots,p_{(s)}\bigr).
\]
In that setting, simulations found that Stouffer works well when null p-values are uniform and signal is low, Fisher works better when null p-values are conservative, and the minimum method works well when evidence is concentrated in a few non-null components [2104.13081].

In meta-analysis based on combined p-value functions, the choice of combiner affects not only hypothesis testing but also point and interval estimation. Among the methods systematically compared for one-sided p-value functions \(p_i(\mu)\), only Edgington’s sum-of-p-values method was orientation-invariant, because replacing \(p_i(\mu)\) by \(1-p_i(\mu)\) leaves the Irwin–Hall symmetry intact. The same study emphasized that Edgington’s method can yield confidence intervals that are not constrained to be symmetric around the point estimate [2408.08135].

## 3. Cross-hypothesis borrowing in multiple testing

In large-scale testing, compound p-values often mean p-values that borrow strength across hypotheses while preserving the null properties required by standard multiple-testing procedures. The canonical construction uses a multiple decision process \(\Delta=(\Delta_m)\) and defines
\[
P_{\Delta_m}(X)=\inf\{\eta_m\in[0,1]:\delta_m(X;\eta_m)=1\}.
\]
A p-value statistic is called simple if the \(m\)-th p-value depends only on the \(m\)-th row or test, and compound if it depends on the full data matrix. The key technical device is sample splitting: training data \(Y\) are used to estimate shared structure across tests, and test-specific data \(Z_m\) are used to calibrate the conditional size. If
\[
E_F\bigl(\delta_m(Y,Z_m;\eta_m)\mid Y\bigr)=\eta_m
\]
under the null, then the resulting p-values remain null-uniform; with suitable independence assumptions on the null components, they also remain null-independent [1108.4848].

The location-shift example in that framework illustrates the mechanism. Shared training data estimate a directional weight \(h_m(Y)\), and the compound p-value becomes
\[
P_{\Delta_m^{(c)}}(Y,Z_m)
=
\min\left\{
\frac{\Phi\!\left(Z_m/\sqrt{1-\lambda^2}\right)}{h_m(Y)},
\frac{1-\Phi\!\left(Z_m/\sqrt{1-\lambda^2}\right)}{1-h_m(Y)}
\right\}.
\]
This construction uses the full dataset through \(h_m(Y)\) but still satisfies the conditional null calibration needed for standard FDR procedures [1108.4848].

A more recent empirical-partially-Bayes formulation uses the entire collection of nuisance estimates to construct per-hypothesis p-values. In the Gaussian means model with unknown variances, the oracle partially Bayes p-value is
\[
P_G(z,s^2)=\mathbb P_G\bigl(|Z^{H_0}|\ge |z|\mid S^2=s^2\bigr)
=
\mathbb E_G\bigl[2\{1-\Phi(|z|/\sigma)\}\mid S^2=s^2\bigr].
\]
Replacing the unknown variance prior \(G\) by a nonparametric MLE \(\widehat G\) yields \(P_{\widehat G}(Z_i,S_i^2)\), which is compound because the p-value for hypothesis \(i\) depends on all sample variances through \(\widehat G\). The paper proves nearly parametric approximation rates for these p-values and asymptotic BH/FDR control both in the hierarchical model and in a fixed-variance compound setting [2303.02887].

## 4. Replicability and additive compound p-values

A particularly simple and influential two-study construction uses the sum of p-values as both a combination method and a replicability criterion. With one-sided p-values \(p_o\) from the original study and \(p_r\) from the replication study, define
\[
E=p_o+p_r.
\]
Under the intersection null and independence, \(E\) has an Irwin–Hall distribution with parameter \(2\), and the valid combined p-value is
\[
p_E=
\Pr(\mathrm{IH}(2)\le E)
=
\begin{cases}
E^2/2,&0<E\le 1,\\
-1+2E-E^2/2,&1<E\le 2.
\end{cases}
\]
To match the overall type-I error of the two-trials rule at one-sided level \(\alpha\), replication success is declared when
\[
p_E\le \alpha^2,
\]
which in the practically relevant branch is equivalent to
\[
p_o+p_r\le \sqrt{2}\,\alpha.
\]
At \(\alpha=0.025\), this gives the explicit budget \(p_o+p_r\le 0.035\), allowing a strong replication to rescue a borderline original while preserving the same overall false-positive probability \(0.025^2=0.000625\) under the intersection null [2401.13615].

The same work develops a weighted version
\[
E_w=w_op_o+w_rp_r,
\]
intended to downweight the original study when it is viewed as more vulnerable to questionable research practice, publication bias, or effect inflation. For \(w_o=1\) and \(w_r=2\), the practically relevant success condition becomes
\[
p_o+2p_r\le 2\alpha,
\]
so at \(\alpha=0.025\),
\[
p_o+2p_r\le 0.05.
\]
This admits a less convincing original result than the unweighted rule but requires a stricter replication; indeed, success is impossible if \(p_r>\alpha\). The conditional replication thresholds are
\[
p_r\le \sqrt{2}\alpha-p_o
\]
for the unweighted rule and
\[
p_r\le \alpha-p_o/2
\]
for the \((1,2)\)-weighted rule. The unweighted version can reduce replication sample size when the original study is already very convincing, with reported maximum reductions of \(10.6\%\) or \(9.2\%\) for conditional power targets \(80\%\) and \(90\%\), and \(11.2\%\) or \(10.3\%\) under predictive-power planning; the weighted version always requires a larger replication sample than the two-trials rule [2401.13615].

The same paper positions additive compound p-values against three alternatives. Relative to the two-trials rule, the sum criterion is less dichotomous because it uses a linear budget rather than fixed marginal cutoffs. Relative to Fisher’s method and fixed-effect meta-analysis, it enforces the idea that both studies must contribute evidence, rather than allowing one spectacular p-value to overwhelm a weak or even contradictory replication. In the empirical analyses of four major replication projects, the additive method was mostly similar to the two-trials rule but was modestly more permissive in cases where the original p-value was just above \(0.025\) and the replication p-value was very small [2401.13615].

## 5. Average-validity, p\*-values, and calibration to e-values

A newer multiple-testing meaning of compound p-values relaxes individual validity to average validity across the true nulls. If \(H\subseteq[m]\) is the true-null index set, then \(p_1,\dots,p_m\) are compound p-values when
\[
\sum_{i\in H} P(p_i\le t)\le mt,\qquad t\in[0,1].
\]
Under independence, BH applied to compound p-values no longer has exact nominal control, but it still satisfies
\[
FDR\le 1.93\,\alpha.
\]
This cannot be improved to \(\alpha\) in general: there are independent compound p-values for which \(FDR\ge \tfrac76\alpha\). Under the global null, the upper bound improves to
\[
FDR\le \alpha+2\alpha^2,
\]
with a corresponding lower bound \(\alpha+\alpha^2/4\). Under PRDS-type positive dependence, however, BH can suffer inflation of order \(\log m\), and the paper gives a lower bound of \(\tfrac38\min\{\alpha h_m,1\}\) [2507.21465].

Closely related is the notion of average significance controlling p-values, defined by
\[
\sum_{k\in\mathcal N(Q)}Q(P_k\le t)\le Kt,\qquad t\in[0,1].
\]
That framework presents these p-values as the p-value analogue of compound e-values. It also makes explicit that Habiger-style compound p-values and average-significance-valid p-values are not the same object. In that terminology, p-BH need not control FDR even if the p-values are jointly independent and average significance controlling, whereas p-BY controls FDR under arbitrary dependence, and p-BH controls under weak-dependence asymptotics [2409.19812].

The p\*-value framework further broadens the landscape by introducing an intermediate object between p-values and e-values. A p\*-variable is characterized by the stochastic order \(U\le_2 P\), where \(U\sim \mathrm U[0,1]\). Every p-variable is automatically a p\*-variable; every mid p-value is a p\*-value; and a p\*-variable is exactly a convex combination of p-variables, with any p\*-variable representable as the arithmetic average of three p-variables. This makes p\*-values a natural closure class for averaging and coarsening operations. In particular, if \(P_1,\dots,P_K\) are p\*-variables, then \(\frac1K\sum_k P_k\) is again a p\*-variable, and the universal calibrator
\[
p=(2p^*)\wedge 1
\]
turns a combined p\*-value into an ordinary p-value. The same theory provides a dependence-robust geometric merger
\[
p_{\mathrm{geom}}=\left(e\prod_{k=1}^K P_k^{w_k}\right)\wedge 1
\]
and explicit bridges between p\*-values and e-values [2010.14010].

## 6. Composite-null, discrete, and dependence-sensitive regimes

Compound-p-value methodology becomes more delicate when nulls are composite, test statistics are discrete, or dependence is substantial. For composite null models with nuisance parameters, one classical route is to take the worst-case p-value
\[
p_S=\sup_{\theta\in\Theta}p(\theta)
\]
or the Berger–Boos/Silvapulle refinement
\[
p_C=\sup_{\theta\in C_\beta}p(\theta)+\beta.
\]
These are valid in the sense that \(P_{H_0}(p\le a)\le a\) for all \(a\in[0,1]\), but they may be strongly conservative. In fact, in one-observation Kolmogorov–Smirnov examples with an unknown nuisance parameter, the paper shows \(p_S=1\), and the confidence-set version is also noninformative [2001.05126].

A sharper existence theory for powerful exact composite p-values is available for convex-polytopal null and alternative classes. For finite \(\mathcal P=\{P_1,\dots,P_L\}\) and \(\mathcal Q=\{Q_1,\dots,Q_M\}\), exact nontrivial p-values exist if and only if
\[
\mathrm{Span}(P_1,\dots,P_L)\cap \mathrm{Conv}(Q_1,\dots,Q_M)=\varnothing,
\]
whereas ordinary conservative nontrivial p-values exist if and only if
\[
\mathrm{Conv}(P_1,\dots,P_L)\cap \mathrm{Conv}(Q_1,\dots,Q_M)=\varnothing.
\]
A notable implication is that exact powerful p-values may fail to exist in the full data filtration but may appear after coarsening the data or batching observations [2305.16539].

Discrete inputs create a different problem: the standard continuous null references for Fisher, Pearson, George, Stouffer, and Edgington no longer apply. One recent framework addresses this by replacing each transformed discrete p-value with a Wasserstein-optimal adjusted value and approximating the null distribution of the sum by a continuous surrogate with matching first two moments. For Edgington’s statistic, the optimal adjusted component is the midpoint
\[
z_i=\frac{F_i+F_{i-1}}{2},
\]
which coincides with the mid-p-value, and the paper emphasizes the variance ratio \(\mathrm{Var}(Z)/\mathrm{Var}(Y)\) as a practical selector of which combination statistic is likely to be best calibrated for a given discrete null distribution [2508.02647].

A complementary discrete-data literature studies randomized p-values for composite nulls. In one-parameter exponential-family settings, the least-favorable-configuration p-value \(P^{LFC}\), a single-stage randomized version \(P^{rand1}\), a first-stage exact randomized tail probability \(P_T^{rand}\), and a two-stage randomized version \(P^{rand2}\) are all valid. The two-stage procedure is designed to address both discreteness and the extra conservativeness induced by interior null values, and in the binomial examples it is closest to \(\mathrm{Unif}(0,1)\) under the null. The same paper reports that powers based on \(P_T^{rand}\) and \(P^{rand2}\) increase monotonically with sample size, unlike the non-randomized least-favorable p-value [2208.06342].

Dependence can also alter the meaning of compound p-values. Heavy-tailed combination tests such as the Cauchy combination and harmonic-mean-style rules are asymptotically valid for fixed \(n\) and \(\alpha\to 0\) under pairwise Gaussian dependence without perfect correlations, but in that regime they become asymptotically equivalent to weighted Bonferroni whenever the transformed variables are pairwise quasi-asymptotically independent. Under \(t\)-like tail dependence, however, simulations suggest that the same heavy-tailed combinations can remain well calibrated and can substantially outperform Bonferroni [2310.20460].

The literature therefore treats compound p-values less as a single formula than as a family of constructions for preserving p-value-like calibration while borrowing information, combining evidence, or relaxing validity in controlled ways. The unifying theme is that the inferential burden is shifted from a single-test null law to a richer structure: across studies, across hypotheses, across nuisance values, or across nulls on average.

Source: https://www.emergentmind.com/topics/compound-p-values