---
title: Post-Hoc Hypothesis Testing
url: https://www.emergentmind.com/topics/post-hoc-hypothesis-testing
type: topic
---

# Post-Hoc Hypothesis Testing

Post-hoc hypothesis testing denotes inferential settings in which some component that is classically fixed before analysis—most prominently the significance level $\alpha$, but also a rejection set, a variable set, a pathway collection, or a local region of interest—is chosen after inspection of the data. In classical Neyman–Pearson testing this is ordinarily invalid, because size control relies on choosing $\alpha$ independently of the data. Contemporary work therefore treats post-hoc testing in two main ways: by deriving simultaneous guarantees that hold for all admissible post-data choices, or by replacing fixed-$\alpha$ tail-probability control with e-value-based risk control that remains valid under data-dependent choices of $\alpha$ [2410.02306], [1208.2841], [2205.00901].

## 1. Meanings and problem formulations

In the cited literature, the expression *post-hoc hypothesis testing* refers to several distinct but related inferential problems. One concerns choosing $\alpha$ after observing the data, which violates the standard assumption that $\alpha$ is fixed independently of the realized $p$-value. A second concerns choosing the rejected subset $R$ or $S$ after seeing the data, while still wanting a valid statement about the number of false rejections. A third, older usage refers to pairwise or localized follow-up procedures applied after rejection of a global null. These uses are not identical, and much of the modern literature is concerned with separating them conceptually while restoring formal guarantees in each setting [1703.02307], [1505.02288], [2312.17566].

| Formulation | Object chosen after seeing data | Guarantee |
|---|---|---|
| Data-dependent $\alpha$ | Significance level | Type-I risk control via e-values or post-hoc $p$-values |
| User-agnostic multiple testing | Rejection set $R$ or $S$ | Simultaneous bound on false rejections or false positives |
| Classical post-hoc comparison | Pairwise contrasts after omnibus rejection | Family-wise control under a model-specific procedure |

The first formulation is central to "Choosing alpha post hoc: the danger of multiple standard significance thresholds" [2410.02306]. The second is central to the “reversed roles” approach of Goeman and Solari, the JER framework of Blanchard et al., the forest-structured spatial bounds of Durand et al., and exact closed testing for Globaltest pathway analysis [1208.2841], [1703.02307], [1807.01470], [2001.01541]. The third includes both critiques of invalid follow-up procedures and construction of valid ones, as in mean-ranks comparisons after Friedman’s test and Tukey-style post-hoc comparison of Sharpe ratios [1505.02288], [1911.04090].

## 2. Data-dependent significance levels and size inflation

In classical Neyman–Pearson testing, one specifies before seeing the data a significance level $\alpha$. A test $\phi_\alpha(X)$ rejects the null $H_0$ exactly when the data $X$ fall into a pre-defined rejection region of size $\alpha$, so that under $H_0$,
$$
P_0[\phi_\alpha(X)=1]\le \alpha.
$$
The same pre-specification requirement underlies the nominal coverage of confidence intervals obtained by inverting such tests. If $\alpha$ is allowed to depend on the observed data, written $\alpha(X)$, the guarantee no longer has its usual interpretation [2410.02306].

The 2024 analysis of multiple standard thresholds formalizes this distortion by considering a set $A\subseteq(0,1]$ of possible significance levels and defining, for each $a\in A$, the conditional discrepancy
$$
d_a := P_0[\phi_a(X)=1 \mid \alpha(X)=a] - a,
$$
and the relative discrepancy
$$
r_a := P_0[\phi_a=1 \mid \alpha=a]/a.
$$
If $r_a>1$, then conditional on $\alpha=a$ the test rejects too often. The overall expected discrepancy ratio is
$$
E[r_{\alpha(X)}] = E[\phi_{\alpha(X)}(X)/\alpha(X)].
$$
This quantity can exceed $1$ even when each fixed-$\alpha$ test is valid [2410.02306].

Two examples make the point explicit. With two candidate thresholds $a_1<a_2$, suppose the rule is to set $\alpha=a_1$ if $p\le a_1$ and $\alpha=a_2$ if $p>a_1$. Then conditional on $\alpha=a_1$ one must have $p\le a_1$, hence
$$
P_0[\phi=1\mid \alpha=a_1]=1,
$$
so $r_{a_1}=1/a_1\gg 1$. By contrast, conditional on $\alpha=a_2$ the procedure is conservative, yet the overall expected ratio is
$$
E[r_\alpha]=1+(a_2-a_1)/a_2 >1.
$$
For $a_1=0.005$ and $a_2=0.05$, this gives $E[r_\alpha]=1.9$, so on average one rejects $90\%$ more often than intended. In a second example, if one sets $\alpha=\min\{p,C\}$ over infinitely many candidate thresholds, then the expected ratio is infinite, and the test is “disastrously invalid” [2410.02306].

This analysis is directly relevant to proposals to replace a single field-wide threshold by multiple “accepted” thresholds such as $0.05$ and $0.005$. If different journals or venues accept different thresholds, researchers with $p\approx 0.02$ may “shop” for venues accepting $0.05$, whereas those with $p\approx 0.004$ may prefer venues demanding $0.005$. The underlying problem is not merely sociological: the post-hoc choice of $\alpha$ violates the independence assumption, inflates Type I error, and breaks the nominal interpretation of both tests and confidence intervals [2410.02306].

## 3. E-values, post-hoc $p$-values, and decision-theoretic admissibility

The modern solution to post-hoc $\alpha$ choice is based on e-values. An e-value is a nonnegative statistic $S$ such that under every null distribution,
$$
E[S]\le 1.
$$
By Markov’s inequality, the rule “reject if $S\ge 1/\alpha$” has Type I error at most $\alpha$ for every fixed $\alpha$. More importantly, e-values support a generalized Neyman–Pearson formulation in which the downstream decision task, or loss function, may itself be chosen after observation of the data. In that setting, sufficiently rich decision problems have only e-value-based admissible rules [2205.00901].

This perspective yields a precise characterization of post-hoc $p$-values. A statistic $p$ is a post-hoc $p$-value if and only if
$$
E_0[1/p]\le 1,
$$
that is, if and only if $1/p$ is an e-value. Unlike regular $p$-values, such post-hoc $p$-values allow one to “reject at level $p$” while retaining the relevant expectation guarantee. They also combine multiplicatively under independence: if $p_1$ and $p_2$ are independent post-hoc $p$-values, then $p_1p_2$ is again a post-hoc $p$-value because
$$
E\!\left[\frac{1}{p_1p_2}\right]
=
E\!\left[\frac{1}{p_1}\right]
E\!\left[\frac{1}{p_2}\right]
\le 1.
$$
This equivalence also clarifies the role of e-values as the reciprocal evidence measure underlying post-hoc testing [2312.08040].

The decision-theoretic formulation goes further. For a loss budget $\ell>0$, a rule $\delta$ is Type I-risk safe if
$$
\sup_{P\in H_0} E_P[L(\text{null},\delta(Y))]\le \ell.
$$
The generalized Neyman–Pearson theorem states that if $\delta$ is admissible, then there exists an e-variable $S$ such that
$$
L(\text{null},\delta(y))\le \ell\,S(y)
\quad \forall y.
$$
Conversely, maximally compatible rules based on sharp e-variables are admissible. In this sense, e-value-based rules form a complete class for post-hoc decision tasks [2205.00901].

The 2025 theory of $\Gamma$-admissibility sharpens this by modeling an adversary that maps the data to a significance level. If $\Gamma$ contains all constant mappings and $\delta$ is $\Gamma$-admissible, then
$$
\delta(X,b)=\min\{1,E_\delta(X)/L_b(0,1)\}
$$
almost surely for every $b$. In the binary setting this reduces to
$$
\delta(X,b)=1\{E_\delta(X)\ge L_b(0,1)\}.
$$
When $\Gamma$ contains only constant adversaries and the loss span collapses to a singleton, this recovers the classical Neyman–Pearson likelihood-ratio test; when $\Gamma$ allows arbitrary data-dependent adversaries, admissible tests are exactly the canonical tests arising from sharp e-variables [2508.00770].

## 4. Simultaneous post-hoc inference in multiple testing

A second major strand does not try to salvage ordinary fixed-$\alpha$ tests under data-driven threshold choice; instead, it returns a statement that is valid simultaneously for every post-hoc chosen rejection set. In “Multiple Testing for Exploratory Research,” the conventional roles are reversed: the analyst chooses any candidate rejection set $R\subseteq\{1,\dots,n\}$ after seeing the data, and the procedure returns an integer $t_\alpha(R)$ such that, with confidence at least $1-\alpha$,
$$
0 \le \tau(R)\le t_\alpha(R),
$$
where $\tau(R)$ is the number of false rejections in $R$. Equivalently, if $\phi(R)=|R|-\tau(R)$ is the number of correct discoveries, then
$$
\phi(R)\ge |R|-t_\alpha(R)
$$
simultaneously for all $R$. The machinery is classical closed testing, but the crucial insight is that non-consonant rejections, often treated as a nuisance in ordinary FWER theory, are informative because they shrink the upper bound on the number of false rejections [1208.2841].

Blanchard et al. generalized this idea in large-scale multiple testing via the joint-family-wise-error rate (JER). For a sequence of thresholds $\{t_k\}_{k=1}^K$, JER is
$$
\mathrm{JER}(\{t_k\})
=
\Pr\Bigl(\exists\,k\le K\wedge m_0:\;p_{(k:\mathcal H_0)}<t_k\Bigr).
$$
If JER is controlled at level $\alpha$, then
$$
V(R)=\min_{1\le k\le K}\Bigl\{|\{i\in R:p_i\ge t_k\}|+(k-1)\Bigr\}
$$
is a simultaneous post-hoc upper bound on the number of false positives in any user-chosen set $R$, and $S(R):=|R|-V(R)$ is a lower-confidence bound on the number of true discoveries. The framework accommodates both known dependence and permutation-based calibration, and step-down calibration adapts to the unknown quantity of signal [1703.02307].

Durand et al. specialized this perspective to spatially structured hypotheses. For a reference family $\{(R_k,\zeta_k)\}_{k\in K}$ satisfying a joint error-rate control and a forest-structure condition, they construct a data-dependent upper bound $V(X,S)$ such that
$$
\Pr\Bigl[\forall S\subset N:\ \FP(S)\le V(X,S)\Bigr]\ge 1-\alpha.
$$
The coverage is simultaneous over all $2^m$ subsets $S$, so the subset may be chosen arbitrarily, even after repeated inspection of the data. The forest structure yields an explicit interpolation formula and a low-complexity dynamic-programming algorithm; the implementation is available in the R package **sansSouci** [1807.01470].

Exact post-hoc multiple testing has also been developed for metabolomics pathway analysis. In pathway testing with Globaltest, the family $\{H_R:R\subseteq F\}$ is closed under unions, and closed testing therefore controls FWER simultaneously over all possible feature sets. Xu et al. derive a shortcut, based on convex-hull envelopes for the minimum test statistic and majorization envelopes for the maximum critical value, that makes exact closed testing computationally feasible at metabolomics scale. The result is that one may choose the pathway database after seeing the data without jeopardizing error control; the implementation is provided in the R package **ctgt** [2001.01541].

## 5. Pairwise post-hoc procedures after omnibus tests

In a narrower and older usage, *post-hoc tests* are follow-up comparisons performed after a global null has been rejected. This usage is common in algorithm comparison, medicine, psychology, finance, and other applied fields. The modern literature emphasizes that such procedures are not automatically valid merely because they are labeled “post-hoc”: the validity depends on whether the pairwise decision for a given pair depends only on that pair, and whether the multiplicity correction matches the inferential target [1505.02288], [1911.04090].

Benavoli et al. show that the mean-ranks post-hoc test used after Friedman’s test is inconsistent because the outcome of the comparison between algorithms $A$ and $B$ depends on the performance of the other algorithms included in the original experiment. In their example with $n=20$ data sets, algorithms $A$ and $B$ are tied in every direct comparison, so when considered alone they have $\Delta_{AB}=0$ and $p=1$. After adding three auxiliary algorithms $C,D,E$, the same pair yields $\bar R_A=2.0$, $\bar R_B=3.5$, $\Delta_{AB}=1.5$, and
$$
z_{AB}=3.0 \Rightarrow \text{ two-sided } p\approx 0.003,
$$
so $A$ versus $B$ is declared significant at $\alpha=0.05$ with Bonferroni correction for $m=5$. By choosing a different triple of auxiliary algorithms, one can reverse the conclusion again. The recommended alternatives are tests whose outcome depends only on the paired differences, such as the sign test or the Wilcoxon signed-rank test, combined with multiplicity correction [1505.02288].

A contrasting example is Pav’s post hoc test on the Sharpe ratio. After rejection of the global null
$$
H_0:\theta_1=\theta_2=\cdots=\theta_q,
$$
under a Gaussian equi-correlation model for contemporaneous returns, the pairwise difference in sample Sharpe ratios can be compared using a Tukey-style studentized-range cutoff. The honest significant difference is
$$
\mathrm{HSD}=r_\alpha\sqrt{(1-\rho)/n},
$$
where $r_\alpha=q_{tukey}(1-\alpha;k,df)$ with $k=q$ and $df=\infty$ or $df=n-1$. The family-wise guarantee follows from the range distribution, and simulation results indicate that the $df=n-1$ variant achieves nearly correct $\alpha$ over a grid of $n$ and $q$, whereas the $df=\infty$ rule can be anti-conservative when $q$ is large and $n$ is small [1911.04090].

These examples illustrate a persistent misconception. Rejection of a global null does not by itself justify arbitrary follow-up comparisons. Valid post-hoc pairwise inference requires either a pairwise test whose sampling law depends only on the two items being compared, or a joint reference distribution that genuinely controls family-wise error for the entire family of contrasts. The mean-ranks critique and the Sharpe-ratio construction occupy opposite sides of this distinction [1505.02288], [1911.04090].

## 6. Confidence sets, asymptotic theory, and emerging extensions

Because ordinary confidence intervals are obtained by inverting fixed-level tests, post-hoc choice of $\alpha$ invalidates their nominal coverage. The same point appears in both the classical critique of choosing $\alpha$ after seeing the data and the decision-theoretic e-value literature: to ensure advertised $1-\alpha$ coverage, $\alpha$ must be fixed in advance, or one must use e-confidence sets, e-posteriors, or other constructions that remain valid under data-dependent $\alpha$ [2410.02306], [2205.00901].

A 2026 extension develops post-hoc large-sample statistical inference through *asymptotic e-values*. For the mean, one example is the IWR e-variable
$$
E_n^{(\mu;\lambda)}
=
\exp\!\left\{\lambda\,[S_n(\mu)/V_n(\mu)]-\lambda^2/2\right\},
$$
which is an asymptotic e-value under only finite-second-moment or domain-of-attraction conditions. Inversion yields asymptotic post-hoc confidence intervals and asymptotic post-hoc $p$-values. The framework also treats mixture constructions over $\lambda$, and extends to post-hoc confidence sequences via asymptotic e-processes [2603.08002].

The same period has seen post-hoc validity imported into other inferential paradigms. “Doublethink” shows that Bayesian model-averaged hypothesis testing is a closed testing procedure that controls the frequentist familywise error rate in the strong sense. It computes simultaneous posterior odds and asymptotic $p$-values, and permits post-hoc variable selection because the relevant intersection nulls are already tested; the paper also emphasizes finite-sample inflation and the need to mitigate highly correlated covariates, for example by testing groups of correlated variables [2312.17566]. In equivalence testing, post-hoc margin selection is handled by reporting a data-dependent bound $\widehat\Delta_\alpha$ and, more generally, a uniformly valid curve $\alpha\mapsto \widehat\Delta_\alpha$ obtained by inverting a non-decreasing family of e-values [2603.16213].

A further development is local post-hoc detection rather than global rejection. For dependence testing, Ćmiel and Gibas replace a single critical threshold by critical surfaces $l_\alpha(u,v)$ and $u_\alpha(u,v)$ for the quantile dependence function, so that one can identify where dependence occurs while preserving the overall significance level of the test. By Bonferroni, with local level $\alpha/(2n^2)$ on a $k\times k$ grid and $k\le n$, the global probability of any false local detection is at most $\alpha$ [2512.20280].

Taken together, these developments suggest that post-hoc hypothesis testing is no longer a single technical warning about “peeking” at $p$-values. It has become a general theory of inference under analyst adaptivity, with at least two non-equivalent forms of validity: simultaneous setwise guarantees in multiple testing, and e-value-based control of post-hoc risk under data-dependent $\alpha$. The main unresolved point is not whether unrestricted post-hoc choice is harmless—it is not for ordinary $p$-value testing—but which replacement guarantee is scientifically appropriate for a given inferential task [2410.02306], [2205.00901], [2603.08002].

Source: https://www.emergentmind.com/topics/post-hoc-hypothesis-testing