---
title: 'Maximally Adjusted P-Value: Concepts & Applications'
url: https://www.emergentmind.com/topics/maximally-adjusted-p-value
type: topic
---

# Maximally Adjusted P-Value: Concepts & Applications

A maximally adjusted p-value is a p-value obtained by calibrating inference against the most challenging admissible configuration induced by multiplicity, selection, nuisance variation, or sequential monitoring. In the literature, the construction is context dependent rather than universal. In some settings it is literally the largest p-value over a set of candidate analyses or nuisance-parameter values; in others it is the p-value attached to an extremal statistic such as the minimum raw p-value across analyses, the maximum sequential p-value across intersection hypotheses, or the running extremum under optional stopping. The common objective is conservative but valid inference under a larger decision space than a single prespecified test, especially when nominal p-values would otherwise be anti-conservative or selectively reported [2401.11537] [2311.15498] [2507.06451] [1610.00378] [2011.01343].

## 1. Core idea and major formalizations

The literature suggests that “maximally adjusted p-value” is best understood as a family of worst-case adjustments rather than a single standardized statistic. Across the cited settings, the adjustment takes one of three recurrent forms: maximization over candidate p-values, maximization over closed-testing components, or calibration against the null distribution of an extremum. This suggests a unifying principle: the adjusted p-value is designed to remain valid after accounting for a space of admissible analytical choices that is larger than the single analysis ultimately reported.

| Context | Defining adjustment | Stated error property |
|---|---|---|
| Researcher degrees of freedom [2401.11537] | permutation tail probability of the minimum raw p-value | weak FWER under the global null |
| Group sequential multiple testing [2311.15498] | \(p^{aseq}_{i,k}=\max_{J\ni i} p^{seq}_{J,k}\), then \(p^*=\max_{i,k} p^{aseq}_{i,k}\) | strong FWER |
| Paired-control vaccine assays [2507.06451] | \(p_{\max}^{\rm adj}=\sup_{\theta\in\mathcal A(\alpha')} p_\theta+\alpha'\) | asymptotically valid Type I control |
| PC-Max causal discovery [1610.00378] | \(p^*=\max_i p_i\) over candidate conditioning sets | consistency under standard PC assumptions |
| Optional stopping [2011.01343] | \(p_{\rm adj}=\min_{1\le n\le N} H_n\) for an always-valid p-value process | Type I control under arbitrary stopping |
| Functional hypotheses [1912.00360] | pointwise \(p_{\rm adj}(x)\) based on null ranks of domainwise minima | FWER over the function domain |

These definitions are not interchangeable. In the group sequential and vaccine-assay settings, the maximization is explicit in the formula. In researcher-degrees-of-freedom problems, the paper instead uses the phrase “max-adjusted (minP) p-value” for the tail probability of the smallest observed p-value under a permutation-estimated global-null distribution. In optional-stopping work, the final adjusted quantity is a minimum over always-valid p-values, but it is derived from control of the running maximum of a supermartingale. The terminology therefore depends on which object is being maximized or calibrated.

## 2. Researcher degrees of freedom and the minP construction

Mandl et al. formalize researcher degrees of freedom as a multiple-testing problem in which \(K\) analysis strategies are applied to the same dataset, producing raw p-values \(p_1,\dots,p_K\). The observed extremum is

$$
p_{\min}^{\mathrm{obs}}=\min_{1\le k\le K} p_k.
$$

Because the test statistics from alternative analysis strategies are usually highly dependent, the paper argues that Bonferroni is inappropriate owing to an unacceptable loss of power. The proposed remedy is a permutation-based minP adjustment that approximates the null distribution of the minimum p-value while preserving dependence among the candidate analyses [2401.11537].

For permutations \(b=1,\dots,B\), one constructs permuted datasets \(D^{(b)}\) that enforce the global null, recomputes all \(K\) analyses, and records

$$
p_{\min}^{(b)}=\min_{1\le k\le K} p_k^{(b)}.
$$

The adjusted p-value is then

$$
p_{\rm adj}
=
\frac{1}{B}\sum_{b=1}^B
I\bigl(p_{\min}^{(b)}\le p_{\min}^{\mathrm{obs}}\bigr).
$$

By construction, this estimates the tail probability of observing a minimum p-value at least as small as \(p_{\min}^{\mathrm{obs}}\) under the global null. The method provides weak control of the family-wise error rate, with the paper emphasizing that weak control is sufficient for the specific aim of guarding against the case in which all admissible analytical choices correspond to null effects.

The practical rationale is dependence adaptation. Bonferroni replaces each \(p_k\) by \(\tilde p_k=\min(Kp_k,1)\), which is valid under arbitrary dependence but conservative when tests are positively correlated. The permutation-minP method instead estimates the true null distribution of \(\min_k p_k\), and the paper states that this yields a smaller correction and boosts power when analysis strategies are highly correlated.

The empirical illustration concerns perioperative \(paO_2\) and post-operative complications after neurosurgery. The study included \(n=3{,}163\) patients, a binary outcome, and \(K=48\) analysis strategies spanning missing-data handling, surrogate modeling for unobserved \(paO_2\), aggregation of repeated measures, and alternative codings or tests of \(paO_2\). With \(B=1{,}000\) permutations obtained by shuffling the binary outcome, unadjusted testing produced an approximately \(70\)–\(75\%\) chance of at least one false positive among 48 p-values at \(\alpha=0.05\), Bonferroni kept FWER below \(0.05\) but was overly conservative for moderate \(n\), and minP held FWER at approximately \(0.05\) for all sample sizes. Under the real association, subsampling experiments with \(n=50\)–\(500\) and \(1{,}000\) repetitions showed that minP yielded substantially more significant findings than Bonferroni at \(\alpha=1\%,5\%,10\%\), while controlling FWER.

## 3. Closed testing and group sequential multiplicity

In group sequential designs with \(m\) elementary hypotheses and \(K\) looks, Zhao et al. define an adjusted-sequential p-value by combining multiple-hypothesis adjustment with repeated-analysis adjustment through closed testing. For an intersection hypothesis \(H_J=\cap_{j\in J} H_j\), the sequential p-value at look \(k\) is

$$
p^{seq}_{J,k}
=
\inf\Bigl\{\mu\in(0,1):
\bigcup_{k'=1}^k \bigcup_{j\in J}
\{Z_{j,k'}\ge \widetilde Z_{j,k'}(\mu,J)\}
\Bigr\},
$$

or equivalently

$$
p^{seq}_{J,k}
=
\sup\Bigl\{\mu\in(0,1):
\max_{1\le k'\le k,\; j\in J}
\bigl[Z_{j,k'}-\widetilde Z_{j,k'}(\mu,J)\bigr]\le 0
\Bigr\}.
$$

The adjusted-sequential p-value for an elementary hypothesis \(H_i\) is

$$
p^{aseq}_{i,k}
=
\max_{J:\, i\in J} p^{seq}_{J,k}.
$$

This is the closed-testing maximum over all intersection hypotheses containing \(i\). The paper then defines a global maximally adjusted p-value,

$$
p^*=\max_{1\le i\le m,\;1\le k\le K} p^{aseq}_{i,k},
$$

described as the smallest \(\mu\) for which at least one \(H_i\) would be rejected at or before its look. The stated decision rule is that at least one hypothesis is rejected if and only if \(p^*\le \alpha\), and \(H_i\) is rejected at the earliest \(k\) for which \(p^{aseq}_{i,k}\le \alpha\). By the closed-testing principle, this yields strong FWER control [2311.15498].

The construction is algorithmic. One chooses a multiple-testing method, such as weighted Bonferroni with graph-recycling or weighted parametric group sequential design, specifies an \(\alpha\)-spending function such as Hwang–Shih–DeCani, and defines a weighting graph. For each nonempty intersection \(J\) and each look \(k\), one computes local \(\alpha\)-allocations and corresponding boundaries \(\widetilde Z_{j,k'}(\mu,J)\), then solves for the root in \(\mu\) at which the rejection event “just becomes true.” One-dimensional root-finding in \(\mu\) is sufficient because the unknown is scalar.

The worked example has \(m=3\) populations, \(K=2\) looks, weights \((0.3,0.3,0.4)\), and HSD\((\gamma=-4)\) spending. Under weighted Bonferroni, the paper reports \(p^{seq}_{\{1,2,3\},1}=0.2097\) at interim and \(p^{seq}_{\{1,2,3\},2}=0.0266\) at final. The resulting elementary adjusted-sequential p-values are all \(0.2097\) at interim and all \(0.0266\) at final, so

$$
p^*=\max\{0.2097,0.2097,0.2097,0.0266,0.0266,0.0266\}=0.2097.
$$

Because \(p^*=0.2097>\alpha=0.025\), none of the three hypotheses is significant at either look. The example illustrates the intended role of \(p^*\) as a single scalar summary for an entire multi-hypothesis, multi-look trial.

## 4. Worst-case nuisance adjustment in paired-control vaccine assays

Chen et al. introduce a maximally adjusted p-value for intracellular cytokine staining data observed before and after vaccination, with paired control samples used to account for assay variation across time points. The motivating problem is differential misclassification between assay runs: false-positive and false-negative rates may differ between baseline and post-vaccination measurements, and naive comparison of estimated positive-cell proportions can therefore inflate false positives [2507.06451].

For each participant, the framework uses a primary baseline sample \((n_0,N_0)\), a primary post-vaccination sample \((n_1,N_1)\), and paired control samples \((n_0',N_0')\) and \((n_1',N_1')\). Let

$$
\boldsymbol\theta=
(p_{T_0,1\mid 0},\,p_{T_0,0\mid 1},\,p_{T_1,1\mid 0},\,p_{T_1,0\mid 1})
$$

denote the four misclassification rates. From the paired control data one constructs a \(100(1-\alpha')\%\) confidence set

$$
\mathcal A(\alpha')
=
\bigl\{\boldsymbol\theta\in[0,1]^4:
|\bar T^C(\boldsymbol\theta)|\le z_{1-\alpha'/2}
\bigr\},
$$

where \(\bar T^C(\theta)\) is the standardized difference in control proportions at \(T_1\) versus \(T_0\). For a fixed \(\theta\), the primary-sample one-sided p-value for testing \(H_0:p_{T_1,1}=p_{T_0,1}\) against \(H_1:p_{T_1,1}>p_{T_0,1}\) is

$$
p_{\boldsymbol\theta}=1-\Phi\bigl(\bar T(\boldsymbol\theta)\bigr).
$$

The maximally adjusted p-value is then

$$
p_{\max}^{\rm adj}
=
\sup_{\boldsymbol\theta\in\mathcal A(\alpha')}
p_{\boldsymbol\theta}
+\alpha'.
$$

The additive \(\alpha'\) term is the Berger–Boos correction. The interpretation given in the paper is explicitly worst case: one takes the largest p-value over all misclassification-rate scenarios consistent with the paired control data. The method therefore preserves Type I error no matter which compatible batch-effect scenario is true.

Its validity depends on several assumptions stated in the paper: within each assay run, all cells share the same misclassification rates; primary and control samples in that run share these rates; the true positive-cell proportion in the control is the same at \(T_0\) and \(T_1\); \(\bar T(\theta)\) and \(\bar T^C(\theta)\) are approximately standard normal under large-sample asymptotics; and \(\mathcal A(\alpha')\) is effectively compact. Under these conditions, Theorem 3.1 gives asymptotic validity:

$$
\limsup_{N'\to\infty}
P\bigl(p_{\max}^{\rm adj}\le \alpha\mid H_0\bigr)
\le \alpha.
$$

The paper contrasts \(p_{\max}^{\rm adj}\) with a minimally adjusted p-value \(p_{\min}^{\rm adj}\), defined as the member of \(\mathcal P_{\mathcal A(\alpha)}\) closest to the naive unadjusted p-value \(p^*\). By construction \(p_{\min}^{\rm adj}\le p_{\max}^{\rm adj}\), but only the maximally adjusted version is described as always conservative enough to hold Type I error at the nominal level.

The practical consequence is a responder classification criterion that is robust to all control-data-compatible batch effects. The paper explicitly recommends \(p_{\max}^{\rm adj}\) in regulatory or other high-stakes settings where Type I control is paramount, while also noting that weak control-data information can make the confidence set \(\mathcal A(\alpha')\) large and the resulting adjustment very conservative. In the CoVPN 3008 case study, some participants with unadjusted \(p\approx 3\times10^{-4}\) had maximally adjusted p-values of \(8.9\times10^{-3}\), reflecting strong batch-effect signals in the paired controls.

## 5. Algorithmic and sequential uses beyond classical multiple testing

In causal structure learning, the maximally adjusted p-value appears as a selection rule inside the PC-Max algorithm. For a pair of variables \(X\) and \(Y\), one enumerates candidate conditioning sets \(S_1,\dots,S_k\), performs conditional-independence tests for \(H_0:X\perp Y\mid S_i\), and records p-values \(p_i\). The algorithm defines

$$
S^*=\arg\max_{1\le i\le k} p_i,
\qquad
p^*=\max_{1\le i\le k} p_i.
$$

During the adjacency phase, at conditioning-set size \(\ell\),

$$
p^*_{X,Y}
=
\max_{S\subseteq adj_G(X)\cup adj_G(Y),\,|S|=\ell}
p\bigl(X,Y\mid S\bigr),
$$

and the edge \(X-Y\) is removed if \(p^*_{X,Y}>\alpha\). The maximizing separating set is stored and later used in collider orientation. The algorithm also ranks candidate colliders in descending order of \(p^*\), orients them subject to a no-bidirected-edge constraint, and then applies Meek’s rules [1610.00378].

This usage is operational rather than directly inferential: the largest p-value is treated as the least evidence against conditional independence among admissible separating sets. Under causal sufficiency, a DAG data-generating process satisfying Markov and faithfulness, and a consistent conditional-independence test, the paper states a consistency theorem: as \(n\to\infty\), PC-Max recovers the true CP-DAG with probability tending to one. In simulations with \(p=1000\), average degree \(d\in\{2,4\}\), \(n=1000\), Fisher \(Z\) testing, and \(\alpha=0.001\), PC-Max achieved adjacency precision \(0.98\), adjacency recall \(0.96\), arrowhead precision \(0.98\), arrowhead recall \(0.91\), bidirected edges \(0\%\), and runtime \(1.3\) seconds on a 4-core laptop. At \(p=20{,}000\) and \(\alpha=10^{-5}\), it retained adjacency precision \(1.00\), adjacency recall \(0.94\), arrowhead precision \(0.99\), arrowhead recall \(0.87\), and runtime \(226\) seconds.

A different sequential use appears in work on p-value peeking. Here the goal is to make testing robust to arbitrary stopping rules by working with a nonnegative supermartingale \(M_n\) under \(H_0\) and its always-valid p-value process \(H_n=1/M_n\). The fully adjusted p-value reported after observing up to \(N\) samples is

$$
p_{\rm adj}
=
\min_{1\le n\le N} H_n
=
\min_{1\le n\le N}\frac{1}{M_n}.
$$

Although this quantity is a minimum rather than a maximum, its validity comes from control of the running maximum of \(M_n\) via Ville’s inequality,

$$
\Pr_0\Bigl(\max_{n\le N} M_n \ge 1/\alpha\Bigr)\le \alpha.
$$

The paper therefore describes the resulting procedure as a maximally peeking-robust test, and the final \(p_{\rm adj}\) can be interpreted as the p-value adjusted against the most extreme excursion induced by arbitrary peeking [2011.01343].

The implementation can be based on mixture martingales \(M_n=\int \exp(\lambda T_n-\psi_n(\lambda))\,d\Gamma(\lambda)\), updated online with \(O(K)\) computation per new sample when \(\Gamma\) is discretized over \(K\) components. In a one-sided Gaussian mean-test simulation with \(10^5\) repeated experiments and horizon \(N=200\), the empirical rejection rate at level \(0.05\) was \(4.8\%\) for the adjusted \(p_{\rm adj}\), versus over \(25\%\) for the naive peeked p-value.

## 6. Related extremal adjustments, conceptual boundaries, and limitations

Several adjacent literatures rely on the same extremal logic even when they do not use the exact phrase “maximally adjusted p-value.” In functional hypothesis testing, Xu and Reiss define pointwise adjusted p-values by ranking the observed statistic function against simulated null functions and calibrating against the minima of rank processes across the function domain. The single-step adjusted p-value is

$$
p_{\rm adj}(x)
=
\frac{1+\#\{\,b: R_b\le R_0^*(x)\,\}}{B+1},
$$

where \(R_b=\min_{x\in\mathcal X} R_b^*(x)\). This controls FWER over the domain, and step-down and extreme-rank-length refinements can yield lower adjusted p-values while maintaining FWER control [1912.00360].

In survival forests based on maximally selected rank statistics, the inferential problem is different but closely related. For a continuous covariate, one computes

$$
T_{\max}=\max_{\tau\in\mathcal M} |T(\tau)|,
$$

over admissible cutpoints \(\tau\). The naive p-value \(2[1-\Phi(T_{\max})]\) is described as far too liberal because the maximum is taken over many highly correlated statistics. Wright et al. therefore discuss several adjusted approximations, including the Brownian-bridge approximation \(P_1\), the improved Bonferroni approximation \(P_2\), and the conservative rule \(p_{\min}=\min\{P_1,P_2\}\). Their simulations show a trade-off between unbiased split-variable selection and runtime, with \(p_{\min}\) recommended as a practical compromise [1605.03391].

For discrete p-values, optimal-transport methods provide yet another related but distinct notion of adjustment. One line of work formulates an adjusted discrete statistic as the \(W_2\)-optimal projection of a discrete null onto a continuous surrogate, recovering Lancaster’s mid-\(p\) and mean-value \(\chi^2\) constructions as special cases, and then replacing the usual \(\chi^2\) null by a moment-matched Gamma distribution to reduce conservativeness [2309.07692]. A later extension generalizes this framework to Fisher, Pearson, George, Stouffer, and Edgington combinations, matching means and variances within surrogate families such as Gamma or Normal, with asymptotic level accuracy and power close to the corresponding continuous procedures when the likelihood-ratio statistic is monotonic in the combination statistic [2508.02647]. These are p-value adjustments, but not “maximally adjusted” in the sense of maximizing over admissible analyses or nuisance configurations.

Two recurrent limitations run through the literature. First, the strength of error control varies sharply by construction. Mandl et al. emphasize only weak FWER control under the global null for permutation minP [2401.11537], whereas Zhao et al. obtain strong FWER via closed testing in group sequential designs [2311.15498]. Second, the price of maximal robustness is often conservativeness or computational burden. In the vaccine-assay framework, a large control-data-compatible set \(\mathcal A(\alpha')\) can make \(p_{\max}^{\rm adj}\) very conservative [2507.06451]. In minP, computational cost grows with \(K\times B\) because each analysis pipeline must be rerun for every permutation [2401.11537]. In group sequential closed testing, one must compute intersection-wise sequential p-values and carry out root-finding over many sets \(J\) and looks \(k\) [2311.15498]. In maximally selected rank statistics, exact or Monte Carlo corrections can be prohibitively slow [1605.03391].

A common misconception is that “maximum” always refers to stronger evidence against the null. In fact, the defining maximization is usually conservative: it selects the largest p-value compatible with admissible analytical variation, or calibrates significance against the most extreme null fluctuation. The result is not a stronger rejection rule but a more robust one.

Source: https://www.emergentmind.com/topics/maximally-adjusted-p-value