Maximally Adjusted P-Value: Concepts & Applications
- Maximally adjusted p-value is a statistical adjustment method that calibrates p-values against worst-case scenarios from multiplicity, selection, nuisance variation, or sequential monitoring.
- It employs techniques such as permutation-based minP, closed testing, and worst-case nuisance adjustments to maintain valid error control across diverse analytical setups.
- While ensuring stringent Type I error protection, this approach can be computationally intensive and conservative due to evaluation over a wide range of admissible configurations.
A maximally adjusted p-value is a p-value obtained by calibrating inference against the most challenging admissible configuration induced by multiplicity, selection, nuisance variation, or sequential monitoring. In the literature, the construction is context dependent rather than universal. In some settings it is literally the largest p-value over a set of candidate analyses or nuisance-parameter values; in others it is the p-value attached to an extremal statistic such as the minimum raw p-value across analyses, the maximum sequential p-value across intersection hypotheses, or the running extremum under optional stopping. The common objective is conservative but valid inference under a larger decision space than a single prespecified test, especially when nominal p-values would otherwise be anti-conservative or selectively reported (Mandl et al., 2024, Zhao et al., 2023, Chen et al., 8 Jul 2025, Ramsey, 2016, Balsubramani, 2020).
1. Core idea and major formalizations
The literature suggests that “maximally adjusted p-value” is best understood as a family of worst-case adjustments rather than a single standardized statistic. Across the cited settings, the adjustment takes one of three recurrent forms: maximization over candidate p-values, maximization over closed-testing components, or calibration against the null distribution of an extremum. This suggests a unifying principle: the adjusted p-value is designed to remain valid after accounting for a space of admissible analytical choices that is larger than the single analysis ultimately reported.
| Context | Defining adjustment | Stated error property |
|---|---|---|
| Researcher degrees of freedom (Mandl et al., 2024) | permutation tail probability of the minimum raw p-value | weak FWER under the global null |
| Group sequential multiple testing (Zhao et al., 2023) | , then | strong FWER |
| Paired-control vaccine assays (Chen et al., 8 Jul 2025) | asymptotically valid Type I control | |
| PC-Max causal discovery (Ramsey, 2016) | over candidate conditioning sets | consistency under standard PC assumptions |
| Optional stopping (Balsubramani, 2020) | for an always-valid p-value process | Type I control under arbitrary stopping |
| Functional hypotheses (Xu et al., 2019) | pointwise based on null ranks of domainwise minima | FWER over the function domain |
These definitions are not interchangeable. In the group sequential and vaccine-assay settings, the maximization is explicit in the formula. In researcher-degrees-of-freedom problems, the paper instead uses the phrase “max-adjusted (minP) p-value” for the tail probability of the smallest observed p-value under a permutation-estimated global-null distribution. In optional-stopping work, the final adjusted quantity is a minimum over always-valid p-values, but it is derived from control of the running maximum of a supermartingale. The terminology therefore depends on which object is being maximized or calibrated.
2. Researcher degrees of freedom and the minP construction
Mandl et al. formalize researcher degrees of freedom as a multiple-testing problem in which analysis strategies are applied to the same dataset, producing raw p-values . The observed extremum is
Because the test statistics from alternative analysis strategies are usually highly dependent, the paper argues that Bonferroni is inappropriate owing to an unacceptable loss of power. The proposed remedy is a permutation-based minP adjustment that approximates the null distribution of the minimum p-value while preserving dependence among the candidate analyses (Mandl et al., 2024).
For permutations , one constructs permuted datasets 0 that enforce the global null, recomputes all 1 analyses, and records
2
The adjusted p-value is then
3
By construction, this estimates the tail probability of observing a minimum p-value at least as small as 4 under the global null. The method provides weak control of the family-wise error rate, with the paper emphasizing that weak control is sufficient for the specific aim of guarding against the case in which all admissible analytical choices correspond to null effects.
The practical rationale is dependence adaptation. Bonferroni replaces each 5 by 6, which is valid under arbitrary dependence but conservative when tests are positively correlated. The permutation-minP method instead estimates the true null distribution of 7, and the paper states that this yields a smaller correction and boosts power when analysis strategies are highly correlated.
The empirical illustration concerns perioperative 8 and post-operative complications after neurosurgery. The study included 9 patients, a binary outcome, and 0 analysis strategies spanning missing-data handling, surrogate modeling for unobserved 1, aggregation of repeated measures, and alternative codings or tests of 2. With 3 permutations obtained by shuffling the binary outcome, unadjusted testing produced an approximately 4–5 chance of at least one false positive among 48 p-values at 6, Bonferroni kept FWER below 7 but was overly conservative for moderate 8, and minP held FWER at approximately 9 for all sample sizes. Under the real association, subsampling experiments with 0–1 and 2 repetitions showed that minP yielded substantially more significant findings than Bonferroni at 3, while controlling FWER.
3. Closed testing and group sequential multiplicity
In group sequential designs with 4 elementary hypotheses and 5 looks, Zhao et al. define an adjusted-sequential p-value by combining multiple-hypothesis adjustment with repeated-analysis adjustment through closed testing. For an intersection hypothesis 6, the sequential p-value at look 7 is
8
or equivalently
9
The adjusted-sequential p-value for an elementary hypothesis 0 is
1
This is the closed-testing maximum over all intersection hypotheses containing 2. The paper then defines a global maximally adjusted p-value,
3
described as the smallest 4 for which at least one 5 would be rejected at or before its look. The stated decision rule is that at least one hypothesis is rejected if and only if 6, and 7 is rejected at the earliest 8 for which 9. By the closed-testing principle, this yields strong FWER control (Zhao et al., 2023).
The construction is algorithmic. One chooses a multiple-testing method, such as weighted Bonferroni with graph-recycling or weighted parametric group sequential design, specifies an 0-spending function such as Hwang–Shih–DeCani, and defines a weighting graph. For each nonempty intersection 1 and each look 2, one computes local 3-allocations and corresponding boundaries 4, then solves for the root in 5 at which the rejection event “just becomes true.” One-dimensional root-finding in 6 is sufficient because the unknown is scalar.
The worked example has 7 populations, 8 looks, weights 9, and HSD0 spending. Under weighted Bonferroni, the paper reports 1 at interim and 2 at final. The resulting elementary adjusted-sequential p-values are all 3 at interim and all 4 at final, so
5
Because 6, none of the three hypotheses is significant at either look. The example illustrates the intended role of 7 as a single scalar summary for an entire multi-hypothesis, multi-look trial.
4. Worst-case nuisance adjustment in paired-control vaccine assays
Chen et al. introduce a maximally adjusted p-value for intracellular cytokine staining data observed before and after vaccination, with paired control samples used to account for assay variation across time points. The motivating problem is differential misclassification between assay runs: false-positive and false-negative rates may differ between baseline and post-vaccination measurements, and naive comparison of estimated positive-cell proportions can therefore inflate false positives (Chen et al., 8 Jul 2025).
For each participant, the framework uses a primary baseline sample 8, a primary post-vaccination sample 9, and paired control samples 0 and 1. Let
2
denote the four misclassification rates. From the paired control data one constructs a 3 confidence set
4
where 5 is the standardized difference in control proportions at 6 versus 7. For a fixed 8, the primary-sample one-sided p-value for testing 9 against 0 is
1
The maximally adjusted p-value is then
2
The additive 3 term is the Berger–Boos correction. The interpretation given in the paper is explicitly worst case: one takes the largest p-value over all misclassification-rate scenarios consistent with the paired control data. The method therefore preserves Type I error no matter which compatible batch-effect scenario is true.
Its validity depends on several assumptions stated in the paper: within each assay run, all cells share the same misclassification rates; primary and control samples in that run share these rates; the true positive-cell proportion in the control is the same at 4 and 5; 6 and 7 are approximately standard normal under large-sample asymptotics; and 8 is effectively compact. Under these conditions, Theorem 3.1 gives asymptotic validity:
9
The paper contrasts 0 with a minimally adjusted p-value 1, defined as the member of 2 closest to the naive unadjusted p-value 3. By construction 4, but only the maximally adjusted version is described as always conservative enough to hold Type I error at the nominal level.
The practical consequence is a responder classification criterion that is robust to all control-data-compatible batch effects. The paper explicitly recommends 5 in regulatory or other high-stakes settings where Type I control is paramount, while also noting that weak control-data information can make the confidence set 6 large and the resulting adjustment very conservative. In the CoVPN 3008 case study, some participants with unadjusted 7 had maximally adjusted p-values of 8, reflecting strong batch-effect signals in the paired controls.
5. Algorithmic and sequential uses beyond classical multiple testing
In causal structure learning, the maximally adjusted p-value appears as a selection rule inside the PC-Max algorithm. For a pair of variables 9 and 00, one enumerates candidate conditioning sets 01, performs conditional-independence tests for 02, and records p-values 03. The algorithm defines
04
During the adjacency phase, at conditioning-set size 05,
06
and the edge 07 is removed if 08. The maximizing separating set is stored and later used in collider orientation. The algorithm also ranks candidate colliders in descending order of 09, orients them subject to a no-bidirected-edge constraint, and then applies Meek’s rules (Ramsey, 2016).
This usage is operational rather than directly inferential: the largest p-value is treated as the least evidence against conditional independence among admissible separating sets. Under causal sufficiency, a DAG data-generating process satisfying Markov and faithfulness, and a consistent conditional-independence test, the paper states a consistency theorem: as 10, PC-Max recovers the true CP-DAG with probability tending to one. In simulations with 11, average degree 12, 13, Fisher 14 testing, and 15, PC-Max achieved adjacency precision 16, adjacency recall 17, arrowhead precision 18, arrowhead recall 19, bidirected edges 20, and runtime 21 seconds on a 4-core laptop. At 22 and 23, it retained adjacency precision 24, adjacency recall 25, arrowhead precision 26, arrowhead recall 27, and runtime 28 seconds.
A different sequential use appears in work on p-value peeking. Here the goal is to make testing robust to arbitrary stopping rules by working with a nonnegative supermartingale 29 under 30 and its always-valid p-value process 31. The fully adjusted p-value reported after observing up to 32 samples is
33
Although this quantity is a minimum rather than a maximum, its validity comes from control of the running maximum of 34 via Ville’s inequality,
35
The paper therefore describes the resulting procedure as a maximally peeking-robust test, and the final 36 can be interpreted as the p-value adjusted against the most extreme excursion induced by arbitrary peeking (Balsubramani, 2020).
The implementation can be based on mixture martingales 37, updated online with 38 computation per new sample when 39 is discretized over 40 components. In a one-sided Gaussian mean-test simulation with 41 repeated experiments and horizon 42, the empirical rejection rate at level 43 was 44 for the adjusted 45, versus over 46 for the naive peeked p-value.
6. Related extremal adjustments, conceptual boundaries, and limitations
Several adjacent literatures rely on the same extremal logic even when they do not use the exact phrase “maximally adjusted p-value.” In functional hypothesis testing, Xu and Reiss define pointwise adjusted p-values by ranking the observed statistic function against simulated null functions and calibrating against the minima of rank processes across the function domain. The single-step adjusted p-value is
47
where 48. This controls FWER over the domain, and step-down and extreme-rank-length refinements can yield lower adjusted p-values while maintaining FWER control (Xu et al., 2019).
In survival forests based on maximally selected rank statistics, the inferential problem is different but closely related. For a continuous covariate, one computes
49
over admissible cutpoints 50. The naive p-value 51 is described as far too liberal because the maximum is taken over many highly correlated statistics. Wright et al. therefore discuss several adjusted approximations, including the Brownian-bridge approximation 52, the improved Bonferroni approximation 53, and the conservative rule 54. Their simulations show a trade-off between unbiased split-variable selection and runtime, with 55 recommended as a practical compromise (Wright et al., 2016).
For discrete p-values, optimal-transport methods provide yet another related but distinct notion of adjustment. One line of work formulates an adjusted discrete statistic as the 56-optimal projection of a discrete null onto a continuous surrogate, recovering Lancaster’s mid-57 and mean-value 58 constructions as special cases, and then replacing the usual 59 null by a moment-matched Gamma distribution to reduce conservativeness (Contador et al., 2023). A later extension generalizes this framework to Fisher, Pearson, George, Stouffer, and Edgington combinations, matching means and variances within surrogate families such as Gamma or Normal, with asymptotic level accuracy and power close to the corresponding continuous procedures when the likelihood-ratio statistic is monotonic in the combination statistic (Contador et al., 4 Aug 2025). These are p-value adjustments, but not “maximally adjusted” in the sense of maximizing over admissible analyses or nuisance configurations.
Two recurrent limitations run through the literature. First, the strength of error control varies sharply by construction. Mandl et al. emphasize only weak FWER control under the global null for permutation minP (Mandl et al., 2024), whereas Zhao et al. obtain strong FWER via closed testing in group sequential designs (Zhao et al., 2023). Second, the price of maximal robustness is often conservativeness or computational burden. In the vaccine-assay framework, a large control-data-compatible set 60 can make 61 very conservative (Chen et al., 8 Jul 2025). In minP, computational cost grows with 62 because each analysis pipeline must be rerun for every permutation (Mandl et al., 2024). In group sequential closed testing, one must compute intersection-wise sequential p-values and carry out root-finding over many sets 63 and looks 64 (Zhao et al., 2023). In maximally selected rank statistics, exact or Monte Carlo corrections can be prohibitively slow (Wright et al., 2016).
A common misconception is that “maximum” always refers to stronger evidence against the null. In fact, the defining maximization is usually conservative: it selects the largest p-value compatible with admissible analytical variation, or calibrates significance against the most extreme null fluctuation. The result is not a stronger rejection rule but a more robust one.