Conditional Power in Adaptive Designs
- Conditional power approach is a set of methods that estimate the probability of eventual success based on interim or conditioned data.
- It is applied in clinical trials and dose-finding studies to adjust sample sizes and decision rules while rigorously controlling type-I error.
- In model-X testing and market risk, the approach optimizes test power and risk measures by embedding power characteristics into statistical models.
Searching arXiv for the cited papers and closely related work on conditional power to ground the article. Conditional power approach denotes a family of methodologies that use the probability of ultimate statistical success, conditional on interim or conditioned data, to guide model construction or design adaptation. In confirmatory and dose-finding clinical trials, it refers to methods that evaluate the probability of ultimately rejecting the null hypothesis given interim data and use that quantity to motivate sample size changes, promising-zone decisions, or other adaptations (Kunzmann et al., 2020). In model-X conditional randomization testing, the phrase can denote training schemes that explicitly optimize models for the power of a conditional independence test rather than for predictive accuracy alone (Shaer et al., 2022). A distinct use of “conditional power” also appears in market-risk modeling, where a conditional autoregressive risk framework is built on the Skew Exponential Power distribution and the power parameter determines whether the conditional location represents a quantile or an expectile (Bottone et al., 2019). Across these settings, the unifying idea is that a conditional probability-of-success or conditional power-family structure is made central to either inference, adaptation, or estimation.
1. Clinical-trial meaning in confirmatory settings
In confirmatory clinical trials, the conditional power approach is the family of methods that looks, at an interim analysis, at the probability of ultimately rejecting the null hypothesis given the interim data, and uses that probability—estimated in one of several ways—to motivate sample size changes such as increasing enrollment for promising-but-weak results or stopping for futility (Kunzmann et al., 2020).
The basic setting used for exposition in Kunzmann et al. is a one-arm normal approximation with i.i.d. outcomes of mean and variance $1$ and null hypothesis
With sample size , the Wald-type statistic is
with a fixed-sample rejection rule for a one-sided design (Kunzmann et al., 2020).
Conditional power is distinguished from unconditional power. Unconditional power is the usual planning quantity,
whereas conditional power is the probability of eventual rejection given interim information. If an interim look at yields 0 and 1, then
2
and the conditional power is
3
This makes conditional power a random quantity because it depends on the interim trajectory rather than only on design parameters (Kunzmann et al., 2020).
Kunzmann et al. treat conditional power as an estimand and compare three estimators: assumed conditional power, observed conditional power, and predictive power. Assumed conditional power fixes the design alternative 4 and sets
5
Observed conditional power plugs in the interim MLE 6,
7
Predictive power averages conditional power over a posterior for 8 restricted to beneficial effects,
9
The paper concludes that predictive power is the posterior mean of conditional power and minimizes posterior MSE under the specified prior, whereas observed conditional power has large MAE and MSE and is hard to justify theoretically (Kunzmann et al., 2020).
A central methodological claim in this literature is that pre-planning a binding interim analysis using methodology for unplanned interim analyses is ineffective and naturally leads instead to optimal two-stage or group-sequential designs. This suggests that, in confirmatory settings, conditional power is best viewed as a diagnostic or a basis for rare unplanned adaptations driven by trial-external information, operational necessity, or post hoc changes in the objective criterion, rather than as a fixed threshold to be enforced in a pre-planned design (Kunzmann et al., 2020).
2. Sample size re-estimation in Phase II dose-finding
In Phase II dose-finding, the conditional power approach is developed for unblinded sample size re-estimation in multi-arm contrast-test designs. The objective is Proof-of-Concept and dose-response exploration in a one-way ANOVA model with continuous efficacy outcome, 0 arms, common known variance 1, and allocation proportions 2 (Liu et al., 2020).
For a single contrast vector 3 satisfying the contrast constraint 4, the null hypothesis is
5
with standardized statistic
6
and contrast effect
7
The one-sided test rejects if 8 (Liu et al., 2020).
The two-stage framework fixes stage sizes 9 and $1$0 and information fraction
$1$1
Without adaptation, the final combination statistic is
$1$2
With sample size re-estimation, stage 2 is changed to $1$3, but the planned weight $1$4 is kept. The reweighted statistic is
$1$5
and the null is rejected if $1$6 (Liu et al., 2020).
For a single contrast, conditional power is defined as the frequentist probability, conditional on the interim data, that the final test will reject the null if the trial continues with stage-2 sample size $1$7 and assumed true effect $1$8:
$1$9
The paper derives an explicit expression and shows that, for fixed 0, conditional power is strictly increasing in 1 (Liu et al., 2020).
A promising-zone decision rule partitions interim outcomes into three regions using 2 and a lower threshold 3. If conditional power is below 4 or 5, the result is unfavorable and sample size is not increased. If conditional power is between 6 and the target power 7, the result is promising and the stage-2 sample size is increased. If conditional power is already at least 8, the result is favorable and the original size is maintained (Liu et al., 2020).
For the single-contrast case, solving the conditional power equation for the required stage-2 size yields
9
used only when sample size re-estimation is triggered (Liu et al., 2020).
The same logic extends to multiple contrasts. The paper forms a matrix 0 of candidate contrasts, uses the maximum of a multivariate normal vector of standardized contrast statistics, and defines conditional power by the probability that the final maximum exceeds a multiplicity-adjusted critical value 1. In the multiple-contrast setting, sample size re-estimation requires multivariate normal integration and numerical solution for 2, but the conceptual structure of unfavorable, promising, and favorable zones remains the same (Liu et al., 2020).
3. Type I error control and adaptive design principles
A defining feature of the clinical-trial conditional power approach is that sample size adaptation must preserve strict type-I error-rate control. Both the dose-finding and confirmatory analyses rely on the conditional error or combination-test principle (Liu et al., 2020, Kunzmann et al., 2020).
In the Phase II dose-finding formulation, type I error is protected by keeping the pre-fixed weights 3 and 4 in the combination statistic. Under 5, changing 6 rescales the stage-2 statistic but does not change its standardized null distribution. As a result, for the single-contrast case,
7
under 8, independent of 9, and in the multiple-contrast case
0
under 1, again independent of 2. Hence the same critical values 3 and 4 guarantee the nominal level regardless of how the sample size re-estimation decision used the interim data (Liu et al., 2020).
Kunzmann et al. formalize the same principle through a conditional error function induced by the original fixed or two-stage design. Any modified second-stage test must have, for any interim data, conditional type-I error not exceeding that original conditional error (Kunzmann et al., 2020). This produces regulatory feasibility for unplanned unblinded recalculations, but the paper emphasizes that regulatory acceptability does not imply efficiency.
The principal criticism advanced in the confirmatory-trial literature is that naïve rules such as “keep conditional power above 80%” can create paradoxical outcomes. When such a rule is made binding for every possible interim outcome, it effectively defines a fully specified two-stage design with a random sample size. Kunzmann et al. show that this can reduce expected power and under-use type I error relative to an optimal two-stage design, because maximum-sample-size constraints implicitly introduce a binding futility boundary (Kunzmann et al., 2020).
A plausible implication is that conditional power thresholds are not, by themselves, optimality criteria. The literature instead separates two tasks: preserving type I error through conditional error or combination-test logic, and choosing sample-size or stopping rules through a distinct objective such as minimizing expected sample size subject to power and type-I error constraints (Kunzmann et al., 2020).
4. Bayesian predictive power and alternative estimands
Conditional power depends on an assumed true effect, and this dependence is treated as a primary methodological issue in both the dose-finding and confirmatory papers. In response, both papers discuss Bayesian predictive power as a posterior expectation of conditional power (Liu et al., 2020, Kunzmann et al., 2020).
In the Phase II dose-finding setting, posterior predictive power is defined for the single-contrast case as
5
Thus conditional power treats 6 as a fixed unknown, whereas predictive power treats 7 as random with a posterior distribution (Liu et al., 2020).
The paper gives closed-form expressions for predictive power under a non-informative prior 8 and under independent conjugate normal priors 9. In both cases, the final predictive distribution of the reweighted combination statistic is normal, so predictive power reduces to a Gaussian tail probability. The authors note that predictive power “shrinks” toward 0 relative to conditional power evaluated at the observed estimate because it incorporates extra uncertainty (Liu et al., 2020).
Kunzmann et al. frame predictive power in a related but not identical Bayesian manner. With prior density 1 and posterior truncated to beneficial effects, predictive power is
2
They emphasize that this conditioning on 3 aligns the Bayesian quantity with the frequentist notion of power, which is defined under alternatives (Kunzmann et al., 2020).
Both papers draw a consistent contrast. Conditional power is simpler and usually leads to smaller sample-size increases, but is sensitive to the assumed effect. Predictive power is more robust to effect uncertainty or prior misspecification, but is more conservative and can require larger sample-size increases for the same target probability level (Liu et al., 2020, Kunzmann et al., 2020).
A common misconception is that predictive power simply replaces conditional power. The papers instead treat predictive power as an estimator or posterior average of conditional power. This suggests that the two notions are not competing definitions of success probability but different inferential treatments of the unknown effect parameter.
5. Optimal design, inefficiency, and controversy
The confirmatory-trial literature places strong emphasis on the distinction between unplanned and pre-planned uses of conditional power. Kunzmann et al. argue that if an interim analysis and adaptation rule are fully specified in advance, then the resulting procedure is simply a two-stage design and should be optimized directly rather than constructed by imposing a conditional-power threshold (Kunzmann et al., 2020).
Their optimization problem seeks a two-stage design with interim at 4, stage-2 sample-size function 5, and critical-value function 6 that minimizes expected sample size under a planning prior while satisfying a type-I error constraint and an expected-power constraint. They compute such designs numerically using the R package adoptr, approximating 7 and 8 via spline functions (Kunzmann et al., 2020).
The optimal designs they report differ qualitatively from conditional-power-threshold rules. Predictive power at interim is not constant across outcomes and can be as low as about 9 near futility. The stage-2 sample-size function is non-convex, with a maximum near the futility boundary, rather than the convex shape induced by rules that force conditional or predictive power above a fixed level (Kunzmann et al., 2020).
This has generated a substantive controversy around the phrase “conditional power approach.” In one usage, the phrase denotes a practical heuristic for mid-trial adaptation; in another, it is a target of criticism because it is not generally aligned with standard design-optimality criteria. The literature resolves this tension by reserving conditional-power-based recalculation primarily for unplanned adaptations caused by new trial-external evidence, operational needs, or revised objectives, while recommending optimal two-stage or group-sequential design theory for pre-planned interim looks (Kunzmann et al., 2020).
Kunzmann et al. also propose two coherent recalculation schemes for reactions to external new evidence after an originally optimal design has been chosen: a fixed conditional type-II error recalculation and a 0-approach preserving the original trade-off between sample size and predictive power. Both preserve type I error through conditional error control and revert to the original design when planning assumptions and interim timing are unchanged (Kunzmann et al., 2020).
This suggests that the mature version of the conditional power approach is not a universal recipe for “promising zone” inflation of sample size, but a broader framework in which conditional success probabilities are diagnostic quantities embedded in error-preserving adaptive design theory.
6. Power-oriented learning in conditional randomization tests
Outside clinical trials, the phrase acquires a different but structurally related meaning in model-X conditional randomization testing. The paper “Learning to Increase the Power of Conditional Randomization Tests” explicitly trains predictive models to improve the power of conditional independence tests while preserving type-I error guarantees (Shaer et al., 2022).
The model-X conditional randomization test considers i.i.d. observations 1 with 2 and tests
3
Assuming knowledge or approximation of 4, the holdout randomization test fits a predictive model 5 on a training set and evaluates a test statistic on a test set. With squared-error loss,
6
Randomized dummy features 7 are drawn from the conditional distribution, and the randomization p-value is
8
Under the null, exchangeability of 9 guarantees
0
The methodological innovation is to replace pure predictive training by training that explicitly targets test power. The paper defines the risk discrepancy for feature 1,
2
If 3 is non-null, one wants this quantity to be large and positive because replacing the true feature by a randomized conditional copy should damage predictive performance. Under the null, the paper shows that
4
for any fixed 5 (Shaer et al., 2022).
Training is then based on an MRD penalized objective:
6
where 7 is the usual prediction objective, 8 and 9 are empirical risks with the original and dummy features, and
00
uses the sigmoid 01 (Shaer et al., 2022).
This is a conditional power approach in the sense that the model is fitted for the randomization test rather than merely for predictive accuracy. The objective is aligned with the separation between the observed statistic and its conditional null randomization distribution. The paper reports that this increases the number of correct discoveries obtained, while maintaining type-I error rates under control, across lasso, elastic net, and deep neural networks (Shaer et al., 2022).
A plausible implication is that the term “conditional power” extends beyond interim monitoring in trials to any setting in which a conditional success probability or power proxy is internalized in the estimation objective. In this machine-learning setting, the power-oriented quantity is not a future rejection probability under interim data, but an expected discrepancy between observed and conditionally randomized test statistics.
7. Conditional power and the exponential power family in market risk
A separate usage appears in market-risk measurement through a unified Bayesian Conditional Autoregressive Risk Measures framework based on the Skew Exponential Power distribution. Here the relevant sense of “power” arises from the exponential power family and, specifically, the power parameter 02 that determines whether the autoregressive location 03 represents a quantile or an expectile (Bottone et al., 2019).
The framework models returns as
04
with autoregressive risk equation
05
where 06 is the News Impact Curve. Conditionally on past information,
07
The Skew Exponential Power likelihood includes a location parameter 08, scale 09, skewness parameter 10, and power parameter 11, with skewness governed by different exponential-decay rates on the left and right of 12 (Bottone et al., 2019).
The paper’s “conditional power approach” arises because the conditional distribution is an exponential power distribution and the power parameter 13 determines the interpretation of 14. With 15, the SEP likelihood reproduces the Asymmetric Laplace structure and 16 is the conditional 17-quantile, yielding a CAViaR specification. With 18, the SEP likelihood reproduces the Asymmetric Gaussian structure and 19 is the conditional 20-expectile, yielding a CARE or CAvE specification (Bottone et al., 2019).
This unifies CAViaR and CARE in a single Bayesian conditional autoregressive framework. The paper further extends the News Impact Curve semiparametrically using penalized B-splines,
21
with a second-order random walk prior on the spline coefficients and an adaptive Independent Metropolis-within-Gibbs MCMC algorithm adapted from Bernardi, Bottone, and Petrella (2018) (Bottone et al., 2019).
Empirically, the methodology is applied to daily returns of five stock indices: Nasdaq, Hang Seng, Korea SE, AEX, and STI, from 1988–2018. The paper compares BNL-CAViaR and BNL-CARE with standard parametric CAViaR and CARE specifications using violation ratios 22, Kupiec’s unconditional coverage test, Christoffersen’s conditional coverage test, the Engle–Manganelli Dynamic Quantile test, and an ES bootstrap test on standardized residuals. It reports that both BNL-CAViaR and BNL-CARE perform comparably to, and sometimes as well as, traditional CAViaR and CARE models and that all models generally fall in the Basel “green” or “yellow” zones (Bottone et al., 2019).
In this financial context, “conditional power approach” does not refer to interim sample size recalculation. Instead, it refers to a conditional autoregressive risk framework built on an exponential power family, where the power parameter governs the risk functional itself. This is a terminological divergence rather than a contradiction.
8. Comparative perspective
The literature supports at least three technically distinct meanings of conditional power approach.
| Setting | Core conditional quantity | Primary use |
|---|---|---|
| Confirmatory and dose-finding trials | Probability of final rejection given interim data | Sample size re-estimation and adaptation |
| Model-X conditional randomization tests | Power-oriented discrepancy under conditional randomization | Training test statistics for higher discovery power |
| Bayesian CARM market-risk models | Conditional SEP specification indexed by power parameter 23 | Unifying quantile- and expectile-based risk measures |
In clinical trials, conditional power is an interim probability of eventual rejection, possibly estimated by assumed conditional power, observed conditional power, or predictive power. The central methodological issues are sensitivity to the assumed effect, strict type-I error control, and the distinction between pre-planned and unplanned adaptations (Liu et al., 2020, Kunzmann et al., 2020).
In model-X testing, the concept is transposed into a “learn to test” paradigm: the conditional null mechanism enters the training loss through risk discrepancy, so the fitted model is optimized for test power rather than prediction alone (Shaer et al., 2022).
In market risk, the term is tied to the exponential power distribution itself. There, conditionality refers to time-varying autoregressive risk functionals and power refers to the shape parameter 24 of the likelihood, not to a rejection probability (Bottone et al., 2019).
A common misconception is that “conditional power approach” has a single universal definition. The cited literature shows instead that the phrase is domain-dependent. In biostatistics it is primarily a design-adaptation concept; in model-X inference it denotes power-aware training under a conditional randomization mechanism; in financial econometrics it refers to a conditional autoregressive framework built on an exponential power family. What unites these uses is the centrality of a conditional structure and a power-related quantity, but the underlying mathematical objects and inferential roles differ substantially (Kunzmann et al., 2020, Liu et al., 2020, Shaer et al., 2022, Bottone et al., 2019).