Probability of Successful Trials (PST)
- Probability of Successful Trials (PST) is a metric that quantifies the likelihood a trial meets a pre-specified success criterion, with definitions varying across research fields.
- It incorporates methodologies ranging from Bayesian sample-sizing and predictive probability designs in clinical trials to approaches in information theory and quantum computing.
- Accurate PST estimation demands careful definition of success criteria, selection of proper modeling techniques, and robust interim monitoring to guide decision-making.
Searching arXiv for the cited papers to ground the article in the relevant literature. Probability of Successful Trials (PST) is a family of quantities that quantify the probability that a pre-specified success criterion will be met, but the term is used in materially different ways across research domains. In clinical trials, PST is most often aligned with the probability of trial success at design or interim analysis, and is closely related to conditional power, predictive power of success, predictive probability, probability of success, and assurance (Kundu et al., 2021). In a proper Bayesian sample-sizing framework, PST is the probability—under the marginal predictive distribution—that the posterior probability of a success region exceeds a chosen threshold (Muirhead et al., 2012). In competing-risks treatment trials, the relevant success probability may instead be the cumulative incidence of a favorable event within a fixed window while accounting for death as a competing event (Beyersmann et al., 2020). Outside biostatistics, the acronym also appears in information theory as the probability of a successful guess (Issa et al., 2015), in sequential inference for Bernoulli success probabilities (Wills et al., 2017), and in quantum computing as a figure of merit based on recovering the initial state after a circuit and its inverse (Roy et al., 2024).
1. Conceptual scope and terminological variants
The literature does not use a single universal definition of PST. A consolidated presentation distinguishes at least four related quantities in clinical investigation: the conditional power (CP) based on interim results, the predictive power of success (PPoS) based on interim results with or without prior distribution, and the probability of success (PoS) for a prospective trial at the design stage; it also distinguishes trial success from clinical success (Kundu et al., 2021). In that usage, trial success refers to achieving a statistically significant result at final analysis, whereas clinical success refers to the observed treatment effect exceeding a clinically meaningful threshold.
A proper Bayesian formulation defines a trial as successful when the posterior probability of a success set exceeds a threshold. If denotes the success region and the required posterior probability, the success event is
Writing
the PST is
that is, the probability under the marginal predictive distribution that the eventual posterior will satisfy the decision criterion (Muirhead et al., 2012).
At interim analysis, a distinct but related definition is the Bayesian predictive probability of success: where is interim data, future data, and the critical region for success (Micoli et al., 2024). This averages conditional success probabilities over posterior uncertainty and is therefore an interim-monitoring quantity rather than a design-stage assurance measure.
This terminological heterogeneity explains why PST sometimes denotes a patient-level event probability, sometimes a design-stage assurance, and sometimes an interim predictive quantity. A plausible implication is that comparisons of “PST” across papers require close attention to the estimand, the stage of trial conduct, and the success criterion.
2. Frequentist and Bayesian formulations in clinical trials
In the frequentist-interim setting, conditional power is the probability of achieving trial or clinical success conditional on observed data so far and a fixed assumed post-interim effect size. With information fraction , interim estimate 0, final-analysis standard error parameter 1, success boundary 2, and assumed post-interim effect 3,
4
If 5,
6
By contrast, PPoS integrates over the uncertainty in the post-interim effect, optionally incorporating a prior distribution (Kundu et al., 2021).
With a normal prior on the treatment effect, the interim PPoS can be written as
7
where
8
and 9 are the prior mean and variance. Without prior information, the same source gives a reduced expression obtained by setting 0 (Kundu et al., 2021).
At the design stage, PoS or assurance averages the probability of final success over the prior: 1 This quantity is unconditional on interim data and is therefore distinct from CP and PPoS (Kundu et al., 2021).
For binary endpoints, exact predictive formulations are also available. With a Beta prior 2, interim 3 responses out of 4, and 5 future responses among the remaining 6,
7
and
8
The same framework extends to continuous and time-to-event endpoints, with endpoint-specific substitutions for the effect estimate and its variance (Kundu et al., 2021).
3. Bayesian sample sizing, assurance, and predictive monitoring
A fully Bayesian sample-sizing approach chooses 9 so that the predictive probability of satisfying the posterior success criterion is large enough. In the normal two-arm superiority model with known precision 0, treatment effect 1, prior mean difference 2, and posterior variance component
3
success is equivalent to
4
The resulting PST is
5
with 6 given by the corresponding marginal predictive variance (Muirhead et al., 2012).
A central feature of this framework is that 7 does not converge to 1 as 8; rather,
9
To normalize by the prior chance that the success region is true, the paper defines
0
which approaches 1 asymptotically (Muirhead et al., 2012). This shows that the prior imposes an upper bound on the attainable probability of success.
Later Bayesian work extends this perspective to richer data structures. For multivariate linear models with correlated endpoints, the probability of success is defined by integrating the indicator of a posterior success event over simulated future data, covariates, and parameters under a validation prior: 1 The associated Bayesian seemingly unrelated regression model accommodates multiple endpoints with different covariates and explicitly models endpoint correlations, while an adjustment is introduced for asymptotically strict family-wise error-rate control under multiplicity (Alt et al., 2020).
For interim monitoring in trials with competing event data, a simulation-based PPoS proceeds in three phases: a modeling phase using Bayesian cause-specific hazard models, a prediction phase that simulates future event times and event types conditional on interim data, and an analysis phase that applies the intended final analysis to each pseudo-final dataset. The Monte Carlo approximation is
2
where 3 records whether the 4-th simulated completion satisfies the success criterion (Micoli et al., 2024). This suggests that “probability of success” in adaptive monitoring is often a computational object defined by the final decision rule rather than by a closed-form formula.
4. Competing risks and endpoint choice in COVID-19 treatment trials
In COVID-19 treatment trials, one formulation aligned with PST is the probability that a patient experiences a favorable event within a prespecified time window while not first experiencing death. The recommended representation is in terms of time to and type of first event: 5 with 6 for improvement or recovery and 7 for death. The primary quantity is the cumulative incidence function for recovery,
8
where 9 is the event-specific hazard for improvement (Beyersmann et al., 2020).
This framing leads to three design principles for a successful COVID-19 treatment trial: increasing the probability of improvement/recovery within a specific time interval, expediting the timing of the favorable event, and not increasing mortality within that period (Beyersmann et al., 2020). The paper therefore recommends endpoints such as recovery by day 28 or time to recovery within 0, while emphasizing that deaths must be handled as competing events rather than as censoring events.
For estimation, the Aalen-Johansen estimator is identified as the gold standard for 1, whereas Kaplan-Meier is only valid if there are no competing risks and otherwise overestimates recovery probabilities. For regression modeling of cumulative incidence, the Fine & Gray model yields a subdistribution hazard ratio directly relevant to increasing the success probability (Beyersmann et al., 2020).
The contrast between hazard-based and cumulative-incidence quantities is explicit. The event-specific hazard ratio is
2
and the subdistribution hazard ratio is
3
Both may exceed 1 for a successful treatment, but they encode different concepts (Beyersmann et al., 2020).
Planning requires attention to event probabilities as well as hazards. Using the Schoenfeld formula,
4
and with
5
sample-size determination explicitly depends on the recovery probability and the competing-event probability. Under constant hazards,
6
with 7 for treatment and control, 8 the recovery hazard, and 9 the death hazard (Beyersmann et al., 2020).
The same source also treats timing of favorable events via the restricted mean time to recovery,
0
and discusses adaptive sample-size re-estimation using interim Aalen-Johansen estimates. A plausible implication is that, in competing-risk settings, PST is inseparable from the estimand and cannot be reduced to a simple hazard-ratio target.
5. Predictive probability designs and interim decision-making
Predictive probability designs use accumulating data to quantify the chance that the planned final success criterion will be met if the trial continues. In randomized biomarker-guided oncology trials, the posterior predictive probability is computed as
1
where 2 is the observed number of responses, 3 future responses among the remaining patients, 4 and 5 the control and experimental response rates, and 6 the posterior probability threshold for declaring efficacy (Zabor et al., 2022). Interim decision rules take the form: continue if 7, and stop early for futility if 8.
An optimal efficiency predictive probability method evaluates posterior and predictive thresholds by balancing power, type I error, and efficiency, with designs compared through expected sample size under the null and under the alternative (Zabor et al., 2022). Three randomized biomarker-guided oncology designs are described: pooled control arm, stratified control arm, and enrichment design, each using predictive probability for futility monitoring. Their reported operating characteristics include type I error control from 0.05 to 0.09 and power 0.80 to 0.86 for detecting a true effect in the predictive biomarker subgroup (Zabor et al., 2022).
A recent approximation paper simplifies interim predictive probability calculation under approximate normality. If 9 is an interim 0-value, 1 an interim posterior probability of superiority, 2 the information fraction, 3 a frequentist success threshold, and 4 a Bayesian threshold, the approximate predictive probabilities are
5
and
6
These approximations are reported to have a high degree of concordance with Monte Carlo imputation across dichotomous, time-to-event, ordinal, historical-borrowing, and longitudinal applications (Marion et al., 2024).
For trials with competing event data, the predictive-probability framework becomes model-based and simulation-driven. Cause-specific hazards
7
define the joint distribution of event time and type, and future event types at simulated event time 8 are sampled with probabilities proportional to the cause-specific hazards: 9 The resulting PPoS supports futility stopping, sample-size optimization, and analysis-time selection (Micoli et al., 2024).
6. Statistical inference on Bernoulli success probabilities
In sequential inference for Bernoulli trials, the phrase “success probability” refers not to trial success in drug development but to the parameter 0 of a Bernoulli sequence. Test supermartingales provide valid 1-value bounds and confidence intervals for hypotheses such as 2, with validity preserved under arbitrary stopping rules (Wills et al., 2017).
Let 3 denote Bernoulli trials, 4, and 5. For the composite null 6, a specific test-factor construction uses
7
and
8
The full test supermartingale is
9
and the reciprocal 0 is a valid 1-value bound for the null (Wills et al., 2017).
Confidence sets follow by inverting the 2-values. For comparison, the exact binomial tail is
3
the Chernoff-Hoeffding bound is
4
and the supermartingale-based 5-value obeys the ordering
6
Thus the exact method is least conservative, Chernoff-Hoeffding is intermediate, and the supermartingale method is most conservative (Wills et al., 2017).
The same paper quantifies the cost of optional-stopping robustness. The supermartingale method yields wider confidence intervals, with endpoint inflation of order 7, whereas exact and Chernoff-Hoeffding intervals are 8 in 9 under the asymptotic scaling reported there (Wills et al., 2017). This does not define PST as a trial-level success metric, but it does show that the probability of success parameter itself can be the object of robust sequential inference.
7. Extensions beyond clinical trials
In information-theoretic secrecy, the operational metric is the probability that an eavesdropper makes a successful guess of the source sequence within an acceptable distortion level. For i.i.d. source 00, encoder 01, and eavesdropper strategy 02, the success event is
03
and the probability of a successful guess is
04
The central asymptotic quantity is the exponent
05
with single-letter characterizations given for both the keyless and key-enabled settings (Issa et al., 2015). This usage shares the phrase “probability of a successful guess” with PST-like terminology, but the object is an adversarial success probability rather than a clinical or statistical design criterion.
In quantum circuit watermarking, Probability of Successful Trials is defined as
06
where 07 is the number of trials with an output identical to that of the initial state, and 08 is the total number of trials (Roy et al., 2024). Here PST measures functionality preservation: the fraction of shots whose outputs match the intended output of the non-watermarked circuit. The reported results state that PST is reduced by a minuscule 09 against the non-watermarked benchmarks and is up to 10 higher compared to existing techniques (Roy et al., 2024).
A closely related quantum-compilation usage defines PST as the probability of obtaining the initial state after running a quantum circuit followed by its inverse. With initial state 11, mirror circuit 12, and 13 shots, if 14 denotes the count of all-zero outcomes, then
15
That work further introduces a weighted version, wPST, in which outcomes are weighted by the fraction of correct zero bits, and reports that machine-learning-predicted FoMs increase the correlation with the true PST or wPST by over 16 (Singh et al., 3 Jul 2026).
These non-clinical usages illustrate that PST is not a domain-invariant technical term. This suggests that the unqualified phrase “Probability of Successful Trials” functions less as a single metric than as a recurring label for probabilities of operational success, with domain-specific sample spaces, decision rules, and loss structures.