- The paper introduces a closed-form correction factor based on tangent-boundary linearization, Brownian first-passage probabilities, and bivariate normal CDFs to estimate sample sizes for always-valid inference.
- The method reaches target power within about 3 percentage points in simulations, reduces sample requirements by roughly 8%–17.5%, and achieves median production savings of 9.5% across 713 Spotify metrics.
- The correction applies to smooth concave boundaries and arbitrary allocation ratios through the burn-in fraction, but users should verify boundary concavity and exercise caution with very low event rates or untested parameter settings.
Motivation and problem statement
Sequential tests that permit continuous monitoring—confidence sequences and mixture sequential probability ratio tests (mSPRTs)—are now standard on large experimentation platforms. Their power calculations, however, resist closed-form treatment: the always-valid (AV) power is the probability of a boundary crossing at any time from burn-in to the planned endpoint, a first-passage problem for which no elementary expression exists for the curved boundaries in practical use. The prevailing workaround is the "last-point rule," which inflates the fixed-sample size nz until the marginal rejection probability at the endpoint alone reaches 1−β. Because this ignores crossings before the endpoint, it is conservative by construction: at (α,1−β)=(0.05,0.80) the paper reports empirical power between $0.86$ and $0.88$, i.e., oversizing by seven to nine percentage points.
The paper's contribution is a closed-form correction factor k∗(α,β,t0), where t0=m/nz is the burn-in fraction, such that setting the total sample size to k∗⋅nz approximately achieves target AV power without simulation. The factor is expressed entirely in elementary functions and the bivariate normal CDF, requires only two Φ2 evaluations per bracket-search step, and depends on the boundary only through its value b(k) and slope 1−β0 at the planned endpoint—so it applies to any smooth concave boundary.
Setup: boundaries and rescaled time
The author considers three widely deployed boundaries under a two-sample model with allocation ratio 1−β1, standardised effect 1−β2, and fixed-sample size
1−β3
WSKR confidence sequence: 1−β4 with 1−β5 from the Robbins–Siegmund limiting distribution (Waudby-Smith et al., 2023). Maharaj confidence sequence: an asymptotic Gaussian-mixture supermartingale boundary with tuning constant 1−β6 involving the Lambert-1−β7 function. mSPRT: the Johari–Pekelis–Walsh boundary with mixing standard deviation set to the MDE, whose rejection rule coincides with a Bayes factor threshold [johariopre2022].
A key normalisation makes the analysis allocation-agnostic. Defining 1−β8 gives unit variance for every 1−β9; under (α,1−β)=(0.05,0.80)0 at (α,1−β)=(0.05,0.80)1, the rescaled process (α,1−β)=(0.05,0.80)2 converges weakly to Brownian motion with drift (α,1−β)=(0.05,0.80)3 and variance rate one, started at (α,1−β)=(0.05,0.80)4. Consequently (α,1−β)=(0.05,0.80)5 depends on (α,1−β)=(0.05,0.80)6 only through (α,1−β)=(0.05,0.80)7; at equal (α,1−β)=(0.05,0.80)8, the factor is identical across allocation ratios. This is confirmed empirically: simulated power at (α,1−β)=(0.05,0.80)9 and $0.86$0 matches the $0.86$1 values within Monte Carlo error when $0.86$2 is computed at the $0.86$3-specific $0.86$4.
All three boundaries are concave on $0.86$5—proved analytically for WSKR and Maharaj, verified numerically for mSPRT. Concavity is essential: it guarantees the tangent at the endpoint lies above the boundary, making the linearised crossing probability a lower bound on the true crossing probability and hence making the resulting $0.86$6 slightly conservative rather than anti-conservative.
The corrected factor
The derivation follows Siegmund's classical strategy: linearise the boundary by its tangent at the planned endpoint $0.86$7, apply the Bachelier first-passage formula to Brownian motion with drift over the linear surrogate, then integrate over the random initial value $0.86$8 using a Gaussian convolution identity that reduces each term to a bivariate normal CDF. The resulting approximation decomposes into three interpretable pieces:
$0.86$9
where $0.88$0 is the immediate-rejection probability at burn-in under the tangent surrogate, $0.88$1 is the crossing probability during $0.88$2, and $0.88$3 is the Bachelier reflection correction. The corrected factor $0.88$4 is the smallest root of $0.88$5.
At $0.88$6 and $0.88$7, representative factors are $0.88$8 (WSKR), $0.88$9 (Maharaj), and k∗(α,β,t0)0 (mSPRT), with savings over the last-point rule of roughly k∗(α,β,t0)1–k∗(α,β,t0)2. Across the full grid (k∗(α,β,t0)3, k∗(α,β,t0)4), savings range from about k∗(α,β,t0)5 to k∗(α,β,t0)6. The author cautions that because the three boundaries arise from different mathematical constructions, the differing factors should not be read as an efficiency ranking.
Two caveats attach to uniqueness and monotonicity. Monotonicity of k∗(α,β,t0)7 in k∗(α,β,t0)8 does not follow automatically—the tangent surrogate changes with k∗(α,β,t0)9, so crossing events are not nested—and uniqueness of the root is a numerical observation on the tested grid, not a theorem. Users applying the formula at untabulated parameters are advised to verify concavity of the mSPRT boundary numerically.
Simulation validation
Monte Carlo validation uses t0=m/nz0 replications per cell (MC SE below t0=m/nz1) across Gaussian, Bernoulli, and log-normal outcomes, allocation ratios t0=m/nz2, and effect sizes t0=m/nz3. Type-I error control is verified first with t0=m/nz4 null replications; empirical one-sided rejection rates stay below t0=m/nz5 for all three boundaries at every grid cell.
Empirical power at the corrected sample size lands close to target: at t0=m/nz6, values cluster around t0=m/nz7–t0=m/nz8 versus t0=m/nz9–k∗⋅nz0 under last-point sizing. On the extended grid—including Maharaj down to k∗⋅nz1—the closed form hits target power within approximately k∗⋅nz2 percentage points, with no degradation at extreme parameters, though savings shrink toward the low-k∗⋅nz3/high-power corner (e.g., k∗⋅nz4 at k∗⋅nz5, k∗⋅nz6). Against the Monte Carlo reference for k∗⋅nz7 itself, the closed form agrees within k∗⋅nz8 relative error (median k∗⋅nz9) for WSKR and Φ20 (median Φ21) for Maharaj.
Production validation and burn-in sensitivity
On Φ22 metrics drawn from the last 283 experiments on Spotify's Confidence platform (Maharaj boundary with Bonferroni corrections for multiple comparisons and metrics), the median saving is Φ23 and the mean Φ24. This is below the Φ25 at the nominal Φ26 cell because Bonferroni adjustment pushes effective Φ27 and Φ28 toward the lower-saving corner of the grid—an honest reconciliation between laboratory and production numbers.
Sensitivity analysis over burn-ins Φ29 reveals a practically important trade-off. Decreasing b(k)0 widens the WSKR and Maharaj boundaries through the b(k)1 term, and this widening more than offsets the longer monitoring window: b(k)2 increases as b(k)3 decreases. At realistic burn-ins (b(k)4), the total budget b(k)5 is b(k)6–b(k)7 the fixed-sample size—considerably more than the factor near b(k)8 obtained if monitoring begins only at b(k)9. The mSPRT boundary, which does not depend on 1−β00, yields 1−β01 essentially constant (variation under 1−β02 across 1−β03). The implication is that the baseline cost of anytime-valid inference with realistic early monitoring is substantially higher than naive extrapolation from late-start monitoring suggests, even after the correction.
Limitations and open questions
Several limitations are stated plainly. The Gaussian approximation replaces the estimated 1−β04 in the WSKR construction with known 1−β05, moving from the distribution-free guarantee to a CLT-based one; at low base rates this degrades measurably—with fewer than roughly ten expected successes per arm, empirical power falls below target (gap of 1−β06 at 1−β07 and 1−β08 at 1−β09). Concavity of the mSPRT boundary is verified only numerically on the tabulated grid. Uniqueness of the root and monotonicity of 1−β10 remain unproven in general. Finally, the Bernoulli robustness checks involve unequal arm variances under 1−β11, outside the common-variance model on which the theory rests. Whether a rigorous monotonicity proof, a finite-population or non-Gaussian analogue of the correction, or an extension beyond concave boundaries is possible remains open.
Conclusion
This paper replaces simulation-dependent sizing for always-valid sequential tests with a closed-form factor requiring two bivariate normal CDF evaluations, valid for any smooth concave boundary and any allocation ratio through the single parameter 1−β12. It corrects a seven-to-nine-percentage-point conservatism in the last-point rule, saves 1−β13–1−β14 of the sample budget across the operating range (median 1−β15 in production), and quantifies the previously underappreciated cost of wide burn-in windows: realistic anytime-valid testing requires budgets two to three times the fixed-sample size.