Papers
Topics
Authors
Recent
Search
2000 character limit reached

A closed-form sample size correction for always-valid inference with optional stopping

Published 16 Jun 2026 in stat.ME | (2606.18366v1)

Abstract: Sequential tests that allow continuous monitoring are common in A/B experimentation. Power calculations for these tests require simulations that are hard to scale across many metrics on an experimentation platform. Instead, a common sizing heuristic inflates the fixed-sample size until the marginal rejection probability at the planned endpoint reaches $1-β$. This last-point rule is conservative because always-valid (AV) power is the probability of a boundary crossing at any time during the run, not at the endpoint alone. We give a closed-form correction factor k<sup>α,β,t0</sup>k<sup>α, β, t_0</sup> expressed in elementary functions and the bivariate normal CDF, where t0=m/nzt_0 = m/n_z is the burn-in fraction. The closed-form approximation depends on the boundary only through its value and slope at the planned endpoint and can be evaluated for any smooth concave boundary. We work out three cases: the confidence sequences of Waudby-Smith et al. (2023) and Maharaj et al. (2023), and the mixture sequential probability ratio test of Johari et al. (2022). Setting the total sample size to knzk^ \cdot n_z, where nzn_z is the fixed-sample size for allocation ratio rr, hits empirical power within approximately 3 percentage points of target in Gaussian simulations. The correction factor depends on the allocation ratio rr only through t0=m/nz(r)t_0 = m/n_z(r). We study sensitivity to the burn-in parameter and show that the correction saves 8--20% of the last-point sample budget across the operating range.

Authors (1)

Summary

  • The paper introduces a closed-form correction factor based on tangent-boundary linearization, Brownian first-passage probabilities, and bivariate normal CDFs to estimate sample sizes for always-valid inference.
  • The method reaches target power within about 3 percentage points in simulations, reduces sample requirements by roughly 8%–17.5%, and achieves median production savings of 9.5% across 713 Spotify metrics.
  • The correction applies to smooth concave boundaries and arbitrary allocation ratios through the burn-in fraction, but users should verify boundary concavity and exercise caution with very low event rates or untested parameter settings.

Motivation and problem statement

Sequential tests that permit continuous monitoring—confidence sequences and mixture sequential probability ratio tests (mSPRTs)—are now standard on large experimentation platforms. Their power calculations, however, resist closed-form treatment: the always-valid (AV) power is the probability of a boundary crossing at any time from burn-in to the planned endpoint, a first-passage problem for which no elementary expression exists for the curved boundaries in practical use. The prevailing workaround is the "last-point rule," which inflates the fixed-sample size nzn_z until the marginal rejection probability at the endpoint alone reaches 1β1-\beta. Because this ignores crossings before the endpoint, it is conservative by construction: at (α,1β)=(0.05,0.80)(\alpha, 1-\beta) = (0.05, 0.80) the paper reports empirical power between $0.86$ and $0.88$, i.e., oversizing by seven to nine percentage points.

The paper's contribution is a closed-form correction factor k(α,β,t0)k^*(\alpha, \beta, t_0), where t0=m/nzt_0 = m/n_z is the burn-in fraction, such that setting the total sample size to knzk^* \cdot n_z approximately achieves target AV power without simulation. The factor is expressed entirely in elementary functions and the bivariate normal CDF, requires only two Φ2\Phi_2 evaluations per bracket-search step, and depends on the boundary only through its value b(k)b(k) and slope 1β1-\beta0 at the planned endpoint—so it applies to any smooth concave boundary.

Setup: boundaries and rescaled time

The author considers three widely deployed boundaries under a two-sample model with allocation ratio 1β1-\beta1, standardised effect 1β1-\beta2, and fixed-sample size

1β1-\beta3

WSKR confidence sequence: 1β1-\beta4 with 1β1-\beta5 from the Robbins–Siegmund limiting distribution (Waudby-Smith et al., 2023). Maharaj confidence sequence: an asymptotic Gaussian-mixture supermartingale boundary with tuning constant 1β1-\beta6 involving the Lambert-1β1-\beta7 function. mSPRT: the Johari–Pekelis–Walsh boundary with mixing standard deviation set to the MDE, whose rejection rule coincides with a Bayes factor threshold [johariopre2022].

A key normalisation makes the analysis allocation-agnostic. Defining 1β1-\beta8 gives unit variance for every 1β1-\beta9; under (α,1β)=(0.05,0.80)(\alpha, 1-\beta) = (0.05, 0.80)0 at (α,1β)=(0.05,0.80)(\alpha, 1-\beta) = (0.05, 0.80)1, the rescaled process (α,1β)=(0.05,0.80)(\alpha, 1-\beta) = (0.05, 0.80)2 converges weakly to Brownian motion with drift (α,1β)=(0.05,0.80)(\alpha, 1-\beta) = (0.05, 0.80)3 and variance rate one, started at (α,1β)=(0.05,0.80)(\alpha, 1-\beta) = (0.05, 0.80)4. Consequently (α,1β)=(0.05,0.80)(\alpha, 1-\beta) = (0.05, 0.80)5 depends on (α,1β)=(0.05,0.80)(\alpha, 1-\beta) = (0.05, 0.80)6 only through (α,1β)=(0.05,0.80)(\alpha, 1-\beta) = (0.05, 0.80)7; at equal (α,1β)=(0.05,0.80)(\alpha, 1-\beta) = (0.05, 0.80)8, the factor is identical across allocation ratios. This is confirmed empirically: simulated power at (α,1β)=(0.05,0.80)(\alpha, 1-\beta) = (0.05, 0.80)9 and $0.86$0 matches the $0.86$1 values within Monte Carlo error when $0.86$2 is computed at the $0.86$3-specific $0.86$4.

All three boundaries are concave on $0.86$5—proved analytically for WSKR and Maharaj, verified numerically for mSPRT. Concavity is essential: it guarantees the tangent at the endpoint lies above the boundary, making the linearised crossing probability a lower bound on the true crossing probability and hence making the resulting $0.86$6 slightly conservative rather than anti-conservative.

The corrected factor

The derivation follows Siegmund's classical strategy: linearise the boundary by its tangent at the planned endpoint $0.86$7, apply the Bachelier first-passage formula to Brownian motion with drift over the linear surrogate, then integrate over the random initial value $0.86$8 using a Gaussian convolution identity that reduces each term to a bivariate normal CDF. The resulting approximation decomposes into three interpretable pieces:

$0.86$9

where $0.88$0 is the immediate-rejection probability at burn-in under the tangent surrogate, $0.88$1 is the crossing probability during $0.88$2, and $0.88$3 is the Bachelier reflection correction. The corrected factor $0.88$4 is the smallest root of $0.88$5.

At $0.88$6 and $0.88$7, representative factors are $0.88$8 (WSKR), $0.88$9 (Maharaj), and k(α,β,t0)k^*(\alpha, \beta, t_0)0 (mSPRT), with savings over the last-point rule of roughly k(α,β,t0)k^*(\alpha, \beta, t_0)1–k(α,β,t0)k^*(\alpha, \beta, t_0)2. Across the full grid (k(α,β,t0)k^*(\alpha, \beta, t_0)3, k(α,β,t0)k^*(\alpha, \beta, t_0)4), savings range from about k(α,β,t0)k^*(\alpha, \beta, t_0)5 to k(α,β,t0)k^*(\alpha, \beta, t_0)6. The author cautions that because the three boundaries arise from different mathematical constructions, the differing factors should not be read as an efficiency ranking.

Two caveats attach to uniqueness and monotonicity. Monotonicity of k(α,β,t0)k^*(\alpha, \beta, t_0)7 in k(α,β,t0)k^*(\alpha, \beta, t_0)8 does not follow automatically—the tangent surrogate changes with k(α,β,t0)k^*(\alpha, \beta, t_0)9, so crossing events are not nested—and uniqueness of the root is a numerical observation on the tested grid, not a theorem. Users applying the formula at untabulated parameters are advised to verify concavity of the mSPRT boundary numerically.

Simulation validation

Monte Carlo validation uses t0=m/nzt_0 = m/n_z0 replications per cell (MC SE below t0=m/nzt_0 = m/n_z1) across Gaussian, Bernoulli, and log-normal outcomes, allocation ratios t0=m/nzt_0 = m/n_z2, and effect sizes t0=m/nzt_0 = m/n_z3. Type-I error control is verified first with t0=m/nzt_0 = m/n_z4 null replications; empirical one-sided rejection rates stay below t0=m/nzt_0 = m/n_z5 for all three boundaries at every grid cell.

Empirical power at the corrected sample size lands close to target: at t0=m/nzt_0 = m/n_z6, values cluster around t0=m/nzt_0 = m/n_z7–t0=m/nzt_0 = m/n_z8 versus t0=m/nzt_0 = m/n_z9–knzk^* \cdot n_z0 under last-point sizing. On the extended grid—including Maharaj down to knzk^* \cdot n_z1—the closed form hits target power within approximately knzk^* \cdot n_z2 percentage points, with no degradation at extreme parameters, though savings shrink toward the low-knzk^* \cdot n_z3/high-power corner (e.g., knzk^* \cdot n_z4 at knzk^* \cdot n_z5, knzk^* \cdot n_z6). Against the Monte Carlo reference for knzk^* \cdot n_z7 itself, the closed form agrees within knzk^* \cdot n_z8 relative error (median knzk^* \cdot n_z9) for WSKR and Φ2\Phi_20 (median Φ2\Phi_21) for Maharaj.

Production validation and burn-in sensitivity

On Φ2\Phi_22 metrics drawn from the last 283 experiments on Spotify's Confidence platform (Maharaj boundary with Bonferroni corrections for multiple comparisons and metrics), the median saving is Φ2\Phi_23 and the mean Φ2\Phi_24. This is below the Φ2\Phi_25 at the nominal Φ2\Phi_26 cell because Bonferroni adjustment pushes effective Φ2\Phi_27 and Φ2\Phi_28 toward the lower-saving corner of the grid—an honest reconciliation between laboratory and production numbers.

Sensitivity analysis over burn-ins Φ2\Phi_29 reveals a practically important trade-off. Decreasing b(k)b(k)0 widens the WSKR and Maharaj boundaries through the b(k)b(k)1 term, and this widening more than offsets the longer monitoring window: b(k)b(k)2 increases as b(k)b(k)3 decreases. At realistic burn-ins (b(k)b(k)4), the total budget b(k)b(k)5 is b(k)b(k)6–b(k)b(k)7 the fixed-sample size—considerably more than the factor near b(k)b(k)8 obtained if monitoring begins only at b(k)b(k)9. The mSPRT boundary, which does not depend on 1β1-\beta00, yields 1β1-\beta01 essentially constant (variation under 1β1-\beta02 across 1β1-\beta03). The implication is that the baseline cost of anytime-valid inference with realistic early monitoring is substantially higher than naive extrapolation from late-start monitoring suggests, even after the correction.

Limitations and open questions

Several limitations are stated plainly. The Gaussian approximation replaces the estimated 1β1-\beta04 in the WSKR construction with known 1β1-\beta05, moving from the distribution-free guarantee to a CLT-based one; at low base rates this degrades measurably—with fewer than roughly ten expected successes per arm, empirical power falls below target (gap of 1β1-\beta06 at 1β1-\beta07 and 1β1-\beta08 at 1β1-\beta09). Concavity of the mSPRT boundary is verified only numerically on the tabulated grid. Uniqueness of the root and monotonicity of 1β1-\beta10 remain unproven in general. Finally, the Bernoulli robustness checks involve unequal arm variances under 1β1-\beta11, outside the common-variance model on which the theory rests. Whether a rigorous monotonicity proof, a finite-population or non-Gaussian analogue of the correction, or an extension beyond concave boundaries is possible remains open.

Conclusion

This paper replaces simulation-dependent sizing for always-valid sequential tests with a closed-form factor requiring two bivariate normal CDF evaluations, valid for any smooth concave boundary and any allocation ratio through the single parameter 1β1-\beta12. It corrects a seven-to-nine-percentage-point conservatism in the last-point rule, saves 1β1-\beta13–1β1-\beta14 of the sample budget across the operating range (median 1β1-\beta15 in production), and quantifies the previously underappreciated cost of wide burn-in windows: realistic anytime-valid testing requires budgets two to three times the fixed-sample size.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.