Sequential Model Confidence Sets
- Sequential model confidence sets are sequential extensions of classical MCS that maintain a random subset of best models with time-uniform, nonasymptotic coverage under repeated peeking.
- They leverage e-processes and inversions of confidence sequences to ensure validity during continuous monitoring and data-dependent stopping.
- The framework distinguishes strong, uniformly weak, and weak notions of model superiority, each addressing different inferential objectives in streaming data analysis.
Searching arXiv for papers on sequential model confidence sets and related confidence-sequence/e-process methodology. Sequential model confidence sets are a sequential extension of the classical model confidence set framework of Hansen et al. (2011) for comparing multiple statistical models under streaming data. The objective is to maintain, at each time , a random subset of candidate models that still contains the best models with a prescribed confidence level, while allowing continuous monitoring and data-dependent stopping. In contrast to fixed-sample model confidence sets, sequential model confidence sets are designed to remain valid under repeated peeking and optional stopping by relying on e-processes and confidence sequences, thereby providing time-uniform, nonasymptotic coverage guarantees (Arnold et al., 2024).
1. Classical model confidence sets and the sequential motivation
The classical model confidence set (MCS) starts from a finite collection of competing models indexed by . For a fixed sample of size , model incurs losses under a loss function , with smaller values preferred. Pairwise loss differences are defined by
Hansen et al. test, for each model , the null hypothesis
Using pairwise -tests with bootstrap-estimated variances together with a step-down multiple-testing procedure, the method returns a random subset 0 satisfying
1
where 2 is the set of best models, namely those minimizing 3 (Arnold et al., 2024).
The central limitation of the classical MCS is that its validity is tied to a fixed sample size 4. This conflicts with many forecasting and model-evaluation settings in which data arrive one observation at a time, interim analyses are routine, and the decision to continue or stop depends on previous outcomes. A direct application of the static MCS at each 5 would not preserve the overall error guarantee, because repeated peeking or optional stopping inflates the probability of excluding a truly best model. Sequential model confidence sets address this by targeting the time-uniform guarantee
6
where 7 denotes the set of best models up to time 8 (Arnold et al., 2024).
This establishes the essential distinction between static and sequential formulations: the latter are not merely repeated fixed-sample procedures, but objects constructed to be valid under indefinite monitoring.
2. E-processes, confidence sequences, and time-uniform validity
The sequential construction is built on e-processes and confidence sequences. Let 9 be a filtered probability space, and let the loss differences 0 be 1-measurable. Their conditional expectations are denoted
2
An e-process for a composite null hypothesis 3 is a nonnegative adapted process 4 such that for every law 5 and every stopping time 6,
7
Equivalently, 8 is upper-bounded by a test supermartingale under each 9. By Ville’s inequality, any e-process yields a sequential test at level 0 through the rejection rule
1
with
2
A time-uniform 3-confidence sequence for a scalar parameter process 4 is a sequence of random sets 5 such that
6
for every law 7. One way to obtain such sets is to invert a family of e-processes 8 over 9: 0 Coverage then follows from Ville’s inequality (Arnold et al., 2024).
The broader confidence-sequence literature formalizes this time-uniform perspective as an anytime-valid sequential inference primitive. In particular, a confidence sequence is a sequence of adapted sets for a predictable parameter sequence with a time-uniform coverage guarantee (Mineiro, 2022). This suggests that sequential model confidence sets can be understood as a multiple-model, model-comparison instantiation of the same general inferential principle.
A common misconception is that sequential validity can be recovered by simply applying a fixed-sample method repeatedly with adjusted significance levels. The SMCS formulation instead uses processes whose validity is stable under optional stopping by construction, which is a substantially stronger requirement than pointwise validity at each time (Arnold et al., 2024).
3. Formal definitions of superiority and the three SMCS targets
The framework distinguishes several notions of “best” model over time. These notions lead to different sequential model confidence sets and correspond to different scientific questions.
For the strongly superior notion, model 1 is retained if it is not conditionally worse than any competitor at every previous time: 2 The inferential target is then a sequence 3 satisfying
4
For the uniformly weakly superior notion, comparison is based on running averages of conditional expected losses. Define
5
and set
6
This criterion is weaker than the strong version because it allows temporary inferiority, provided the cumulative average performance remains competitive at all times (Arnold et al., 2024).
For the weakly superior at 7 notion, the requirement is imposed only at the current horizon: 8 In this case the identity of the best models may change over time, and the procedure permits inclusion, exclusion, and re-inclusion of models as performance evolves (Arnold et al., 2024).
These three targets are not interchangeable. In one simulation described in the source paper, a forecaster is slightly worse every seventh day but better on average; the strong SMCS excludes it, whereas the uniformly weak SMCS never excludes it, reflecting the fact that the two procedures answer different questions (Arnold et al., 2024). A plausible implication is that the appropriate SMCS definition depends less on statistical convenience than on whether the application prioritizes per-round dominance, uniformly good cumulative behavior, or current cumulative standing.
4. Construction of sequential model confidence sets
Strong hypothesis construction
Assume the pairwise loss differences satisfy 9, where 0 is predictable. For any predictable sequence 1, define
2
This is an e-process for the pairwise null
3
To test whether model 4 is strong against all competitors, merge the pairwise e-processes over 5: 6 The arithmetic mean remains an e-process for the intersection null 7. A closure-principle adjustment is then applied to the family 8, producing 9 that control family-wise error simultaneously over 0. The strong sequential model confidence set is
1
Under this construction,
2
for any data-generating law 3 (Arnold et al., 2024).
Uniformly weak hypothesis construction
Under 4, and for any fixed 5, an empirical-Bernstein-type e-process can be constructed as
6
where
7
This process is valid for
8
After merging over 9 and applying closure, the uniformly weak SMCS is again
0
with time-uniform coverage (Arnold et al., 2024).
Weak hypothesis construction via confidence sequences
For the weak target, the procedure uses lower confidence bounds 1 on 2, obtained by inverting the empirical-Bernstein-type e-process
3
These lower bounds satisfy
4
After a union-over-pairs adjustment, one defines
5
This gives a time-uniform SMCS for the weak notion of superiority (Arnold et al., 2024).
The architecture is therefore modular: pairwise e-process construction, aggregation across competitors, multiplicity correction, and set inversion.
5. Algorithmic form, computational trade-offs, and practical use
For the strong hypothesis, the source paper gives high-level pseudocode:
- initialize 6 for all 7;
- at each round 8, compute all pairwise loss differences 9;
- choose predictable steps 0;
- update 1;
- merge over 2 using 3;
- apply closure to obtain 4;
- output 5 (Arnold et al., 2024).
Each round requires 6 pairwise updates plus the cost of closure. The paper notes that best-possible closure is exponential in 7, but simpler stepdown or step-up procedures can often be used in 8. For moderate 9, such as dozens of models, the method is described as feasible in real time (Arnold et al., 2024).
The same source discusses tuning of the betting parameters 0. A mixing or mixture over 1 is theoretically attractive but can be computationally expensive. In practice, one may take 2 with a data-driven 3, or use the default 4 (Arnold et al., 2024).
This algorithmic pattern aligns with a broader theme in sequential inference: valid sequential uncertainty quantification often emerges from test supermartingales or likelihood-ratio constructions that are then inverted into sets. Related work constructs anytime-valid confidence sequences from sequential likelihood ratios (Emmenegger et al., 2023), sequential likelihood mixing (Kirschner et al., 20 Feb 2025), or online-to-confidence-set conversions via regret bounds (Clerico et al., 23 Apr 2025). This suggests a general methodological analogy: SMCS uses e-processes over pairwise model comparisons in the same way that parameter confidence sequences use martingales or supermartingales over parameter values.
6. Interpretation, misconceptions, and relation to adjacent literatures
The interpretation of 5 is operational rather than purely descriptive: it is the set of models that could still be the best in an anytime-valid sense. If 6, one may stop and declare a unique winner while controlling the overall type-I error of wrongly excluding a truly best model at level 7. If 8 remains large, the evidence is weak and additional data are needed (Arnold et al., 2024).
A common misunderstanding is to treat SMCS as a ranking procedure. The framework instead produces a set-valued inferential object whose size is flexible and driven by the amount of statistically defensible separation among models. This flexibility is inherited from the classical MCS, where the set is allowed to contain multiple best models (Arnold et al., 2024).
Another misunderstanding is that there is a single canonical notion of “best.” The distinction among strong, uniformly weak, and weak superiority shows that the answer depends on whether performance is evaluated pointwise in time, uniformly through running averages, or only at the current horizon. The simulation examples make clear that different SMCS variants can legitimately retain or exclude the same model under the same data because they encode different null hypotheses (Arnold et al., 2024).
The relation to confidence-sequence methodology is direct. Confidence sequences are defined as adapted sets that cover a predictable parameter sequence uniformly over time (Mineiro, 2022). In SMCS, the inferential target is not a scalar parameter but a subset of a finite model class, and the construction combines pairwise sequential tests with multiplicity correction. A plausible implication is that SMCS occupies a boundary zone between sequential estimation and sequential multiple testing: it is simultaneously a confidence-sequence construction, an e-value aggregation procedure, and a sequential model-elimination rule.
The framework also sits adjacent to several modern lines of work on sequential confidence sets. Sequential kernel regression uses martingale tail inequalities to obtain tighter confidence bounds and then plug them into KernelUCB (Flynn et al., 2024). Likelihood-ratio confidence sets provide anytime-valid sets for parametric and RKHS settings under adaptive data collection (Emmenegger et al., 2023). Sequential likelihood mixing further gives a universal recipe for confidence sequences under arbitrary conditional density models and even approximate inference schemes (Kirschner et al., 20 Feb 2025). These developments do not define SMCS themselves, but they reinforce the same foundational principle: time-uniform validity is achieved by martingale-based or supermartingale-based evidence accumulation, not by repeated fixed-horizon recalibration.
7. Scope, applications, and extensions
The intended applications include forecasting, online model evaluation, A/B testing, and any setting with streaming loss data. Because the procedure only requires loss observations and suitable e-process constructions, it is agnostic to the internal structure of the models being compared; what matters is the sequential behavior of their losses (Arnold et al., 2024).
The source paper emphasizes several practical and methodological extensions. The same machinery can be extended to unbounded losses via parametric e-processes, to alternative multiple-comparison criteria, and to sequential selection among infinitely many models via e-doob martingales or betting on model indices (Arnold et al., 2024). It also notes that false discovery rate control can be pursued through the e-BH procedure of Wang and Ramdas (2022), which may yield larger SMCS but controls the expected proportion of wrong exclusions rather than the family-wise error rate (Arnold et al., 2024).
In this sense, SMCS is not a single fixed algorithm but a family of sequential set-valued procedures built from three ingredients: a notion of superiority, valid pairwise e-processes or confidence sequences, and a multiplicity-control mechanism. Its principal contribution is to transfer the model-confidence-set idea from fixed-horizon inference to an anytime-valid sequential regime while preserving flexible set size and nonasymptotic guarantees (Arnold et al., 2024).