Papers
Topics
Authors
Recent
Search
2000 character limit reached

Sequential Model Confidence Sets

Updated 12 July 2026
  • Sequential model confidence sets are sequential extensions of classical MCS that maintain a random subset of best models with time-uniform, nonasymptotic coverage under repeated peeking.
  • They leverage e-processes and inversions of confidence sequences to ensure validity during continuous monitoring and data-dependent stopping.
  • The framework distinguishes strong, uniformly weak, and weak notions of model superiority, each addressing different inferential objectives in streaming data analysis.

Searching arXiv for papers on sequential model confidence sets and related confidence-sequence/e-process methodology. Sequential model confidence sets are a sequential extension of the classical model confidence set framework of Hansen et al. (2011) for comparing multiple statistical models under streaming data. The objective is to maintain, at each time tt, a random subset of candidate models that still contains the best models with a prescribed confidence level, while allowing continuous monitoring and data-dependent stopping. In contrast to fixed-sample model confidence sets, sequential model confidence sets are designed to remain valid under repeated peeking and optional stopping by relying on e-processes and confidence sequences, thereby providing time-uniform, nonasymptotic coverage guarantees (Arnold et al., 2024).

1. Classical model confidence sets and the sequential motivation

The classical model confidence set (MCS) starts from a finite collection of competing models indexed by M0={1,,m}M_0=\{1,\dots,m\}. For a fixed sample of size TT, model ii incurs losses Li,1,,Li,TL_{i,1},\dots,L_{i,T} under a loss function L(,)L(\cdot,\cdot), with smaller values preferred. Pairwise loss differences are defined by

dij,t=Li,tLj,t,dˉij=1Tt=1Tdij,t.d_{ij,t}=L_{i,t}-L_{j,t}, \qquad \bar d_{ij}=\frac1T\sum_{t=1}^T d_{ij,t}.

Hansen et al. test, for each model ii, the null hypothesis

Hi:maxjidˉij0.H_i:\max_{j\neq i}\bar d_{ij}\le 0.

Using pairwise tt-tests with bootstrap-estimated variances together with a step-down multiple-testing procedure, the method returns a random subset M0={1,,m}M_0=\{1,\dots,m\}0 satisfying

M0={1,,m}M_0=\{1,\dots,m\}1

where M0={1,,m}M_0=\{1,\dots,m\}2 is the set of best models, namely those minimizing M0={1,,m}M_0=\{1,\dots,m\}3 (Arnold et al., 2024).

The central limitation of the classical MCS is that its validity is tied to a fixed sample size M0={1,,m}M_0=\{1,\dots,m\}4. This conflicts with many forecasting and model-evaluation settings in which data arrive one observation at a time, interim analyses are routine, and the decision to continue or stop depends on previous outcomes. A direct application of the static MCS at each M0={1,,m}M_0=\{1,\dots,m\}5 would not preserve the overall error guarantee, because repeated peeking or optional stopping inflates the probability of excluding a truly best model. Sequential model confidence sets address this by targeting the time-uniform guarantee

M0={1,,m}M_0=\{1,\dots,m\}6

where M0={1,,m}M_0=\{1,\dots,m\}7 denotes the set of best models up to time M0={1,,m}M_0=\{1,\dots,m\}8 (Arnold et al., 2024).

This establishes the essential distinction between static and sequential formulations: the latter are not merely repeated fixed-sample procedures, but objects constructed to be valid under indefinite monitoring.

2. E-processes, confidence sequences, and time-uniform validity

The sequential construction is built on e-processes and confidence sequences. Let M0={1,,m}M_0=\{1,\dots,m\}9 be a filtered probability space, and let the loss differences TT0 be TT1-measurable. Their conditional expectations are denoted

TT2

An e-process for a composite null hypothesis TT3 is a nonnegative adapted process TT4 such that for every law TT5 and every stopping time TT6,

TT7

Equivalently, TT8 is upper-bounded by a test supermartingale under each TT9. By Ville’s inequality, any e-process yields a sequential test at level ii0 through the rejection rule

ii1

with

ii2

(Arnold et al., 2024).

A time-uniform ii3-confidence sequence for a scalar parameter process ii4 is a sequence of random sets ii5 such that

ii6

for every law ii7. One way to obtain such sets is to invert a family of e-processes ii8 over ii9: Li,1,,Li,TL_{i,1},\dots,L_{i,T}0 Coverage then follows from Ville’s inequality (Arnold et al., 2024).

The broader confidence-sequence literature formalizes this time-uniform perspective as an anytime-valid sequential inference primitive. In particular, a confidence sequence is a sequence of adapted sets for a predictable parameter sequence with a time-uniform coverage guarantee (Mineiro, 2022). This suggests that sequential model confidence sets can be understood as a multiple-model, model-comparison instantiation of the same general inferential principle.

A common misconception is that sequential validity can be recovered by simply applying a fixed-sample method repeatedly with adjusted significance levels. The SMCS formulation instead uses processes whose validity is stable under optional stopping by construction, which is a substantially stronger requirement than pointwise validity at each time (Arnold et al., 2024).

3. Formal definitions of superiority and the three SMCS targets

The framework distinguishes several notions of “best” model over time. These notions lead to different sequential model confidence sets and correspond to different scientific questions.

For the strongly superior notion, model Li,1,,Li,TL_{i,1},\dots,L_{i,T}1 is retained if it is not conditionally worse than any competitor at every previous time: Li,1,,Li,TL_{i,1},\dots,L_{i,T}2 The inferential target is then a sequence Li,1,,Li,TL_{i,1},\dots,L_{i,T}3 satisfying

Li,1,,Li,TL_{i,1},\dots,L_{i,T}4

(Arnold et al., 2024).

For the uniformly weakly superior notion, comparison is based on running averages of conditional expected losses. Define

Li,1,,Li,TL_{i,1},\dots,L_{i,T}5

and set

Li,1,,Li,TL_{i,1},\dots,L_{i,T}6

This criterion is weaker than the strong version because it allows temporary inferiority, provided the cumulative average performance remains competitive at all times (Arnold et al., 2024).

For the weakly superior at Li,1,,Li,TL_{i,1},\dots,L_{i,T}7 notion, the requirement is imposed only at the current horizon: Li,1,,Li,TL_{i,1},\dots,L_{i,T}8 In this case the identity of the best models may change over time, and the procedure permits inclusion, exclusion, and re-inclusion of models as performance evolves (Arnold et al., 2024).

These three targets are not interchangeable. In one simulation described in the source paper, a forecaster is slightly worse every seventh day but better on average; the strong SMCS excludes it, whereas the uniformly weak SMCS never excludes it, reflecting the fact that the two procedures answer different questions (Arnold et al., 2024). A plausible implication is that the appropriate SMCS definition depends less on statistical convenience than on whether the application prioritizes per-round dominance, uniformly good cumulative behavior, or current cumulative standing.

4. Construction of sequential model confidence sets

Strong hypothesis construction

Assume the pairwise loss differences satisfy Li,1,,Li,TL_{i,1},\dots,L_{i,T}9, where L(,)L(\cdot,\cdot)0 is predictable. For any predictable sequence L(,)L(\cdot,\cdot)1, define

L(,)L(\cdot,\cdot)2

This is an e-process for the pairwise null

L(,)L(\cdot,\cdot)3

To test whether model L(,)L(\cdot,\cdot)4 is strong against all competitors, merge the pairwise e-processes over L(,)L(\cdot,\cdot)5: L(,)L(\cdot,\cdot)6 The arithmetic mean remains an e-process for the intersection null L(,)L(\cdot,\cdot)7. A closure-principle adjustment is then applied to the family L(,)L(\cdot,\cdot)8, producing L(,)L(\cdot,\cdot)9 that control family-wise error simultaneously over dij,t=Li,tLj,t,dˉij=1Tt=1Tdij,t.d_{ij,t}=L_{i,t}-L_{j,t}, \qquad \bar d_{ij}=\frac1T\sum_{t=1}^T d_{ij,t}.0. The strong sequential model confidence set is

dij,t=Li,tLj,t,dˉij=1Tt=1Tdij,t.d_{ij,t}=L_{i,t}-L_{j,t}, \qquad \bar d_{ij}=\frac1T\sum_{t=1}^T d_{ij,t}.1

Under this construction,

dij,t=Li,tLj,t,dˉij=1Tt=1Tdij,t.d_{ij,t}=L_{i,t}-L_{j,t}, \qquad \bar d_{ij}=\frac1T\sum_{t=1}^T d_{ij,t}.2

for any data-generating law dij,t=Li,tLj,t,dˉij=1Tt=1Tdij,t.d_{ij,t}=L_{i,t}-L_{j,t}, \qquad \bar d_{ij}=\frac1T\sum_{t=1}^T d_{ij,t}.3 (Arnold et al., 2024).

Uniformly weak hypothesis construction

Under dij,t=Li,tLj,t,dˉij=1Tt=1Tdij,t.d_{ij,t}=L_{i,t}-L_{j,t}, \qquad \bar d_{ij}=\frac1T\sum_{t=1}^T d_{ij,t}.4, and for any fixed dij,t=Li,tLj,t,dˉij=1Tt=1Tdij,t.d_{ij,t}=L_{i,t}-L_{j,t}, \qquad \bar d_{ij}=\frac1T\sum_{t=1}^T d_{ij,t}.5, an empirical-Bernstein-type e-process can be constructed as

dij,t=Li,tLj,t,dˉij=1Tt=1Tdij,t.d_{ij,t}=L_{i,t}-L_{j,t}, \qquad \bar d_{ij}=\frac1T\sum_{t=1}^T d_{ij,t}.6

where

dij,t=Li,tLj,t,dˉij=1Tt=1Tdij,t.d_{ij,t}=L_{i,t}-L_{j,t}, \qquad \bar d_{ij}=\frac1T\sum_{t=1}^T d_{ij,t}.7

This process is valid for

dij,t=Li,tLj,t,dˉij=1Tt=1Tdij,t.d_{ij,t}=L_{i,t}-L_{j,t}, \qquad \bar d_{ij}=\frac1T\sum_{t=1}^T d_{ij,t}.8

After merging over dij,t=Li,tLj,t,dˉij=1Tt=1Tdij,t.d_{ij,t}=L_{i,t}-L_{j,t}, \qquad \bar d_{ij}=\frac1T\sum_{t=1}^T d_{ij,t}.9 and applying closure, the uniformly weak SMCS is again

ii0

with time-uniform coverage (Arnold et al., 2024).

Weak hypothesis construction via confidence sequences

For the weak target, the procedure uses lower confidence bounds ii1 on ii2, obtained by inverting the empirical-Bernstein-type e-process

ii3

These lower bounds satisfy

ii4

After a union-over-pairs adjustment, one defines

ii5

This gives a time-uniform SMCS for the weak notion of superiority (Arnold et al., 2024).

The architecture is therefore modular: pairwise e-process construction, aggregation across competitors, multiplicity correction, and set inversion.

5. Algorithmic form, computational trade-offs, and practical use

For the strong hypothesis, the source paper gives high-level pseudocode:

  1. initialize ii6 for all ii7;
  2. at each round ii8, compute all pairwise loss differences ii9;
  3. choose predictable steps Hi:maxjidˉij0.H_i:\max_{j\neq i}\bar d_{ij}\le 0.0;
  4. update Hi:maxjidˉij0.H_i:\max_{j\neq i}\bar d_{ij}\le 0.1;
  5. merge over Hi:maxjidˉij0.H_i:\max_{j\neq i}\bar d_{ij}\le 0.2 using Hi:maxjidˉij0.H_i:\max_{j\neq i}\bar d_{ij}\le 0.3;
  6. apply closure to obtain Hi:maxjidˉij0.H_i:\max_{j\neq i}\bar d_{ij}\le 0.4;
  7. output Hi:maxjidˉij0.H_i:\max_{j\neq i}\bar d_{ij}\le 0.5 (Arnold et al., 2024).

Each round requires Hi:maxjidˉij0.H_i:\max_{j\neq i}\bar d_{ij}\le 0.6 pairwise updates plus the cost of closure. The paper notes that best-possible closure is exponential in Hi:maxjidˉij0.H_i:\max_{j\neq i}\bar d_{ij}\le 0.7, but simpler stepdown or step-up procedures can often be used in Hi:maxjidˉij0.H_i:\max_{j\neq i}\bar d_{ij}\le 0.8. For moderate Hi:maxjidˉij0.H_i:\max_{j\neq i}\bar d_{ij}\le 0.9, such as dozens of models, the method is described as feasible in real time (Arnold et al., 2024).

The same source discusses tuning of the betting parameters tt0. A mixing or mixture over tt1 is theoretically attractive but can be computationally expensive. In practice, one may take tt2 with a data-driven tt3, or use the default tt4 (Arnold et al., 2024).

This algorithmic pattern aligns with a broader theme in sequential inference: valid sequential uncertainty quantification often emerges from test supermartingales or likelihood-ratio constructions that are then inverted into sets. Related work constructs anytime-valid confidence sequences from sequential likelihood ratios (Emmenegger et al., 2023), sequential likelihood mixing (Kirschner et al., 20 Feb 2025), or online-to-confidence-set conversions via regret bounds (Clerico et al., 23 Apr 2025). This suggests a general methodological analogy: SMCS uses e-processes over pairwise model comparisons in the same way that parameter confidence sequences use martingales or supermartingales over parameter values.

6. Interpretation, misconceptions, and relation to adjacent literatures

The interpretation of tt5 is operational rather than purely descriptive: it is the set of models that could still be the best in an anytime-valid sense. If tt6, one may stop and declare a unique winner while controlling the overall type-I error of wrongly excluding a truly best model at level tt7. If tt8 remains large, the evidence is weak and additional data are needed (Arnold et al., 2024).

A common misunderstanding is to treat SMCS as a ranking procedure. The framework instead produces a set-valued inferential object whose size is flexible and driven by the amount of statistically defensible separation among models. This flexibility is inherited from the classical MCS, where the set is allowed to contain multiple best models (Arnold et al., 2024).

Another misunderstanding is that there is a single canonical notion of “best.” The distinction among strong, uniformly weak, and weak superiority shows that the answer depends on whether performance is evaluated pointwise in time, uniformly through running averages, or only at the current horizon. The simulation examples make clear that different SMCS variants can legitimately retain or exclude the same model under the same data because they encode different null hypotheses (Arnold et al., 2024).

The relation to confidence-sequence methodology is direct. Confidence sequences are defined as adapted sets that cover a predictable parameter sequence uniformly over time (Mineiro, 2022). In SMCS, the inferential target is not a scalar parameter but a subset of a finite model class, and the construction combines pairwise sequential tests with multiplicity correction. A plausible implication is that SMCS occupies a boundary zone between sequential estimation and sequential multiple testing: it is simultaneously a confidence-sequence construction, an e-value aggregation procedure, and a sequential model-elimination rule.

The framework also sits adjacent to several modern lines of work on sequential confidence sets. Sequential kernel regression uses martingale tail inequalities to obtain tighter confidence bounds and then plug them into KernelUCB (Flynn et al., 2024). Likelihood-ratio confidence sets provide anytime-valid sets for parametric and RKHS settings under adaptive data collection (Emmenegger et al., 2023). Sequential likelihood mixing further gives a universal recipe for confidence sequences under arbitrary conditional density models and even approximate inference schemes (Kirschner et al., 20 Feb 2025). These developments do not define SMCS themselves, but they reinforce the same foundational principle: time-uniform validity is achieved by martingale-based or supermartingale-based evidence accumulation, not by repeated fixed-horizon recalibration.

7. Scope, applications, and extensions

The intended applications include forecasting, online model evaluation, A/B testing, and any setting with streaming loss data. Because the procedure only requires loss observations and suitable e-process constructions, it is agnostic to the internal structure of the models being compared; what matters is the sequential behavior of their losses (Arnold et al., 2024).

The source paper emphasizes several practical and methodological extensions. The same machinery can be extended to unbounded losses via parametric e-processes, to alternative multiple-comparison criteria, and to sequential selection among infinitely many models via e-doob martingales or betting on model indices (Arnold et al., 2024). It also notes that false discovery rate control can be pursued through the e-BH procedure of Wang and Ramdas (2022), which may yield larger SMCS but controls the expected proportion of wrong exclusions rather than the family-wise error rate (Arnold et al., 2024).

In this sense, SMCS is not a single fixed algorithm but a family of sequential set-valued procedures built from three ingredients: a notion of superiority, valid pairwise e-processes or confidence sequences, and a multiplicity-control mechanism. Its principal contribution is to transfer the model-confidence-set idea from fixed-horizon inference to an anytime-valid sequential regime while preserving flexible set size and nonasymptotic guarantees (Arnold et al., 2024).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Sequential Model Confidence Sets.