Papers
Topics
Authors
Recent
Search
2000 character limit reached

Shtarkov Solution in Minimax Prediction

Updated 7 July 2026
  • Shtarkov Solution is the normalized maximum likelihood distribution that achieves the exact minimax log-loss regret by normalizing pointwise maximum likelihoods.
  • It underpins sequential prediction, universal coding, and MDL by converting worst-case decision problems into effective probabilistic forecasts.
  • Recent extensions adapt the framework to adversarial contexts and use geometric entropy bounds to link redundancy with scale-sensitive complexity.

Searching arXiv for papers on the Shtarkov solution and related minimax log-loss regret. arXiv search query: "Shtarkov solution normalized maximum likelihood minimax regret sequential probability assignment" The Shtarkov solution is the exact minimax solution to sequential probability assignment under logarithmic loss in the context-free setting. For a class of sequence distributions QΔ(Yn)Q\subset\Delta(Y^n), it is the Normalized Maximum Likelihood (NML) distribution

p(y)=supqQq(y)Sn(Q),Sn(Q)yYnsupqQq(y),p^*(y)=\frac{\sup_{q\in Q} q(y)}{S_n(Q)}, \qquad S_n(Q)\coloneqq \sum_{y\in Y^n}\sup_{q\in Q} q(y),

and the corresponding minimax regret is exactly logSn(Q)\log S_n(Q) (Jia et al., 22 Mar 2025). In parametric notation, the same construction is written

qNML(y)=p(yθ^(y))Cn,θ^(y)=argmaxθp(yθ),q_{\rm NML}(y)=\frac{p(y\mid \hat\theta(y))}{C_n}, \qquad \hat\theta(y)=\arg\max_\theta p(y\mid\theta),

with CnC_n the Shtarkov sum (Clarke et al., 28 Jul 2025). The construction occupies a central position in online log-loss prediction, universal coding, and MDL, and recent work extends it from the classical context-free setting to adversarial contextual forecasting and to entropy-based characterizations of minimax regret.

1. Classical minimax characterization

In the non-contextual formulation, one fixes a finite alphabet YY and considers sequences y=(y1,,yn)Yny=(y_1,\dots,y_n)\in Y^n. A forecaster chooses pΔ(Yn)p\in\Delta(Y^n) and incurs log-loss logp(y)-\log p(y). The worst-case regret against a reference class QΔ(Yn)Q\subset\Delta(Y^n) is

p(y)=supqQq(y)Sn(Q),Sn(Q)yYnsupqQq(y),p^*(y)=\frac{\sup_{q\in Q} q(y)}{S_n(Q)}, \qquad S_n(Q)\coloneqq \sum_{y\in Y^n}\sup_{q\in Q} q(y),0

Shtarkov’s theorem shows that this optimization has a closed form:

p(y)=supqQq(y)Sn(Q),Sn(Q)yYnsupqQq(y),p^*(y)=\frac{\sup_{q\in Q} q(y)}{S_n(Q)}, \qquad S_n(Q)\coloneqq \sum_{y\in Y^n}\sup_{q\in Q} q(y),1

and the minimizing distribution is the NML predictor p(y)=supqQq(y)Sn(Q),Sn(Q)yYnsupqQq(y),p^*(y)=\frac{\sup_{q\in Q} q(y)}{S_n(Q)}, \qquad S_n(Q)\coloneqq \sum_{y\in Y^n}\sup_{q\in Q} q(y),2 defined by normalizing p(y)=supqQq(y)Sn(Q),Sn(Q)yYnsupqQq(y),p^*(y)=\frac{\sup_{q\in Q} q(y)}{S_n(Q)}, \qquad S_n(Q)\coloneqq \sum_{y\in Y^n}\sup_{q\in Q} q(y),3 over all sequences (Jia et al., 22 Mar 2025).

An equivalent parametric presentation starts from a family of joint distributions p(y)=supqQq(y)Sn(Q),Sn(Q)yYnsupqQq(y),p^*(y)=\frac{\sup_{q\in Q} q(y)}{S_n(Q)}, \qquad S_n(Q)\coloneqq \sum_{y\in Y^n}\sup_{q\in Q} q(y),4 over sequences p(y)=supqQq(y)Sn(Q),Sn(Q)yYnsupqQq(y),p^*(y)=\frac{\sup_{q\in Q} q(y)}{S_n(Q)}, \qquad S_n(Q)\coloneqq \sum_{y\in Y^n}\sup_{q\in Q} q(y),5. The sequence-wise regret of a predictor p(y)=supqQq(y)Sn(Q),Sn(Q)yYnsupqQq(y),p^*(y)=\frac{\sup_{q\in Q} q(y)}{S_n(Q)}, \qquad S_n(Q)\coloneqq \sum_{y\in Y^n}\sup_{q\in Q} q(y),6 is

p(y)=supqQq(y)Sn(Q),Sn(Q)yYnsupqQq(y),p^*(y)=\frac{\sup_{q\in Q} q(y)}{S_n(Q)}, \qquad S_n(Q)\coloneqq \sum_{y\in Y^n}\sup_{q\in Q} q(y),7

and the minimax value is

p(y)=supqQq(y)Sn(Q),Sn(Q)yYnsupqQq(y),p^*(y)=\frac{\sup_{q\in Q} q(y)}{S_n(Q)}, \qquad S_n(Q)\coloneqq \sum_{y\in Y^n}\sup_{q\in Q} q(y),8

In this notation, Shtarkov’s theorem gives

p(y)=supqQq(y)Sn(Q),Sn(Q)yYnsupqQq(y),p^*(y)=\frac{\sup_{q\in Q} q(y)}{S_n(Q)}, \qquad S_n(Q)\coloneqq \sum_{y\in Y^n}\sup_{q\in Q} q(y),9

with logSn(Q)\log S_n(Q)0 (Liu et al., 2024).

The significance of the theorem is that it converts a worst-case sequential decision problem into the normalization of a pointwise maximum likelihood. The Shtarkov sum is therefore both a coding-theoretic normalizer and the exact minimax individual-sequence log-loss regret.

2. Normalized maximum likelihood and the equalized-regret property

For log-loss, the NML construction has a distinctive equalization property. In the generalized form studied in learning theory, the strategy

logSn(Q)\log S_n(Q)1

equalizes the regret over all logSn(Q)\log S_n(Q)2, and its constant value is the logarithm of the corresponding Shtarkov integral (Grünwald et al., 2017). In the ordinary log-loss specialization with logSn(Q)\log S_n(Q)3 and logSn(Q)\log S_n(Q)4, this recovers the classical NML density.

The same point is expressed in coding-length language in the parametric streaming formulation. Under the logarithmic score, a forecaster using density logSn(Q)\log S_n(Q)5 pays a code-length of logSn(Q)\log S_n(Q)6 nats, whereas the best expert at logSn(Q)\log S_n(Q)7 pays logSn(Q)\log S_n(Q)8. The NML predictor achieves the minimax worst-case excess code-length

logSn(Q)\log S_n(Q)9

and no other qNML(y)=p(yθ^(y))Cn,θ^(y)=argmaxθp(yθ),q_{\rm NML}(y)=\frac{p(y\mid \hat\theta(y))}{C_n}, \qquad \hat\theta(y)=\arg\max_\theta p(y\mid\theta),0 can guarantee smaller maximum regret (Clarke et al., 28 Jul 2025).

This characterization is compatible with sequential one-step-ahead prediction. If

qNML(y)=p(yθ^(y))Cn,θ^(y)=argmaxθp(yθ),q_{\rm NML}(y)=\frac{p(y\mid \hat\theta(y))}{C_n}, \qquad \hat\theta(y)=\arg\max_\theta p(y\mid\theta),1

then the one-step-ahead density is

qNML(y)=p(yθ^(y))Cn,θ^(y)=argmaxθp(yθ),q_{\rm NML}(y)=\frac{p(y\mid \hat\theta(y))}{C_n}, \qquad \hat\theta(y)=\arg\max_\theta p(y\mid\theta),2

In streaming prediction, the minimax guarantee holds at each qNML(y)=p(yθ^(y))Cn,θ^(y)=argmaxθp(yθ),q_{\rm NML}(y)=\frac{p(y\mid \hat\theta(y))}{C_n}, \qquad \hat\theta(y)=\arg\max_\theta p(y\mid\theta),3 (Clarke et al., 28 Jul 2025).

A common misunderstanding is that any low-regret expert-advice method is equivalent to the Shtarkov solution. The available results distinguish the notions sharply: exponential-weights forecasters can bound regret in expectation or with high probability under worst-case sequences, but they do not in general achieve the absolute minimax code-length (Clarke et al., 28 Jul 2025).

3. Contextual and adversarial sequential extensions

Recent work extends the Shtarkov construction to the fully contextual online setting. Here, at each round qNML(y)=p(yθ^(y))Cn,θ^(y)=argmaxθp(yθ),q_{\rm NML}(y)=\frac{p(y\mid \hat\theta(y))}{C_n}, \qquad \hat\theta(y)=\arg\max_\theta p(y\mid\theta),4, the learner observes a context qNML(y)=p(yθ^(y))Cn,θ^(y)=argmaxθp(yθ),q_{\rm NML}(y)=\frac{p(y\mid \hat\theta(y))}{C_n}, \qquad \hat\theta(y)=\arg\max_\theta p(y\mid\theta),5 that may depend adversarially on past labels qNML(y)=p(yθ^(y))Cn,θ^(y)=argmaxθp(yθ),q_{\rm NML}(y)=\frac{p(y\mid \hat\theta(y))}{C_n}, \qquad \hat\theta(y)=\arg\max_\theta p(y\mid\theta),6. Equivalently, one fixes a qNML(y)=p(yθ^(y))Cn,θ^(y)=argmaxθp(yθ),q_{\rm NML}(y)=\frac{p(y\mid \hat\theta(y))}{C_n}, \qquad \hat\theta(y)=\arg\max_\theta p(y\mid\theta),7-ary context tree

qNML(y)=p(yθ^(y))Cn,θ^(y)=argmaxθp(yθ),q_{\rm NML}(y)=\frac{p(y\mid \hat\theta(y))}{C_n}, \qquad \hat\theta(y)=\arg\max_\theta p(y\mid\theta),8

where each qNML(y)=p(yθ^(y))Cn,θ^(y)=argmaxθp(yθ),q_{\rm NML}(y)=\frac{p(y\mid \hat\theta(y))}{C_n}, \qquad \hat\theta(y)=\arg\max_\theta p(y\mid\theta),9. An expert CnC_n0 is a sequence of conditional probability mappings

CnC_n1

with joint likelihood

CnC_n2

The contextual Shtarkov sum for a fixed context tree is then

CnC_n3

The minimax regret against CnC_n4 and adaptive contexts is

CnC_n5

(Liu et al., 2024).

The corresponding minimax-optimal strategy is contextual Normalized Maximum Likelihood (cNML). At time CnC_n6, after past labels CnC_n7 and context CnC_n8, it predicts

CnC_n9

These fractions form a valid probability mass function over YY0, and by construction the strategy achieves the value YY1 (Liu et al., 2024).

The contextual extension is technically significant because the usual sequential YY2 entropy does not characterize minimax risk in general, whereas the contextual Shtarkov sum does. The framework also applies to general finite label alphabets, not only binary labels, and allows expert classes of mappings from YY3, including nonparametric and combinatorial classes (Liu et al., 2024).

4. Sequential square-root entropy and geometric upper bounds

A major recent development is the control of the Shtarkov sum by a geometric complexity defined through Hellinger-type sequential covers. For YY4, a finite set YY5 is an YY6 sequential square-root YY7-cover if for every YY8 and every history YY9 one can choose y=(y1,,yn)Yny=(y_1,\dots,y_n)\in Y^n0 such that, at each time y=(y1,,yn)Yny=(y_1,\dots,y_n)\in Y^n1 and symbol y=(y1,,yn)Yny=(y_1,\dots,y_n)\in Y^n2,

y=(y1,,yn)Yny=(y_1,\dots,y_n)\in Y^n3

The smallest cover size is denoted y=(y1,,yn)Yny=(y_1,\dots,y_n)\in Y^n4, and

y=(y1,,yn)Yny=(y_1,\dots,y_n)\in Y^n5

At each time-step, this controls the Hellinger distance between the conditional laws y=(y1,,yn)Yny=(y_1,\dots,y_n)\in Y^n6 and y=(y1,,yn)Yny=(y_1,\dots,y_n)\in Y^n7 uniformly in the history y=(y1,,yn)Yny=(y_1,\dots,y_n)\in Y^n8 (Jia et al., 22 Mar 2025).

Jia, Polyanskiy, and Rakhlin prove that this entropy yields a general upper bound for the non-contextual minimax regret. For any class y=(y1,,yn)Yny=(y_1,\dots,y_n)\in Y^n9 and alphabet size pΔ(Yn)p\in\Delta(Y^n)0, there are absolute constants pΔ(Yn)p\in\Delta(Y^n)1 such that for all pΔ(Yn)p\in\Delta(Y^n)2,

pΔ(Yn)p\in\Delta(Y^n)3

Up to pΔ(Yn)p\in\Delta(Y^n)4 factors, this may be written as

pΔ(Yn)p\in\Delta(Y^n)5

(Jia et al., 22 Mar 2025).

The proof route is itself structurally important. It passes through a dual form of regret, a truncation step that clips conditionals into pΔ(Yn)p\in\Delta(Y^n)6 at cost pΔ(Yn)p\in\Delta(Y^n)7, a log-to-Hellinger transform using a function pΔ(Yn)p\in\Delta(Y^n)8 with pΔ(Yn)p\in\Delta(Y^n)9, sequential symmetrization with Rademacher signs and a ghost-sample construction, and finally chaining over a sequence of covers. The resulting Dudley-type integral shows that the exact Shtarkov/NML redundancy admits a geometric upper bound even without side information (Jia et al., 22 Mar 2025).

5. Tightness regimes and Donsker-type behavior

The square-root entropy bounds are nearly tight in a broad class of regimes. If

logp(y)-\log p(y)0

then choosing logp(y)-\log p(y)1 and logp(y)-\log p(y)2 yields

logp(y)-\log p(y)3

For classes with logp(y)-\log p(y)4, described as the classical Donsker regime, the lower bounds match the upper bounds up to logarithmic factors (Jia et al., 22 Mar 2025).

Two examples illustrate the regime distinction. A Hilbert-ball infinite-dimensional linear class has

logp(y)-\log p(y)5

so logp(y)-\log p(y)6 and therefore logp(y)-\log p(y)7. A one-dimensional Lipschitz class has logp(y)-\log p(y)8, so again logp(y)-\log p(y)9 (Jia et al., 22 Mar 2025).

These results matter because they connect an exact minimax quantity, QΔ(Yn)Q\subset\Delta(Y^n)0, to scale-sensitive entropy. A plausible implication is that the Shtarkov solution is not merely an extremal coding construction; it also indexes the rate structure of nonparametric sequential prediction. In the cited work, the approach extends to contextual forecasting and is presented as resolving a long-standing open problem in online log-loss prediction and universal coding (Jia et al., 22 Mar 2025).

6. Existence, asymptotics, and computation

The Shtarkov solution is exact only when the normalizer exists. In the parametric presentation, the normalizer is

QΔ(Yn)Q\subset\Delta(Y^n)1

Finite-sample existence is guaranteed under mild regularity conditions. In the compact-support case, if QΔ(Yn)Q\subset\Delta(Y^n)2 is bounded and convex, the joint density in QΔ(Yn)Q\subset\Delta(Y^n)3 is continuous and strictly convex for large QΔ(Yn)Q\subset\Delta(Y^n)4, the outcomes have common bounded support, and suitable continuity and prior conditions hold, then QΔ(Yn)Q\subset\Delta(Y^n)5, so the frequentist NML and Bayesian NML exist (Clarke et al., 28 Jul 2025).

A separate asymptotic regularity theorem assumes uniform boundedness and convergence of Fisher information matrices, uniform local asymptotic normality around the maximizer, and continuity and equicontinuity conditions for second derivatives; under these assumptions, again QΔ(Yn)Q\subset\Delta(Y^n)6 for each QΔ(Yn)Q\subset\Delta(Y^n)7 (Clarke et al., 28 Jul 2025). Under conditions ensuring uniform convergence of the MLE and Fisher-information regularity,

QΔ(Yn)Q\subset\Delta(Y^n)8

where QΔ(Yn)Q\subset\Delta(Y^n)9, so p(y)=supqQq(y)Sn(Q),Sn(Q)yYnsupqQq(y),p^*(y)=\frac{\sup_{q\in Q} q(y)}{S_n(Q)}, \qquad S_n(Q)\coloneqq \sum_{y\in Y^n}\sup_{q\in Q} q(y),00. Clarke and Chanda cite Barron–Rissanen–Yu and S. Watanabe for asymptotic p(y)=supqQq(y)Sn(Q),Sn(Q)yYnsupqQq(y),p^*(y)=\frac{\sup_{q\in Q} q(y)}{S_n(Q)}, \qquad S_n(Q)\coloneqq \sum_{y\in Y^n}\sup_{q\in Q} q(y),01 expansions (Clarke et al., 28 Jul 2025).

Computationally, the normalizer can be available in closed or semi-closed form in some families and require approximation in others. The paper reports that, for the one-parameter exponential family, p(y)=supqQq(y)Sn(Q),Sn(Q)yYnsupqQq(y),p^*(y)=\frac{\sup_{q\in Q} q(y)}{S_n(Q)}, \qquad S_n(Q)\coloneqq \sum_{y\in Y^n}\sup_{q\in Q} q(y),02 diverges and NML fails. For larger models, the listed approaches are asymptotic approximation via p(y)=supqQq(y)Sn(Q),Sn(Q)yYnsupqQq(y),p^*(y)=\frac{\sup_{q\in Q} q(y)}{S_n(Q)}, \qquad S_n(Q)\coloneqq \sum_{y\in Y^n}\sup_{q\in Q} q(y),03, Monte Carlo with simulated maximization and importance-sampling or MCMC for the integral, and Bayesian mixture fallback

p(y)=supqQq(y)Sn(Q),Sn(Q)yYnsupqQq(y),p^*(y)=\frac{\sup_{q\in Q} q(y)}{S_n(Q)}, \qquad S_n(Q)\coloneqq \sum_{y\in Y^n}\sup_{q\in Q} q(y),04

(Clarke et al., 28 Jul 2025).

In streaming applications, only the ratio p(y)=supqQq(y)Sn(Q),Sn(Q)yYnsupqQq(y),p^*(y)=\frac{\sup_{q\in Q} q(y)}{S_n(Q)}, \qquad S_n(Q)\coloneqq \sum_{y\in Y^n}\sup_{q\in Q} q(y),05 is needed for one-step-ahead prediction, and this ratio can be approximated via Laplace’s method in smooth problems (Clarke et al., 28 Jul 2025). The same source works through classical examples including normal mean, normal variance, binomial, and Gamma/Inv-Gamma models, showing that existence and tractability depend strongly on model structure rather than on the minimax principle alone.

7. Relation to MDL, PAC-Bayesian theory, and broader predictive methodology

The Shtarkov solution also appears as a special case of a broader learning-theoretic complexity developed in the PAC-Bayesian-Rademacher-Shtarkov-MDL framework. For entropified densities

p(y)=supqQq(y)Sn(Q),Sn(Q)yYnsupqQq(y),p^*(y)=\frac{\sup_{q\in Q} q(y)}{S_n(Q)}, \qquad S_n(Q)\coloneqq \sum_{y\in Y^n}\sup_{q\in Q} q(y),06

the Shtarkov integral for a deterministic estimator p(y)=supqQq(y)Sn(Q),Sn(Q)yYnsupqQq(y),p^*(y)=\frac{\sup_{q\in Q} q(y)}{S_n(Q)}, \qquad S_n(Q)\coloneqq \sum_{y\in Y^n}\sup_{q\in Q} q(y),07 is

p(y)=supqQq(y)Sn(Q),Sn(Q)yYnsupqQq(y),p^*(y)=\frac{\sup_{q\in Q} q(y)}{S_n(Q)}, \qquad S_n(Q)\coloneqq \sum_{y\in Y^n}\sup_{q\in Q} q(y),08

and the corresponding complexity is its logarithm scaled by p(y)=supqQq(y)Sn(Q),Sn(Q)yYnsupqQq(y),p^*(y)=\frac{\sup_{q\in Q} q(y)}{S_n(Q)}, \qquad S_n(Q)\coloneqq \sum_{y\in Y^n}\sup_{q\in Q} q(y),09. In the log-loss special case with p(y)=supqQq(y)Sn(Q),Sn(Q)yYnsupqQq(y),p^*(y)=\frac{\sup_{q\in Q} q(y)}{S_n(Q)}, \qquad S_n(Q)\coloneqq \sum_{y\in Y^n}\sup_{q\in Q} q(y),10 and a model containing p(y)=supqQq(y)Sn(Q),Sn(Q)yYnsupqQq(y),p^*(y)=\frac{\sup_{q\in Q} q(y)}{S_n(Q)}, \qquad S_n(Q)\coloneqq \sum_{y\in Y^n}\sup_{q\in Q} q(y),11, one has p(y)=supqQq(y)Sn(Q),Sn(Q)yYnsupqQq(y),p^*(y)=\frac{\sup_{q\in Q} q(y)}{S_n(Q)}, \qquad S_n(Q)\coloneqq \sum_{y\in Y^n}\sup_{q\in Q} q(y),12 and recovers the classical NML density (Grünwald et al., 2017).

Within this framework, the old Shtarkov/NML complexity is exactly the special case with luckiness weight p(y)=supqQq(y)Sn(Q),Sn(Q)yYnsupqQq(y),p^*(y)=\frac{\sup_{q\in Q} q(y)}{S_n(Q)}, \qquad S_n(Q)\coloneqq \sum_{y\in Y^n}\sup_{q\in Q} q(y),13 and deterministic posterior. The cited results then connect minimax log-loss regret to excess-risk bounds: Theorem 3.1 gives an exact exponential-moment identity for the generalized complexity, and Corollary 3.2 specializes to deterministic ERM to bound expected excess risk in terms of the NML complexity under bounded-loss plus a central/Bernstein condition (Grünwald et al., 2017).

The same paper places the Shtarkov code inside classical MDL. Two-part codelengths of the form

p(y)=supqQq(y)Sn(Q),Sn(Q)yYnsupqQq(y),p^*(y)=\frac{\sup_{q\in Q} q(y)}{S_n(Q)}, \qquad S_n(Q)\coloneqq \sum_{y\in Y^n}\sup_{q\in Q} q(y),14

use the NML normalization p(y)=supqQq(y)Sn(Q),Sn(Q)yYnsupqQq(y),p^*(y)=\frac{\sup_{q\in Q} q(y)}{S_n(Q)}, \qquad S_n(Q)\coloneqq \sum_{y\in Y^n}\sup_{q\in Q} q(y),15 for submodels p(y)=supqQq(y)Sn(Q),Sn(Q)yYnsupqQq(y),p^*(y)=\frac{\sup_{q\in Q} q(y)}{S_n(Q)}, \qquad S_n(Q)\coloneqq \sum_{y\in Y^n}\sup_{q\in Q} q(y),16, and the resulting p(y)=supqQq(y)Sn(Q),Sn(Q)yYnsupqQq(y),p^*(y)=\frac{\sup_{q\in Q} q(y)}{S_n(Q)}, \qquad S_n(Q)\coloneqq \sum_{y\in Y^n}\sup_{q\in Q} q(y),17-generalized MDL estimator attains the same excess-risk rates as ERM on the best submodel (Grünwald et al., 2017). This identifies the Shtarkov solution not only as a predictor but also as a complexity term that mediates between MDL, PAC-Bayesian information complexity, and empirical-process quantities such as Rademacher complexity.

Two limitations follow directly from the cited literature. First, NML is not automatically available: existence can fail, as in the one-parameter exponential example, or can demand nontrivial approximation. Second, minimax optimality is not a guarantee of empirical dominance under misspecification or computational constraints. In a review of streaming observational prediction, Clarke and Chanda report that although Shtarkov is conceptually minimax-optimal, it can be outperformed in practice when models are misspecified or computing is constrained (Clarke et al., 28 Jul 2025). In that sense, the Shtarkov solution is best understood as an exact benchmark for worst-case log-loss regret, with extensions that now cover contextual forecasting and geometric regret analysis, rather than as a universal prescription independent of model class or computational regime.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Shtarkov Solution.