Papers
Topics
Authors
Recent
Search
2000 character limit reached

Martingale Posteriors: Predictive Uncertainty

Updated 14 July 2026
  • Martingale posteriors are posterior-like constructs derived from sequential predictive distributions, providing a predictive-first approach to uncertainty quantification.
  • They enforce coherence via a martingale structure, ensuring that sequential updates remain consistent and recover standard Bayesian results under conditionally i.i.d. settings.
  • Various algorithmic variants, including copula updates, score-based recursions, and federated methods, extend the framework to density estimation, neural processes, and high-dimensional models.

Martingale posteriors are posterior-like distributions constructed from a sequence of predictive distributions for unobserved data rather than from an explicit prior–likelihood pair. In the formulation introduced by Fong, Holmes, and Walker, uncertainty is attributed to the missing continuation Yn+1:Y_{n+1:\infty} of the observed sample Y1:nY_{1:n}, with the target quantity becoming determined once the relevant full data are available. A martingale posterior is then the law induced on the target by a coherent predictive mechanism for those missing observations. In the conditionally i.i.d. Bayesian setting, Doob’s theorem implies that choosing the Bayesian posterior predictive returns the conventional posterior, so ordinary Bayes appears as a special case of the predictive construction (Fong et al., 2021).

1. Historical origin and basic construction

The foundational construction begins with observed data Y1:n=(Y1,,Yn)Y_{1:n}=(Y_1,\dots,Y_n), a target θ\theta that can be represented as a functional of the completed population or of a limiting empirical distribution, and a predictive law for the unobserved continuation Yn+1:Y_{n+1:\infty}. In this setup, the completed-data target is written as

θ=θ(Y1:)=θ(F),\theta_\infty=\theta(Y_{1:\infty})=\theta(F_\infty),

where FF_\infty is a limiting empirical or predictive distribution. The martingale posterior is the induced law

$\Pi_{\infty}(\theta_\infty \in A \mid y_{1:n}) = \int \mathbbm{1}\{\theta(F_{\infty}) \in A\}\, d\Pi(F_{\infty}\mid y_{1:n}),$

and a finite version is

$\Pi_N(\theta_N \in A \mid y_{1:n}) = \int \mathbbm{1}\{\theta(y_{1:N}) \in A\}\, p(y_{n+1:N}\mid y_{1:n})\,dy_{n+1:N}.$

The predictive joint law factors as

p(yn+1:Ny1:n)=i=n+1Np(yiy1:i1).p(y_{n+1:N}\mid y_{1:n})=\prod_{i=n+1}^N p(y_i\mid y_{1:i-1}).

These definitions formalize the predictive-first view in which posterior uncertainty is generated by imputing the missing future rather than by directly updating a prior on Y1:nY_{1:n}0 (Fong et al., 2021).

This construction is motivated by the claim that, if the complete population were observed, the inferential target would be known exactly. The predictive sequence is therefore treated as primitive. In practice, posterior sampling is implemented by predictive resampling: compute the current predictive, simulate future observations recursively from it, update the predictive after each simulated observation, and compute the target from the synthetic completed sample. In the original framework, this yields posterior samples without Markov chain Monte Carlo and permits direct uncertainty quantification for means, medians, density functionals, regression functionals, and likelihood-based targets defined through losses such as Y1:nY_{1:n}1 (Fong et al., 2021).

2. Martingale structure, coherence, and relation to Bayes

The defining structural requirement is a martingale or sequential-coherence condition on the predictive or posterior sequence. In the original nonparametric formulation, one writes

Y1:nY_{1:n}2

with filtration Y1:nY_{1:n}3, and imposes

Y1:nY_{1:n}4

Under the conditionally identically distributed setting used in the original theory, Y1:nY_{1:n}5 almost surely, and the limiting empirical distribution satisfies

Y1:nY_{1:n}6

This establishes existence of the martingale posterior measure and justifies computing Y1:nY_{1:n}7 from either the limiting predictive or the limiting empirical distribution (Fong et al., 2021).

In the ordinary Bayesian conditionally i.i.d. model,

Y1:nY_{1:n}8

the posterior mean Y1:nY_{1:n}9 is itself a martingale, with

Y1:n=(Y1,,Yn)Y_{1:n}=(Y_1,\dots,Y_n)0

Doob’s theorem gives Y1:n=(Y1,,Yn)Y_{1:n}=(Y_1,\dots,Y_n)1 almost surely under standard conditions, so if future observations are drawn from the Bayesian posterior predictive, the induced martingale posterior coincides with the usual posterior. Later asymptotic work on parametric martingale posteriors makes this relation more explicit through a predictive central limit theorem and a Bernstein–von Mises theorem. In the parametric plug-in construction,

Y1:n=(Y1,,Yn)Y_{1:n}=(Y_1,\dots,Y_n)2

and, under regularity conditions, one has

Y1:n=(Y1,,Yn)Y_{1:n}=(Y_1,\dots,Y_n)3

with Y1:n=(Y1,,Yn)Y_{1:n}=(Y_1,\dots,Y_n)4 (Fong et al., 2024).

A distinct but related line studies posterior probabilities themselves as stochastic processes in sample size. There the posterior mass Y1:n=(Y1,,Yn)Y_{1:n}=(Y_1,\dots,Y_n)5 is a martingale under the Bayesian marginal law, a submartingale under the true sampling law Y1:n=(Y1,,Yn)Y_{1:n}=(Y_1,\dots,Y_n)6, and generally neither monotone nor conditionally decreasing under a false sampling law Y1:n=(Y1,,Yn)Y_{1:n}=(Y_1,\dots,Y_n)7. This sharpens the distinction between martingale structure under a mixture law and posterior behavior under fixed-data-generating laws (Hart et al., 2022).

3. Foundational debates and predictive completeness

A major discussion in the literature concerns whether martingale posteriors are genuinely “prior-free.” Rossell argues that they are not. Given a likelihood and a posterior, “there is a posterior and a likelihood, hence the prior is proportional to their ratio,” namely

Y1:n=(Y1,,Yn)Y_{1:n}=(Y_1,\dots,Y_n)8

On that reading, the framework is best understood as using a data-dependent prior. Rossell regards this as “an interesting avenue to develop objective Bayes methods,” but also stresses the cost: “loosing the coherence property in belief updating.” He recommends inspecting the implied prior as a diagnostic, and gives two illustrations from Figure 1: in a Bernoulli(Y1:n=(Y1,,Yn)Y_{1:n}=(Y_1,\dots,Y_n)9) example with θ\theta0, the implied prior “places little mass around that value,” while in a Normal(5,1) example with θ\theta1, “the prior is centered around the sample mean.” He further writes that such behavior “might be problematic for model choice via Bayes factors, e.g. returning a very small integrated likelihood in the Bernoulli example” (Rossell, 2023).

Rossell’s discussion also challenges stronger practical claims about elicitation, computation, and asymptotics. He states that “sometimes it is easier to elicit a predictive than a prior,” but that “in my experience the reverse is often true,” gives regression and θ\theta2 as an example, says “I am afraid I disagree on the frameworks’ computational convenience,” and argues that “assuming that at θ\theta3 there is no uncertainty left” can fail in high-dimensional regression with θ\theta4. He concludes that “The proposed framework does not account for such uncertainty, unless suitable adjustments are made” (Rossell, 2023).

A complementary discussion by Draper and Guo reframes bootstrap comparisons. They argue that the frequentist bootstrap is “actually an instance of Bayesian nonparametric inference,” using exchangeability, de Finetti’s theorem, a Dirichlet process prior θ\theta5, the low-information limit θ\theta6, and the posterior

θ\theta7

which reduces to θ\theta8 as θ\theta9. Their theorem states that frequentist bootstrap samples of size Yn+1:Y_{n+1:\infty}0 are “asymptotically stochastically indistinguishable from stick-breaking samples of the same size from Yn+1:Y_{n+1:\infty}1” (Draper et al., 2023).

A more recent theoretical critique focuses on exchangeable Bernoulli sequences and asks what predictive structure is identified by a martingale of posterior means. The answer is that one-step prediction is pinned down by the mean, but multi-step prediction is not. For

Yn+1:Y_{n+1:\infty}2

the binomial expansion shows dependence on posterior moments up to order Yn+1:Y_{n+1:\infty}3. The paper proves that for every Yn+1:Y_{n+1:\infty}4 the map from posterior mean to Yn+1:Y_{n+1:\infty}5-step predictive is set-valued rather than single-valued, that the plug-in predictive is strictly dominated by the Bayes predictive under any strictly proper scoring rule whenever the posterior is non-degenerate, and that predictive completeness holds if and only if the conditional law of the terminal value is uniquely determined. Hill’s Yn+1:Y_{n+1:\infty}6 rule under the Jeffreys prior Yn+1:Y_{n+1:\infty}7 is given as a positive example because it specifies the full conditional law rather than only the first moment (Polson et al., 28 Feb 2026).

4. Principal methodological variants

The original nonparametric predictive construction uses recursive copula updates. In the univariate continuous case,

Yn+1:Y_{n+1:\infty}8

equivalently

Yn+1:Y_{n+1:\infty}9

with

θ=θ(Y1:)=θ(F),\theta_\infty=\theta(Y_{1:\infty})=\theta(F_\infty),0

This yields c.i.d. coherence and predictive-resampling algorithms for density estimation, regression, and classification (Fong et al., 2021).

A parametric branch replaces the predictive CDF by a parameter process. In “parametric martingale posteriors,” the recursively updated estimate is a martingale under the predictive-resampling law, and a hybrid algorithm replaces the long unsimulated tail by a Gaussian approximation: θ=θ(Y1:)=θ(F),\theta_\infty=\theta(Y_{1:\infty})=\theta(F_\infty),1 with

θ=θ(Y1:)=θ(F),\theta_\infty=\theta(Y_{1:\infty})=\theta(F_\infty),2

The predictive central limit theorem justifies this acceleration, while the Bernstein–von Mises theorem supplies the large-sample normal approximation centered at the initial estimator θ=θ(Y1:)=θ(F),\theta_\infty=\theta(Y_{1:\infty})=\theta(F_\infty),3 (Fong et al., 2024).

A closely related variant is the score-based martingale posterior. There the recursion is

θ=θ(Y1:)=θ(F),\theta_\infty=\theta(Y_{1:\infty})=\theta(F_\infty),4

with θ=θ(Y1:)=θ(F),\theta_\infty=\theta(Y_{1:\infty})=\theta(F_\infty),5. Because θ=θ(Y1:)=θ(F),\theta_\infty=\theta(Y_{1:\infty})=\theta(F_\infty),6, the parameter sequence is a martingale, and under regularity conditions it converges almost surely to a finite random limit θ=θ(Y1:)=θ(F),\theta_\infty=\theta(Y_{1:\infty})=\theta(F_\infty),7. The resulting martingale posterior is prior-free in the paper’s terminology and does not rely on MCMC (Cui et al., 3 Jan 2025).

Semiparametric predictive Bayes has produced a “moment martingale posterior.” Its predictive mixture is

θ=θ(Y1:)=θ(F),\theta_\infty=\theta(Y_{1:\infty})=\theta(F_\infty),8

and the method of moments is used to choose θ=θ(Y1:)=θ(F),\theta_\infty=\theta(Y_{1:\infty})=\theta(F_\infty),9 so that

FF_\infty0

Under this constraint, the first FF_\infty1 empirical moments become martingales: FF_\infty2 This produces a semiparametric martingale posterior with regularization when the sample size is small and robustness to misspecification when the sample size is large (Yung et al., 24 Jul 2025).

For quantile inference, the quantile martingale posterior updates a possibly non-monotone quantile estimate FF_\infty3 by

FF_\infty4

while the corresponding proper quantile function is recovered by increasing rearrangement. This yields posterior samples for quantile functions and quantile regression without an explicit likelihood-prior specification and permits a Gaussian-process approximation for fast sampling (Fong et al., 2024).

Other specialized variants include martingale posterior distributions for log-concave density functions, where the starting point is the log-concave NPMLE and uncertainty is generated by repeatedly sampling from the current fitted density and refitting the NPMLE, with convergence proved through submartingale arguments (Cui et al., 2024).

5. Applications and computational profile

The martingale-posterior framework has been adapted to several modern machine-learning settings. For prior-data fitted networks, the problem is that PFNs approximate a posterior predictive distribution but do not supply a posterior distribution over predictive means, quantiles, or similar summaries. The proposed solution uses the PFN output only as the initial predictive FF_\infty5, then enforces martingale coherence through the Gaussian-copula update

FF_\infty6

producing posterior samples of functionals FF_\infty7. The paper proves almost sure convergence to a random limit FF_\infty8, gives the bound

FF_\infty9

and reports complexity

$\Pi_{\infty}(\theta_\infty \in A \mid y_{1:n}) = \int \mathbbm{1}\{\theta(F_{\infty}) \in A\}\, d\Pi(F_{\infty}\mid y_{1:n}),$0

excluding the initial PFN call (Nagler et al., 16 May 2025).

In neural processes, martingale posterior uncertainty replaces explicit latent-variable posterior assumptions. “Martingale Posterior Neural Processes” construct an exchangeable predictive distribution $\Pi_{\infty}(\theta_\infty \in A \mid y_{1:n}) = \int \mathbbm{1}\{\theta(F_{\infty}) \in A\}\, d\Pi(F_{\infty}\mid y_{1:n}),$1, define the finite martingale posterior

$\Pi_{\infty}(\theta_\infty \in A \mid y_{1:n}) = \int \mathbbm{1}\{\theta(F_{\infty}) \in A\}\, d\Pi(F_{\infty}\mid y_{1:n}),$2

and amortize $\Pi_{\infty}(\theta_\infty \in A \mid y_{1:n}) = \int \mathbbm{1}\{\theta(F_{\infty}) \in A\}\, d\Pi(F_{\infty}\mid y_{1:n}),$3 with the encoder. The resulting MPNP and MPANP use pseudo-context generation in representation space rather than a fixed Gaussian latent posterior (Lee et al., 2023).

Federated learning has produced a one-shot “Federated Martingale Posterior” protocol. Centralized martingale posterior sampling would require pooled data, so each client instead uploads a compressed set

$\Pi_{\infty}(\theta_\infty \in A \mid y_{1:n}) = \int \mathbbm{1}\{\theta(F_{\infty}) \in A\}\, d\Pi(F_{\infty}\mid y_{1:n}),$4

the server forms

$\Pi_{\infty}(\theta_\infty \in A \mid y_{1:n}) = \int \mathbbm{1}\{\theta(F_{\infty}) \in A\}\, d\Pi(F_{\infty}\mid y_{1:n}),$5

samples pseudo-data with a set transformer, and solves

$\Pi_{\infty}(\theta_\infty \in A \mid y_{1:n}) = \int \mathbbm{1}\{\theta(F_{\infty}) \in A\}\, d\Pi(F_{\infty}\mid y_{1:n}),$6

Experiments on MNIST, CIFAR-10, and CIFAR-100 show that FMP “closely matches the centralized counterpart and significantly improves calibration over consensus-style baselines” (Zhang et al., 18 May 2026).

The score-based approach has also been explored directly for deep neural networks. There the recursion is

$\Pi_{\infty}(\theta_\infty \in A \mid y_{1:n}) = \int \mathbbm{1}\{\theta(F_{\infty}) \in A\}\, d\Pi(F_{\infty}\mid y_{1:n}),$7

with $\Pi_{\infty}(\theta_\infty \in A \mid y_{1:n}) = \int \mathbbm{1}\{\theta(F_{\infty}) \in A\}\, d\Pi(F_{\infty}\mid y_{1:n}),$8 simulated under the current parameter and $\Pi_{\infty}(\theta_\infty \in A \mid y_{1:n}) = \int \mathbbm{1}\{\theta(F_{\infty}) \in A\}\, d\Pi(F_{\infty}\mid y_{1:n}),$9 a Fisher-based preconditioner. The resulting score-based martingale posterior can be competitive with NUTS in a small neural-network example, but large-scale behavior is highly sensitive to preconditioning; on MNIST, unpreconditioned SMP is stable but nearly deterministic, whereas diagonal Fisher variants can be over-dispersed and miscalibrated (Zhumekenov et al., 14 Jun 2026).

For discretely observed diffusions, the difficulty is that transition densities are unavailable and naive discretization of the score is unstable. The proposed MPD algorithm uses guided diffusion bridges and a difference-of-two-log-$\Pi_N(\theta_N \in A \mid y_{1:n}) = \int \mathbbm{1}\{\theta(y_{1:N}) \in A\}\, p(y_{n+1:N}\mid y_{1:n})\,dy_{n+1:N}.$0 score increment,

$\Pi_N(\theta_N \in A \mid y_{1:n}) = \int \mathbbm{1}\{\theta(y_{1:N}) \in A\}\, p(y_{n+1:N}\mid y_{1:n})\,dy_{n+1:N}.$1

to build a practical martingale-posterior recursion. Its main theorem proves that, under assumptions (A1)–(A2),

$\Pi_N(\theta_N \in A \mid y_{1:n}) = \int \mathbbm{1}\{\theta(y_{1:N}) \in A\}\, p(y_{n+1:N}\mid y_{1:n})\,dy_{n+1:N}.$2

so the discretized algorithm approximates the ideal continuous-time martingale posterior with $\Pi_N(\theta_N \in A \mid y_{1:n}) = \int \mathbbm{1}\{\theta(y_{1:N}) \in A\}\, p(y_{n+1:N}\mid y_{1:n})\,dy_{n+1:N}.$3 mean-square error. The paper reports “orders of magnitude speed up versus state-of-the-art MCMC algorithms” (Yao et al., 30 Apr 2026).

6. Limitations, asymptotics, and terminology

Martingale posteriors remain a heterogeneous family of posterior-like constructions rather than a single closed formalism. Several limitations recur across the literature. Rossell’s critique emphasizes loss of dynamic coherence, sensitivity of the implied prior, possible distortion of Bayes factors, unresolved computational advantage, and the fact that high-dimensional settings may retain posterior uncertainty even as $\Pi_N(\theta_N \in A \mid y_{1:n}) = \int \mathbbm{1}\{\theta(y_{1:N}) \in A\}\, p(y_{n+1:N}\mid y_{1:n})\,dy_{n+1:N}.$4 unless the framework is suitably adjusted (Rossell, 2023). The Bernoulli predictive-completeness results likewise show that first-moment martingale coherence identifies one-step prediction but not the full hierarchy of multi-step predictive distributions, posterior variances, or other nonlinear functionals (Polson et al., 28 Feb 2026).

At the same time, recent asymptotic work has clarified where martingale posteriors recover familiar large-sample behavior. Parametric martingale posteriors now have predictive central limit theory and a Bernstein–von Mises theorem (Fong et al., 2024). Quantile martingale posteriors admit posterior consistency and contraction results, together with a Gaussian-process asymptotic approximation (Fong et al., 2024). This suggests that, in regular settings, predictive-first uncertainty quantification can be studied with tools parallel to those used for standard Bayes, even though the inferential object is constructed differently.

A final terminological distinction is necessary. “M-posteriors” are not martingale posteriors. In that literature, the “M” stands for M-estimation: the posterior is

$\Pi_N(\theta_N \in A \mid y_{1:n}) = \int \mathbbm{1}\{\theta(y_{1:N}) \in A\}\, p(y_{n+1:N}\mid y_{1:n})\,dy_{n+1:N}.$5

and there is “no martingale structure, optional stopping argument, filtration-based update rule, or posterior process indexed by time” (Marusic et al., 1 Oct 2025). The overlap is therefore only at the level of generalized or posterior-like inference.

Taken together, the literature presents martingale posteriors as predictive-sequential uncertainty distributions whose defining coherence is martingale structure rather than prior-to-posterior updating. Standard Bayes is recovered when the predictive sequence is the Bayesian posterior predictive; outside that case, martingale posteriors form a broader class of predictive Bayes procedures, ranging from copula-based nonparametrics and score-driven recursions to semiparametric, quantile, neural, federated, and diffusion-specific constructions (Fong et al., 2021).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Martingale Posteriors.