---
title: 'Martingale Posteriors: Predictive Uncertainty'
url: https://www.emergentmind.com/topics/martingale-posteriors
type: topic
---

# Martingale Posteriors: Predictive Uncertainty

Martingale posteriors are posterior-like distributions constructed from a sequence of predictive distributions for unobserved data rather than from an explicit prior–likelihood pair. In the formulation introduced by Fong, Holmes, and Walker, uncertainty is attributed to the missing continuation \(Y_{n+1:\infty}\) of the observed sample \(Y_{1:n}\), with the target quantity becoming determined once the relevant full data are available. A martingale posterior is then the law induced on the target by a coherent predictive mechanism for those missing observations. In the conditionally i.i.d. Bayesian setting, Doob’s theorem implies that choosing the Bayesian posterior predictive returns the conventional posterior, so ordinary Bayes appears as a special case of the predictive construction [2103.15671].

## 1. Historical origin and basic construction

The foundational construction begins with observed data \(Y_{1:n}=(Y_1,\dots,Y_n)\), a target \(\theta\) that can be represented as a functional of the completed population or of a limiting empirical distribution, and a predictive law for the unobserved continuation \(Y_{n+1:\infty}\). In this setup, the completed-data target is written as
\[
\theta_\infty=\theta(Y_{1:\infty})=\theta(F_\infty),
\]
where \(F_\infty\) is a limiting empirical or predictive distribution. The martingale posterior is the induced law
\[
\Pi_{\infty}(\theta_\infty \in A \mid y_{1:n}) = \int \mathbbm{1}\{\theta(F_{\infty}) \in A\}\, d\Pi(F_{\infty}\mid y_{1:n}),
\]
and a finite version is
\[
\Pi_N(\theta_N \in A \mid y_{1:n}) = \int \mathbbm{1}\{\theta(y_{1:N}) \in A\}\, p(y_{n+1:N}\mid y_{1:n})\,dy_{n+1:N}.
\]
The predictive joint law factors as
\[
p(y_{n+1:N}\mid y_{1:n})=\prod_{i=n+1}^N p(y_i\mid y_{1:i-1}).
\]
These definitions formalize the predictive-first view in which posterior uncertainty is generated by imputing the missing future rather than by directly updating a prior on \(\theta\) [2103.15671].

This construction is motivated by the claim that, if the complete population were observed, the inferential target would be known exactly. The predictive sequence is therefore treated as primitive. In practice, posterior sampling is implemented by predictive resampling: compute the current predictive, simulate future observations recursively from it, update the predictive after each simulated observation, and compute the target from the synthetic completed sample. In the original framework, this yields posterior samples without Markov chain Monte Carlo and permits direct uncertainty quantification for means, medians, density functionals, regression functionals, and likelihood-based targets defined through losses such as \(\ell(\theta,y)=-\log f_\theta(y)\) [2103.15671].

## 2. Martingale structure, coherence, and relation to Bayes

The defining structural requirement is a martingale or sequential-coherence condition on the predictive or posterior sequence. In the original nonparametric formulation, one writes
\[
P_i(y)=P(Y_{i+1}\le y\mid y_{1:i}),
\]
with filtration \(\mathcal F_i=\sigma(Y_1,\dots,Y_i)\), and imposes
\[
E[P_i(y)\mid y_{1:i-1}] = P_{i-1}(y).
\]
Under the conditionally identically distributed setting used in the original theory, \(P_i(y)\to P_\infty(y)\) almost surely, and the limiting empirical distribution satisfies
\[
F_\infty(y)=P_\infty(y)\quad \text{a.s.}
\]
This establishes existence of the martingale posterior measure and justifies computing \(\theta\) from either the limiting predictive or the limiting empirical distribution [2103.15671].

In the ordinary Bayesian conditionally i.i.d. model,
\[
p(\theta,y_{1:N})=\pi(\theta)\prod_{i=1}^N f_\theta(y_i),
\]
the posterior mean \(\bar\theta_N=E[\Theta\mid Y_{1:N}]\) is itself a martingale, with
\[
E[\bar\theta_N\mid Y_{1:N-1}] = \bar\theta_{N-1}.
\]
Doob’s theorem gives \(\bar\theta_N\to \Theta\) almost surely under standard conditions, so if future observations are drawn from the Bayesian posterior predictive, the induced martingale posterior coincides with the usual posterior. Later asymptotic work on parametric martingale posteriors makes this relation more explicit through a predictive central limit theorem and a Bernstein–von Mises theorem. In the parametric plug-in construction,
\[
\theta_N=\theta_{N-1}+N^{-1}\,\mathcal I(\theta_{N-1})^{-1}\,s(\theta_{N-1},Y_N),
\]
and, under regularity conditions, one has
\[
r_n^{-1}(\theta_{n\infty}-\theta_n)\to \mathcal N\{0,\mathcal I(\theta^*)^{-1}\},
\]
with \(r_n^2=\sum_{i=n+1}^\infty i^{-2}\asymp n^{-1}\) [2410.17692].

A distinct but related line studies posterior probabilities themselves as stochastic processes in sample size. There the posterior mass \(q_n^{\theta_0}=\Pr(\theta=\theta_0\mid \mathcal F_n)\) is a martingale under the Bayesian marginal law, a submartingale under the true sampling law \(\mathbb P_{\theta_0}\), and generally neither monotone nor conditionally decreasing under a false sampling law \(\mathbb P_{\theta_1}\). This sharpens the distinction between martingale structure under a mixture law and posterior behavior under fixed-data-generating laws [2209.11728].

## 3. Foundational debates and predictive completeness

A major discussion in the literature concerns whether martingale posteriors are genuinely “prior-free.” Rossell argues that they are not. Given a likelihood and a posterior, “there is a posterior and a likelihood, hence the prior is proportional to their ratio,” namely
\[
\pi(\theta\mid y_{1:n}) \propto \frac{\pi_n(\theta\mid y_{1:n})}{L_n(\theta;y_{1:n})}.
\]
On that reading, the framework is best understood as using a data-dependent prior. Rossell regards this as “an interesting avenue to develop objective Bayes methods,” but also stresses the cost: “loosing the coherence property in belief updating.” He recommends inspecting the implied prior as a diagnostic, and gives two illustrations from Figure 1: in a Bernoulli(\(0.5\)) example with \(n=100\), the implied prior “places little mass around that value,” while in a Normal(5,1) example with \(n=100\), “the prior is centered around the sample mean.” He further writes that such behavior “might be problematic for model choice via Bayes factors, e.g. returning a very small integrated likelihood in the Bernoulli example” [2303.02403].

Rossell’s discussion also challenges stronger practical claims about elicitation, computation, and asymptotics. He states that “sometimes it is easier to elicit a predictive than a prior,” but that “in my experience the reverse is often true,” gives regression and \(R^2\) as an example, says “I am afraid I disagree on the frameworks’ computational convenience,” and argues that “assuming that at \(n=\infty\) there is no uncertainty left” can fail in high-dimensional regression with \(p\gg n\). He concludes that “The proposed framework does not account for such uncertainty, unless suitable adjustments are made” [2303.02403].

A complementary discussion by Draper and Guo reframes bootstrap comparisons. They argue that the frequentist bootstrap is “actually an instance of Bayesian nonparametric inference,” using exchangeability, de Finetti’s theorem, a Dirichlet process prior \(DP(a,F_0)\), the low-information limit \(DP(0)\), and the posterior
\[
DP\!\left(a+n,\, \frac{aF_0+n\hat F_n}{a+n} \right),
\]
which reduces to \(DP(n,\hat F_n)\) as \(a\to0\). Their theorem states that frequentist bootstrap samples of size \(n\) are “asymptotically stochastically indistinguishable from stick-breaking samples of the same size from \(DP(n,\hat F_n)\)” [2302.07779].

A more recent theoretical critique focuses on exchangeable Bernoulli sequences and asks what predictive structure is identified by a martingale of posterior means. The answer is that one-step prediction is pinned down by the mean, but multi-step prediction is not. For
\[
\Pr(X_{n+1}=\cdots=X_{n+k}=0\mid \mathcal F_n)=E[(1-\theta)^k\mid \mathcal F_n],
\]
the binomial expansion shows dependence on posterior moments up to order \(k\). The paper proves that for every \(k\ge2\) the map from posterior mean to \(k\)-step predictive is set-valued rather than single-valued, that the plug-in predictive is strictly dominated by the Bayes predictive under any strictly proper scoring rule whenever the posterior is non-degenerate, and that predictive completeness holds if and only if the conditional law of the terminal value is uniquely determined. Hill’s \(A_{(n)}\) rule under the Jeffreys prior \(Beta(\tfrac12,\tfrac12)\) is given as a positive example because it specifies the full conditional law rather than only the first moment [2603.00661].

## 4. Principal methodological variants

The original nonparametric predictive construction uses recursive copula updates. In the univariate continuous case,
\[
p_{i+1}(y)=\left(1-\alpha_{i+1}\right)p_i(y)+\alpha_{i+1} c_\rho\{P_i(y),P_i(y_{i+1})\}p_i(y),
\]
equivalently
\[
P_{i+1}(y)=(1-\alpha_{i+1})P_i(y)+\alpha_{i+1}H_\rho\{P_i(y),P_i(y_{i+1})\},
\]
with
\[
H_\rho(u,v)= \Phi\left(\frac{\Phi^{-1}(u)-\rho\Phi^{-1}(v)}{\sqrt{1-\rho^2}}\right),\qquad
\alpha_i=\left(2-\frac{1}{i}\right)\frac{1}{i+1}.
\]
This yields c.i.d. coherence and predictive-resampling algorithms for density estimation, regression, and classification [2103.15671].

A parametric branch replaces the predictive CDF by a parameter process. In “parametric martingale posteriors,” the recursively updated estimate is a martingale under the predictive-resampling law, and a hybrid algorithm replaces the long unsimulated tail by a Gaussian approximation:
\[
\widehat\theta_\infty=\theta_N+\widehat V_N\varepsilon,\qquad \varepsilon\sim\mathcal N(0,1),
\]
with
\[
\widehat V_N^2=\mathcal I(\theta_N)^{-1}\sum_{i=N}^\infty i^{-2}.
\]
The predictive central limit theorem justifies this acceleration, while the Bernstein–von Mises theorem supplies the large-sample normal approximation centered at the initial estimator \(\theta_n\) [2410.17692].

A closely related variant is the score-based martingale posterior. There the recursion is
\[
X_{m+1}\sim p(\cdot\mid \widehat\theta_m),\qquad
\widehat\theta_{m+1}=\widehat\theta_m+\epsilon_m s(X_{m+1},\widehat\theta_m),
\]
with \(s(x,\theta)=\nabla_\theta\log p(x\mid\theta)\). Because \(E[s(X_{m+1},\widehat\theta_m)\mid \widehat\theta_m]=0\), the parameter sequence is a martingale, and under regularity conditions it converges almost surely to a finite random limit \(\widehat\theta_\infty\). The resulting martingale posterior is prior-free in the paper’s terminology and does not rely on MCMC [2501.01890].

Semiparametric predictive Bayes has produced a “moment martingale posterior.” Its predictive mixture is
\[
p_i(y)=\frac{c}{c+i}f_{\theta_i}(y) + \frac{i}{c+i}\mathbbm{P}_i,
\]
and the method of moments is used to choose \(\theta_i\) so that
\[
\mu_{\theta_i}^{(k)}=\mu_i^{(k)},\qquad k=1,\dots,p.
\]
Under this constraint, the first \(p\) empirical moments become martingales:
\[
E\left[\mu_{i+1}^{(k)}\mid y_{1:i}\right]=\mu_i^{(k)}.
\]
This produces a semiparametric martingale posterior with regularization when the sample size is small and robustness to misspecification when the sample size is large [2507.18148].

For quantile inference, the quantile martingale posterior updates a possibly non-monotone quantile estimate \(Q_N\) by
\[
Q_{N+1}(u)=Q_N(u)+\alpha_{N+1}\left[u-H_{\rho_{N+1}(u,V_{N+1})\right],\qquad V_{N+1}\sim\mathcal U(0,1),
\]
while the corresponding proper quantile function is recovered by increasing rearrangement. This yields posterior samples for quantile functions and quantile regression without an explicit likelihood-prior specification and permits a Gaussian-process approximation for fast sampling [2406.03358].

Other specialized variants include martingale posterior distributions for log-concave density functions, where the starting point is the log-concave NPMLE and uncertainty is generated by repeatedly sampling from the current fitted density and refitting the NPMLE, with convergence proved through submartingale arguments [2401.14515].

## 5. Applications and computational profile

The martingale-posterior framework has been adapted to several modern machine-learning settings. For prior-data fitted networks, the problem is that PFNs approximate a posterior predictive distribution but do not supply a posterior distribution over predictive means, quantiles, or similar summaries. The proposed solution uses the PFN output only as the initial predictive \(P_0\), then enforces martingale coherence through the Gaussian-copula update
\[
P_k(y) = (1 - \alpha_{n + k})P_{k - 1}(y) + \alpha_{n + k} H_\rho\!\bigl(P_{k - 1}(y), P_{k - 1}(y_{n + k})\bigr),
\]
producing posterior samples of functionals \(\theta(P_{\infty,x})\). The paper proves almost sure convergence to a random limit \(P_{\infty,x}\), gives the bound
\[
\sup_y \limsup_{M \to \infty} \Pr\left(|P_{M}(y)-P_{N}(y)| \ge \epsilon\right) \le 2\exp\left(-\epsilon^2 (n + N)/8\right),
\]
and reports complexity
\[
O(BN),
\]
excluding the initial PFN call [2505.11325].

In neural processes, martingale posterior uncertainty replaces explicit latent-variable posterior assumptions. “Martingale Posterior Neural Processes” construct an exchangeable predictive distribution \(p(Z'\mid Z_c)\), define the finite martingale posterior
\[
\pi_N(\theta \in A \mid Z) = \int \mathbbm{1}(\theta(g_N) \in A) \, p(dZ' \mid Z),
\]
and amortize \(\theta(g_N)\) with the encoder. The resulting MPNP and MPANP use pseudo-context generation in representation space rather than a fixed Gaussian latent posterior [2304.09431].

Federated learning has produced a one-shot “Federated Martingale Posterior” protocol. Centralized martingale posterior sampling would require pooled data, so each client instead uploads a compressed set
\[
\tilde{\mathcal Z}_m = h_\phi(\mathcal Z_m)=\{\tilde z_{m,i}\}_{i=1}^s,
\]
the server forms
\[
\tilde{\mathcal Z} = \bigcup_{m=1}^M \tilde{\mathcal Z}_m,
\]
samples pseudo-data with a set transformer, and solves
\[
\theta^{\text{FMP}}_{\mathcal{Z}}
= \arg \min_{\theta} \sum_{z \in \tilde{\mathcal{Z}} \cup \tilde{\mathcal{Z}}'} \ell(z, \theta).
\]
Experiments on MNIST, CIFAR-10, and CIFAR-100 show that FMP “closely matches the centralized counterpart and significantly improves calibration over consensus-style baselines” [2605.18554].

The score-based approach has also been explored directly for deep neural networks. There the recursion is
\[
\theta_k = \theta_{k-1} + \gamma_k P_k^{-1}\nabla \log f_{\theta_{k-1}}(Z_k),\qquad \gamma_k=\frac{\tau}{N+k},
\]
with \(Z_k\) simulated under the current parameter and \(P_k\) a Fisher-based preconditioner. The resulting score-based martingale posterior can be competitive with NUTS in a small neural-network example, but large-scale behavior is highly sensitive to preconditioning; on MNIST, unpreconditioned SMP is stable but nearly deterministic, whereas diagonal Fisher variants can be over-dispersed and miscalibrated [2606.15725].

For discretely observed diffusions, the difficulty is that transition densities are unavailable and naive discretization of the score is unstable. The proposed MPD algorithm uses guided diffusion bridges and a difference-of-two-log-\(R\) score increment,
\[
\mathsf S_{k-1,k}^l = \nabla \log(R_{\theta,k-1,k}^l(\cdot))-\nabla \log(R_{\theta,k-1,k}^l(\cdot)),
\]
to build a practical martingale-posterior recursion. Its main theorem proves that, under assumptions (A1)–(A2),
\[
\sup_{n\ge 1}\check{\mathbb E}\left[\|\theta_n^l-\theta_n\|^2\right]\le C\Delta_l,
\]
so the discretized algorithm approximates the ideal continuous-time martingale posterior with \(\mathcal O(\Delta_l)\) mean-square error. The paper reports “orders of magnitude speed up versus state-of-the-art MCMC algorithms” [2604.27603].

## 6. Limitations, asymptotics, and terminology

Martingale posteriors remain a heterogeneous family of posterior-like constructions rather than a single closed formalism. Several limitations recur across the literature. Rossell’s critique emphasizes loss of dynamic coherence, sensitivity of the implied prior, possible distortion of Bayes factors, unresolved computational advantage, and the fact that high-dimensional settings may retain posterior uncertainty even as \(n\to\infty\) unless the framework is suitably adjusted [2303.02403]. The Bernoulli predictive-completeness results likewise show that first-moment martingale coherence identifies one-step prediction but not the full hierarchy of multi-step predictive distributions, posterior variances, or other nonlinear functionals [2603.00661].

At the same time, recent asymptotic work has clarified where martingale posteriors recover familiar large-sample behavior. Parametric martingale posteriors now have predictive central limit theory and a Bernstein–von Mises theorem [2410.17692]. Quantile martingale posteriors admit posterior consistency and contraction results, together with a Gaussian-process asymptotic approximation [2406.03358]. This suggests that, in regular settings, predictive-first uncertainty quantification can be studied with tools parallel to those used for standard Bayes, even though the inferential object is constructed differently.

A final terminological distinction is necessary. “M-posteriors” are not martingale posteriors. In that literature, the “M” stands for M-estimation: the posterior is
\[
\pi_n^\rho(\theta \mid F_n)
=
\frac{\exp\!\left(-\sum_{i=1}^n \rho(X_i,\theta)\right)\pi(\theta)}
{\int_\Theta \exp\!\left(-\sum_{i=1}^n \rho(X_i,\theta')\right)\pi(\theta')\,d\theta'},
\]
and there is “no martingale structure, optional stopping argument, filtration-based update rule, or posterior process indexed by time” [2510.01358]. The overlap is therefore only at the level of generalized or posterior-like inference.

Taken together, the literature presents martingale posteriors as predictive-sequential uncertainty distributions whose defining coherence is martingale structure rather than prior-to-posterior updating. Standard Bayes is recovered when the predictive sequence is the Bayesian posterior predictive; outside that case, martingale posteriors form a broader class of predictive Bayes procedures, ranging from copula-based nonparametrics and score-driven recursions to semiparametric, quantile, neural, federated, and diffusion-specific constructions [2103.15671].

Source: https://www.emergentmind.com/topics/martingale-posteriors