Papers
Topics
Authors
Recent
Search
2000 character limit reached

Misspecified Bayesianism

Updated 7 July 2026
  • Misspecified Bayesianism is Bayesian updating with a model that excludes the true data-generating process, leading to inference based on the closest approximation.
  • It emphasizes the role of pseudo-true parameters and altered uncertainty quantification, challenging classical consistency measures.
  • The field integrates techniques like variational, generalized, and calibrated posteriors to address and mitigate the effects of model error.

Searching arXiv for recent and foundational papers on misspecified Bayesianism and closely related misspecified Bayesian inference. Misspecified Bayesianism denotes Bayesian updating or inference conducted under a subjective or statistical model that is not the true data-generating mechanism. In one formulation, an agent is a misspecified Bayesian if she updates her belief using Bayes’ rule given a subjective, possibly misspecified model of her signals; in another, the true distribution P0P_0 lies outside the inferential model class, so the posterior no longer targets truth but a best approximation within that class (Molavi, 30 Jul 2025, Bochkina, 2022). Across asymptotic statistics, econometric learning, variational inference, generalized Bayes, and simulation-based inference, the central theme is stable: Bayes’ rule can remain internally coherent under model error, but its target, dynamics, and uncertainty quantification are altered by misspecification.

1. Conceptual domain

In the statistical literature, misspecification is the case in which P0{P(θ),θΘ}P_0 \notin \{P(\cdot\mid\theta),\theta\in\Theta\}, so the model class does not contain the truth (Bochkina, 2022). In the decision-theoretic formulation of misspecified Bayesianism, one starts with a true prior μ0\mu_0^*, a random posterior μ1\mu_1, and asks whether there exists some subjective joint distribution Q\mathbb{Q} over states and signals such that the observed belief change is generated by Bayesian conditioning under Q\mathbb{Q} (Molavi, 30 Jul 2025). The subject therefore includes both inferential misspecification of a statistical model and behavioral misspecification of an agent’s signal model.

The practical motivation is explicit throughout the literature. Models are “rarely well-specified in practice” for variational Bayes (Wang et al., 2019); approximate models are used “for computational convenience” in the review of misspecified Bernstein–von Mises theory (Bochkina, 2022); and misspecification is “unavoidable” in model selection when one has no knowledge of the true model or omits true predictors (Lv et al., 2010). A recurring implication is that the scientific meaning of a posterior must be separated from the formal act of Bayesian updating itself.

2. Pseudo-truth, posterior concentration, and asymptotic limits

A standard object under misspecification is the pseudo-true parameter, defined as a Kullback–Leibler projection. One prominent formulation is

θ=argminθKL ⁣(p0(x)p(xθ)),\theta^*=\arg\min_\theta \mathrm{KL}\!\left(p_0(x)\,\|\,p(x\mid \theta)\right),

which is the point targeted by the exact posterior under misspecified Bernstein–von Mises theory (Wang et al., 2019). The review of misspecified asymptotics describes the same idea as the “best parametric approximation” and emphasizes that posterior concentration at θ\theta^\star is weaker than consistency for a true parameter (Bochkina, 2022).

Under regular LAN-type conditions, misspecified Bernstein–von Mises results still yield local Gaussianity around the pseudo-true point. The posterior can converge to a point mass at θ\theta^\star and, after centering and scaling, become asymptotically normal, but the covariance is no longer the Fisher inverse in the classical well-specified sense (Bochkina, 2022). This is the core asymptotic statement behind much of misspecified Bayesianism: when the model is wrong, Bayesian learning can still stabilize, but it stabilizes around the least-wrong element of the model.

The same idea extends beyond iid regular parametric settings. For dependent and possibly non-Markovian data, the key quantity is the KL divergence rate

h(θ)=limt1tlogP(X1t)Pθ(X1t),h(\theta)=\lim_{t\to\infty}\frac{1}{t}\log \frac{P(X_1^t)}{P_\theta(X_1^t)},

and posterior mass concentrates on hypotheses whose divergence rates are minimal among those with prior mass (Shalizi, 2009). If a measurable set P0{P(θ),θΘ}P_0 \notin \{P(\cdot\mid\theta),\theta\in\Theta\}0 satisfies P0{P(θ),θΘ}P_0 \notin \{P(\cdot\mid\theta),\theta\in\Theta\}1, then P0{P(θ),θΘ}P_0 \notin \{P(\cdot\mid\theta),\theta\in\Theta\}2 almost surely. This establishes a general “least-wrong model” principle even when all hypotheses are false and the data are dependent.

3. Uncertainty quantification, efficiency, and model selection

The main technical difficulty after concentration is calibration. Under misspecification, the curvature of the log-likelihood and the variability of the score typically differ:

P0{P(θ),θΘ}P_0 \notin \{P(\cdot\mid\theta),\theta\in\Theta\}3

In the well-specified case P0{P(θ),θΘ}P_0 \notin \{P(\cdot\mid\theta),\theta\in\Theta\}4, but under misspecification this equality generally fails, and the relevant frequentist covariance is the sandwich form P0{P(θ),θΘ}P_0 \notin \{P(\cdot\mid\theta),\theta\in\Theta\}5 rather than the posterior covariance P0{P(θ),θΘ}P_0 \notin \{P(\cdot\mid\theta),\theta\in\Theta\}6 (Bochkina, 2022). The standard consequence is that credible sets can be asymptotically invalid or inefficient even when the posterior center is sensible.

Several correction strategies are surveyed in the literature. These include learning-rate or fractional posteriors, curvature adjustment, loss-likelihood bootstrap, bagged posterior constructions, and composite-likelihood or pseudo-likelihood recalibration (Bochkina, 2022). A related performance-bound perspective is given by the misspecified Bayesian Cramér–Rao bound, which defines a Bayesian pseudotrue parameter P0{P(θ),θΘ}P_0 \notin \{P(\cdot\mid\theta),\theta\in\Theta\}7 as a KL projection of the true model onto the assumed model and bounds the mean-square error of an estimator relative to P0{P(θ),θΘ}P_0 \notin \{P(\cdot\mid\theta),\theta\in\Theta\}8 rather than to the true parameter directly (Tang et al., 2023). This formalizes the idea that mismatch changes the estimand itself.

Misspecification also alters information criteria. In misspecified generalized linear models, the covariance contrast matrix

P0{P(θ),θΘ}P_0 \notin \{P(\cdot\mid\theta),\theta\in\Theta\}9

measures the discrepancy between model-implied curvature and score covariance, and asymptotic expansions of the Bayesian and KL principles lead to μ0\mu_0^*0, μ0\mu_0^*1, and especially

μ0\mu_0^*2

(Lv et al., 2010). The criterion μ0\mu_0^*3 is explicitly interpreted as a decomposition into goodness of fit, model complexity, and model misspecification.

4. Variational, generalized, and calibrated posteriors

Misspecified Bayesianism is not confined to exact posteriors. For variational Bayes, the central asymptotic result is that the VB posterior remains asymptotically normal under misspecification and centers at the same KL projection as the exact posterior (Wang et al., 2019). The centered and rescaled VB posterior has a diagonal precision matrix with the same diagonal entries as the exact posterior precision, which captures the familiar VB under-dispersion. The same paper proves that, in posterior predictive distributions, the model misspecification error dominates the variational approximation error asymptotically, providing one explanation for the empirical fact that VB often matches MCMC in predictive accuracy.

A second line of work replaces ordinary Bayes by a tempered or generalized posterior

μ0\mu_0^*4

In misspecified linear regression with heteroskedastic data, standard Bayes can become inconsistent, placing mass on increasingly high-dimensional models, while the SafeBayes method chooses a learning rate μ0\mu_0^*5 from the data using a prequential criterion designed to detect lack of cumulative posterior concentration (Grünwald et al., 2014). For generalized linear models, μ0\mu_0^*6-generalized Bayes is shown to concentrate around the best approximation of the truth inside the model for specific μ0\mu_0^*7, even under severely misspecified noise, provided the tails of the true distribution are exponential and the conditional mean is correctly modeled (Heide et al., 2019).

A third line modifies the posterior kernel itself to repair calibration. The Q-posterior replaces the usual likelihood or loss by a quadratic-form loss in the empirical score or estimating equation,

μ0\mu_0^*8

and yields asymptotically reliable uncertainty quantification for both likelihood-based and loss-based posteriors (Frazier et al., 2023). In this framework, correct uncertainty quantification is obtained without learning-rate tuning, bootstrapping, or a Gaussian post-processing step.

The predictive objective under misspecification can also depart sharply from the inferential objective. PACμ0\mu_0^*9-Bayes is motivated by the claim that the Bayesian posterior minimizes an inferential risk that only bounds the predictive risk, and that misspecification induces a gap between them (Morningstar et al., 2020). Closely related second-order PAC-Bayes analyses argue that Bayesian model averaging is suboptimal for predictive performance when the model family is misspecified, because Bayes tends to concentrate on the best single model rather than on the posterior distribution that gives the best posterior predictive mixture (Masegosa, 2019).

5. Dynamic learning and sequential decision-making

In economics and learning theory, misspecified Bayesianism is studied as a dynamic process. For a single infinitely lived Bayesian agent who repeatedly chooses actions and updates by Bayes’ rule under a misspecified model, the key state variable is the empirical frequency of past actions,

μ1\mu_10

A uniform law of large numbers shows that the posterior asymptotically depends on history through μ1\mu_11, and the posterior mass concentrates on the set of KLD minimizers μ1\mu_12 corresponding to the current action frequencies (Esponda et al., 2019). The long-run dynamics are characterized by a differential inclusion,

μ1\mu_13

which permits convergence to equilibria, convergence to mixed steady states, or nonconvergent cycles.

History dependence can produce sharper failures. When an agent is wrong about the time lag between actions and feedback, the misspecification creates attribution errors: outcomes are credited to the wrong past actions (Li et al., 2020). If actions converge, the misspecification has no long-run effect and the agent must converge to the optimal action. If actions cycle, however, the same wrong-lag belief can produce arbitrarily large long-run inefficiencies, because repeated action switches keep the posterior from recovering.

Prior misspecification in sequential decision rules admits quantitative sensitivity bounds. For Thompson sampling with a misspecified prior, if the true and assumed priors differ by total variation distance μ1\mu_14 and the prior is μ1\mu_15-bounded, then

μ1\mu_16

(Simchowitz et al., 2021). The same analysis extends to a broader family of Bayesian decision-making algorithms, including a Monte-Carlo implementation of knowledge gradient, and to Bayesian POMDPs.

6. Restricted likelihoods, modularity, simulation-based inference, and testability

One response to misspecification is to change what is updated. A review of misspecified generative models distinguishes three families of meaningful procedures when the analyst is unwilling to act as if the full model is correct: restricted likelihood methods based on a non-sufficient summary, modular inference methods that cut feedback between coupled submodels, and reference-model projection methods that define inference for a simplified model through a well-specified reference model (Nott et al., 2023). In the canonical two-module setting, the cut posterior blocks feedback from a suspect module and replaces the full posterior for μ1\mu_17 by μ1\mu_18.

This modular theme is generalized in a recent treatment of Bayesian networks, where cut-posteriors are defined as a space of belief updates parameterized by a partition of data nodes into modules, a topological ordering of those modules, and per-parameter decisions about where each shared hidden parameter is updated or merely conditioned on (Li, 2023). A notable claim is that any cut-posterior has local computation only, so misspecification can be managed by deliberately preventing unreliable modules from contaminating trusted ones.

Simulation-based Bayesian inference is especially vulnerable because the simulator is often only approximate. A recent review organizes misspecification-robust SBI around three strategies: robust summary statistics, generalized Bayesian inference, and error modelling with adjustment parameters (Kelly et al., 16 Mar 2025). Within that literature, Bayesian synthetic likelihood is a particularly clear cautionary case: misspecification is defined at the summary level, and the BSL posterior can become bimodal, boundary-concentrated, flat, or asymptotically non-Gaussian, even placing substantial mass on parameter values that do not match the observed summaries well (Frazier et al., 2021). Robust BSL is proposed as a mitigation.

Diagnostics for misspecification can also be made external to the posterior. CARMEN uses a probabilistic classifier to discriminate between simulated and observed data, interprets classifier odds as a density-ratio estimate, and uses the expected log ratio as an estimate of the negative KL divergence from the true distribution to the model (Thomas et al., 2019). The same diagnostic can be used both for model comparison and for choosing a tempering level in generalized Bayes.

The most stringent question is whether an observed belief change is falsifiably non-Bayesian once misspecification is allowed. The characterization result in “Misspecified Bayesianism” states that a prior/posterior pair μ1\mu_19 is consistent with misspecified Bayesianism if and only if there exists a measurable partition of posterior space such that, on every cell with positive mass, the prior contains a grain of the cell-average posterior (Molavi, 30 Jul 2025). The grain condition means

Q\mathbb{Q}0

for some Q\mathbb{Q}1 and some probability measure Q\mathbb{Q}2, equivalently a bounded Radon–Nikodym derivative. Under correct specification this reduces to Bayes plausibility. On finite state spaces with full-support priors, essentially every posterior distribution is consistent with misspecified Bayesianism; on compact spaces, support inclusion largely governs consistency; on unbounded spaces, posteriors with heavier tails than the prior can be ruled out. The resulting controversy is explicit: many seemingly non-Bayesian updating rules may be observationally equivalent to Bayesian updating under misspecified beliefs, and testing Bayes’ rule in isolation can therefore be difficult.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (19)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Misspecified Bayesianism.