Papers
Topics
Authors
Recent
Search
2000 character limit reached

Von Mises Functionals: Theory & Applications

Updated 6 July 2026
  • Von Mises functionals are statistical functionals defined on convex sets that use directional derivatives along feasible mixture paths to capture local behavior in probability spaces.
  • They provide a framework for robust and semiparametric statistics by enabling influence functions and local Taylor-type expansions, crucial for asymptotic analysis.
  • Modern applications include bagging-induced smoothing and Bernstein–von Mises theorems that facilitate efficient inference in models such as regression and covariance analysis.

Searching arXiv for recent and foundational papers on von Mises functionals and related Bernstein–von Mises/statistical calculus work. Von Mises functionals are statistical functionals whose local behavior is analyzed through directional expansions on spaces of probability distributions or other convex parameter domains. In the classical statistical setting, the derivative is taken along feasible mixture paths inside the domain rather than along arbitrary directions in an ambient linear space, because spaces of distributions are typically convex, non-open, and often have empty interior. This calculus was introduced by von Mises to study asymptotic behavior and later became central in robust statistics and semiparametric statistics, including the influence-function framework associated with Hampel, Huber, Fernholz, Gill, and Reeds (Cerreia-Vioglio et al., 2024). In modern work, the same perspective underlies first- and second-order functional expansions, semiparametric Bernstein–von Mises theorems, and even smoothing phenomena induced by bagging (Castillo et al., 2013, Buja et al., 2016).

1. Convex-domain differentiability and the classical statistical calculus

A defining feature of von Mises functionals is that their natural domain is not an open subset of a linear space. The relevant domain is instead a convex set CC in a vector space XX, often a space of probability measures. For such f:CRf:C\to\mathbb R, the basic derivative is the directional derivative along line segments,

Df(x;y)=limt0f((1t)x+ty)f(x)t,Df(x;y)=\lim_{t\downarrow 0}\frac{f((1-t)x+ty)-f(x)}{t},

whenever the finite limit exists (Cerreia-Vioglio et al., 2024). This formulation keeps perturbations inside the feasible set and is therefore adapted to statistical models defined by convex constraints.

Two elementary properties organize this calculus. First, Df(x;x)=0Df(x;x)=0. Second, the derivative is homogeneous along segments: Df(x;(1a)x+ay)=aDf(x;y),a[0,1].Df\bigl(x;(1-a)x+ay\bigr)=a\,Df(x;y),\qquad a\in[0,1]. What is not automatic is affinity in the second argument. This leads to the distinction between weak affine differentiability, where Df(x;)Df(x;\cdot) is affine on CC, and affine differentiability, where this affine map extends to the ambient space XX (Cerreia-Vioglio et al., 2024).

If the extension exists, then there are xXx^*\in X^* and XX0 such that

XX1

and, using XX2, equivalently

XX3

The corresponding affine gradient XX4 is generally defined only up to the equivalence class

XX5

This nonuniqueness is intrinsic in probability models because constants annihilate differences of probability measures (Cerreia-Vioglio et al., 2024).

The relation to ordinary Gateaux differentiability is exact on algebraic interior points. If XX6, then affine differentiability at XX7 is equivalent to Gateaux differentiability at XX8, and the affine gradient class collapses to the usual Gateaux gradient. Outside XX9, affine differentiability is the appropriate extension. This matters statistically because distribution spaces usually have empty algebraic interior, so standard Gateaux theory is often unavailable where von Mises calculus still applies (Cerreia-Vioglio et al., 2024).

2. Influence functions and local von Mises expansions

The canonical first-order application of the calculus is the influence function. For a functional f:CRf:C\to\mathbb R0, the influence function at f:CRf:C\to\mathbb R1 is

f:CRf:C\to\mathbb R2

If f:CRf:C\to\mathbb R3 is affinely differentiable, then

f:CRf:C\to\mathbb R4

where f:CRf:C\to\mathbb R5; under the normalization f:CRf:C\to\mathbb R6, one has f:CRf:C\to\mathbb R7 (Cerreia-Vioglio et al., 2024). In this form, the derivative becomes a local sensitivity map to infinitesimal contamination.

The same idea yields the local expansion usually called a von Mises expansion. In the affine setting, the first-order approximation is

f:CRf:C\to\mathbb R8

or, in the Fréchet-type form used in the paper,

f:CRf:C\to\mathbb R9

This is the abstract statement that a statistical functional is locally approximated by an affine functional, with a remainder negligible relative to perturbation size (Cerreia-Vioglio et al., 2024).

In semiparametric theory, the same pattern is written in Hilbert-space notation. A smooth functional Df(x;y)=limt0f((1t)x+ty)f(x)t,Df(x;y)=\lim_{t\downarrow 0}\frac{f((1-t)x+ty)-f(x)}{t},0 is assumed to admit an expansion

Df(x;y)=limt0f((1t)x+ty)f(x)t,Df(x;y)=\lim_{t\downarrow 0}\frac{f((1-t)x+ty)-f(x)}{t},1

with first-order term Df(x;y)=limt0f((1t)x+ty)f(x)t,Df(x;y)=\lim_{t\downarrow 0}\frac{f((1-t)x+ty)-f(x)}{t},2, second-order operator Df(x;y)=limt0f((1t)x+ty)f(x)t,Df(x;y)=\lim_{t\downarrow 0}\frac{f((1-t)x+ty)-f(x)}{t},3, and remainder Df(x;y)=limt0f((1t)x+ty)f(x)t,Df(x;y)=\lim_{t\downarrow 0}\frac{f((1-t)x+ty)-f(x)}{t},4 (Castillo et al., 2013). The literature distinguished in the source data considers both the purely first-order regime and a second-order regime needed for nonlinear functionals and lower-regularity settings.

Several standard examples fit this template. The Mann–Whitney functional has a weak affine differential and becomes affinely differentiable under continuity and bounded-support assumptions, with gradient Df(x;y)=limt0f((1t)x+ty)f(x)t,Df(x;y)=\lim_{t\downarrow 0}\frac{f((1-t)x+ty)-f(x)}{t},5. Quadratic functionals

Df(x;y)=limt0f((1t)x+ty)f(x)t,Df(x;y)=\lim_{t\downarrow 0}\frac{f((1-t)x+ty)-f(x)}{t},6

are affinely differentiable with

Df(x;y)=limt0f((1t)x+ty)f(x)t,Df(x;y)=\lim_{t\downarrow 0}\frac{f((1-t)x+ty)-f(x)}{t},7

linking von Mises calculus directly to Df(x;y)=limt0f((1t)x+ty)f(x)t,Df(x;y)=\lim_{t\downarrow 0}\frac{f((1-t)x+ty)-f(x)}{t},8-statistics and quadratic empirical functionals (Cerreia-Vioglio et al., 2024).

3. Higher-order structure and the smoothing effect of bagging

A modern extension of von Mises functionals arises from bagging. Let Df(x;y)=limt0f((1t)x+ty)f(x)t,Df(x;y)=\lim_{t\downarrow 0}\frac{f((1-t)x+ty)-f(x)}{t},9 be a real-valued statistical functional, and let

Df(x;x)=0Df(x;x)=00

where Df(x;x)=0Df(x;x)=01. This is the generalized bagged functional obtained by averaging the original statistic over an Df(x;x)=0Df(x;x)=02-tuple of i.i.d. draws from Df(x;x)=0Df(x;x)=03 (Buja et al., 2016).

The central result is that the bagged functional is always smooth in the von Mises sense, with an expansion that is exact and finite of length Df(x;x)=0Df(x;x)=04. Specifically,

Df(x;x)=0Df(x;x)=05

where Df(x;x)=0Df(x;x)=06 is the Df(x;x)=0Df(x;x)=07-th Efron–Stein ANOVA interaction term of the raw statistic Df(x;x)=0Df(x;x)=08, and the corresponding von Mises coefficients satisfy

Df(x;x)=0Df(x;x)=09

Thus all influence functions of order Df(x;(1a)x+ay)=aDf(x;y),a[0,1].Df\bigl(x;(1-a)x+ay\bigr)=a\,Df(x;y),\qquad a\in[0,1].0 vanish (Buja et al., 2016).

This establishes a precise structural relation between the von Mises expansion of the bagged functional and the Efron–Stein ANOVA decomposition of the raw statistic. The raw statistic may be rough or unstable, but after bagging the resulting functional has finite-order smoothness. The resample size Df(x;(1a)x+ay)=aDf(x;y),a[0,1].Df\bigl(x;(1-a)x+ay\bigr)=a\,Df(x;y),\qquad a\in[0,1].1 acts as a smoothing parameter: smaller Df(x;(1a)x+ay)=aDf(x;y),a[0,1].Df\bigl(x;(1-a)x+ay\bigr)=a\,Df(x;y),\qquad a\in[0,1].2 gives more smoothing because the exact expansion is shorter. The paper explicitly notes that this is the opposite of the naive intuition that larger resamples should smooth more because they are closer to Df(x;(1a)x+ay)=aDf(x;y),a[0,1].Df\bigl(x;(1-a)x+ay\bigr)=a\,Df(x;y),\qquad a\in[0,1].3 (Buja et al., 2016).

A useful interpretation is that bagging is not merely a prediction-error device. It is also a smoothing operator on statistical functionals: higher-order perturbation sensitivity is suppressed beyond order Df(x;(1a)x+ay)=aDf(x;y),a[0,1].Df\bigl(x;(1-a)x+ay\bigr)=a\,Df(x;y),\qquad a\in[0,1].4, and instability in the original statistic is replaced by a controlled finite influence structure (Buja et al., 2016).

4. Bernstein–von Mises theory and efficient inference for functionals

Much of the contemporary importance of von Mises functionals comes from Bernstein–von Mises theory. In Gaussian white noise, density sampling, and general semiparametric models, posterior asymptotic normality for low-dimensional functionals is obtained by linearizing the functional and showing that the linear term dominates at the Df(x;(1a)x+ay)=aDf(x;y),a[0,1].Df\bigl(x;(1-a)x+ay\bigr)=a\,Df(x;y),\qquad a\in[0,1].5 scale (Castillo et al., 2012, Castillo et al., 2013).

In the Gaussian white noise model, the basic nonlinear-functional expansion is

Df(x;(1a)x+ay)=aDf(x;y),a[0,1].Df\bigl(x;(1-a)x+ay\bigr)=a\,Df(x;y),\qquad a\in[0,1].6

which is the nonparametric analogue of a first-order functional Taylor expansion (Castillo et al., 2012). For linear functionals

Df(x;(1a)x+ay)=aDf(x;y),a[0,1].Df\bigl(x;(1-a)x+ay\bigr)=a\,Df(x;y),\qquad a\in[0,1].7

the induced posterior is asymptotically normal when centered at the efficient estimator Df(x;(1a)x+ay)=aDf(x;y),a[0,1].Df\bigl(x;(1-a)x+ay\bigr)=a\,Df(x;y),\qquad a\in[0,1].8, with variance Df(x;(1a)x+ay)=aDf(x;y),a[0,1].Df\bigl(x;(1-a)x+ay\bigr)=a\,Df(x;y),\qquad a\in[0,1].9 (Castillo et al., 2012). For nonlinear functionals satisfying the same quadratic-remainder condition, posterior credible intervals have asymptotically exact frequentist coverage once the posterior contracts fast enough in Df(x;)Df(x;\cdot)0 (Castillo et al., 2012).

A broader semiparametric formulation uses a local asymptotic normality expansion for the likelihood together with a pathwise perturbation Df(x;)Df(x;\cdot)1 in the efficient influence direction. The decisive assumption is a no-bias or prior-invariance condition: the prior mass near Df(x;)Df(x;\cdot)2 and the shifted path Df(x;)Df(x;\cdot)3 must be asymptotically comparable, otherwise a Bernstein–von Mises limit can fail (Castillo et al., 2013). This is especially important for nonlinear functionals and adaptive priors.

In a multiscale formulation for regression and density estimation, once the posterior satisfies a weak Bernstein–von Mises theorem in the sequence space Df(x;)Df(x;\cdot)4, continuous linear functionals inherit Gaussian asymptotics by the continuous mapping theorem. The same framework yields posterior Donsker and posterior Kolmogorov–Smirnov theorems for posterior cumulative distribution functions, as well as multiscale posterior credible bands that are asymptotically exact frequentist confidence bands (Castillo et al., 2013).

5. Model-specific realizations: sieves, trees, matrices, and diffusions

The general calculus acquires different concrete forms in different models. In Gaussian regression with an increasing number of regressors, the semiparametric result is based on a first-order expansion around the sieve projection or sieve MLE. For a nonlinear functional Df(x;)Df(x;\cdot)5, the leading term is Df(x;)Df(x;\cdot)6, and the covariance of the Gaussian limit is

Df(x;)Df(x;\cdot)7

The second-order remainder must be negligible on the posterior concentration set, which makes the result a genuine posterior functional delta method based on a von Mises/Taylor expansion (Bontemps, 2010).

For BART in fixed-design nonparametric regression, the functional of interest is

Df(x;)Df(x;\cdot)8

with Df(x;)Df(x;\cdot)9 bounded Hölder. The efficient influence direction is the projection CC0 of CC1 onto the tree space, and the key condition is the no-bias bound

CC2

The paper emphasizes that adaptive priors may violate this condition; self-similarity of the truth is imposed to prevent the posterior from selecting partitions that are too smooth relative to the functional (Rockova, 2019).

For covariance and precision matrix functionals, the derivative becomes an influence matrix. A covariance functional CC3 is treated as approximately linear if

CC4

while a precision functional CC5 is linearized by

CC6

This framework yields Bernstein–von Mises theorems for entries of CC7 and CC8, quadratic forms, log-determinant, eigenvalues, and also Bayesian LDA and QDA functionals (Gao et al., 2014). A related finite-sample theory for covariance matrices gives nonasymptotic Gaussian approximations for smooth covariance functionals and for the Frobenius distance between spectral projectors, with the projector linearization governed by the reduced resolvent and the eigengap (Silin, 2017).

In periodic reversible diffusions,

CC9

the one-dimensional functional XX0 is assumed to satisfy

XX1

with negligible remainder on posterior localization sets (Giordano et al., 22 May 2025). The efficient direction is no longer a simple XX2 representer. It is the solution XX3 of the elliptic PDE

XX4

and the semiparametric efficiency bound is

XX5

This PDE-based efficient influence structure is a distinctive feature of the reversible diffusion model (Giordano et al., 22 May 2025).

6. Scope, misconceptions, and neighboring terminology

A recurring misconception is to identify von Mises functionals with ordinary Gateaux differentiability. The modern affine theory shows that this is correct only on the algebraic interior of the domain; on convex domains with empty interior, affine directional differentiability is the relevant notion (Cerreia-Vioglio et al., 2024). A second misconception is that any adaptive Bayesian procedure automatically yields a Bernstein–von Mises theorem for functionals. Several papers show the opposite: adaptive priors can fail because of semiparametric bias, non-negligible remainder terms, or incompatibility between the regularity of the truth and the regularity of the functional (Castillo et al., 2013, Rockova, 2019).

A third misconception is to equate “von Mises” in this context with the von Mises distribution of directional statistics. The eponym appears in distinct areas. One branch concerns functionals, derivatives, and expansions on spaces of distributions. Another concerns circular distributions such as the sine-skewed von Mises law

XX6

which is characterized through truncated conditional moments (Ahsanullah et al., 2019). A further neighboring usage appears in time-series analysis, where the generalized von Mises density is used as a spectral density and shown to be the maximum Shannon-entropy spectral distribution under fixed autocovariance constraints (Gatto, 2021). These are not von Mises functionals, although they share the same historical name.

Across the literature summarized here, the unifying content of a von Mises functional is local approximability by an influence representation on a statistically feasible perturbation path. Depending on the model, the first-order object may be an affine gradient on a convex set, an influence function, a linear representer in a Hilbert or multiscale space, an influence matrix for a covariance functional, a projected tree-space direction, or the solution of an elliptic PDE. The common role is the same: it controls asymptotic sensitivity, determines efficient variance, and supplies the expansion that turns a complex nonparametric or semiparametric functional into a tractable Gaussian limit (Cerreia-Vioglio et al., 2024, Giordano et al., 22 May 2025).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Von Mises Functionals.