Papers
Topics
Authors
Recent
Search
2000 character limit reached

Pre-Averaged Bipower Variation

Updated 26 January 2026
  • Pre-averaged bipower variation is a robust method for estimating volatility-related functionals in noisy, jump-infused financial data.
  • It combines pre-averaging and bipower functionals to suppress large shocks and microstructure noise, ensuring consistent inference.
  • The technique uses threshold filtering to remove extreme increments, optimizing the bias-variance trade-off in high-frequency estimation.

Pre-averaged bipower variation is a technique designed to robustly estimate volatility-related functionals of stochastic processes in the presence of jumps and high-frequency noise. The method forms part of a broader class of estimators in high-frequency financial econometrics, aimed at extracting meaningful integrated variational quantities from observed data streams where large, non-continuous increments (jumps) and microstructure noise can strongly bias conventional realized variation-based statistics. Pre-averaged bipower variation operates by combining pre-averaging—where raw increments are smoothed over small rolling windows—with bipower functionals, which intrinsically suppress the influence of large shocks. This dual mechanism yields estimators that, under appropriate asymptotic regimes and mild noise/jump assumptions, provide consistent and robust inference for integrated volatility and related quantities, even when the underlying process exhibits heavy tails or infinite activity.

1. Stochastic Setting and Noise Model

Consider a dd-dimensional process XεX^\varepsilon on [0,1][0, 1] defined by the stochastic differential equation:

Xtε=x+0tb(Xsε,θ0)ds+εQtε,X_t^\varepsilon = x + \int_0^t b(X_s^\varepsilon, \theta_0)\, ds + \varepsilon Q_t^\varepsilon,

where xRdx \in \mathbb{R}^d, θ0Θ0Rp\theta_0 \in \Theta_0 \subset \mathbb{R}^p is an unknown drift parameter, b:Rd×ΘRdb: \mathbb{R}^d \times \Theta \to \mathbb{R}^d, and QεQ^\varepsilon is a semimartingale noise that converges uniformly in tt to a limiting semimartingale QQ as XεX^\varepsilon0, with XεX^\varepsilon1 (Doob-Meyer decomposition: XεX^\varepsilon2 finite variation, XεX^\varepsilon3 local martingale). A common instance is XεX^\varepsilon4 as a Lévy process with characteristic exponent:

XεX^\varepsilon5

subject to conditions ensuring uniform convergence of increments and finiteness of moments. These settings cover processes with both diffusion and jump components.

The drift function XεX^\varepsilon6 is assumed to satisfy regularity, growth, and identifiability conditions, in particular ensuring a positive definite information matrix:

XεX^\varepsilon7

where XεX^\varepsilon8 solves the deterministic ODE XεX^\varepsilon9 (Shimizu, 2015).

2. Threshold and Filtering Strategy

High-frequency increments are contaminated by rare but large jumps and noise. To suppress these effects, a thresholding filter is applied. For discrete sampled observations at times [0,1][0, 1]0 ([0,1][0, 1]1), set:

[0,1][0, 1]2

A threshold sequence [0,1][0, 1]3 is chosen with:

  • [0,1][0, 1]4,
  • [0,1][0, 1]5, ensuring [0,1][0, 1]6 but not excessively large, relative to [0,1][0, 1]7.

The (hard) increment filter is then

[0,1][0, 1]8

excluding increments "too large" to be explained by the continuous or small-noise part of the process, and thus likely to arise from jumps or extreme microstructure noise (Shimizu, 2015).

3. Definition of the Filtered (Pre-Averaged) Bipower Variation

While the provided data focuses on threshold-filtered least squares, the same mathematical principle underlies pre-averaged bipower statistics, which are constructed as follows: the pre-averaging step smooths increments to mitigate the influence of noise, and bipower variation is calculated using products of (possibly non-overlapping) absolute increments, respecting the filter:

[0,1][0, 1]9

where Xtε=x+0tb(Xsε,θ0)ds+εQtε,X_t^\varepsilon = x + \int_0^t b(X_s^\varepsilon, \theta_0)\, ds + \varepsilon Q_t^\varepsilon,0 denotes pre-averaged (smoothed) increments with the chosen window width. This estimator is robust to both infrequent large increments (jumps) and continuous small-noise contamination.

The filtered least squares estimator is the minimizer of the contrast function:

Xtε=x+0tb(Xsε,θ0)ds+εQtε,X_t^\varepsilon = x + \int_0^t b(X_s^\varepsilon, \theta_0)\, ds + \varepsilon Q_t^\varepsilon,1

with Xtε=x+0tb(Xsε,θ0)ds+εQtε,X_t^\varepsilon = x + \int_0^t b(X_s^\varepsilon, \theta_0)\, ds + \varepsilon Q_t^\varepsilon,2, and the threshold-type estimator

Xtε=x+0tb(Xsε,θ0)ds+εQtε,X_t^\varepsilon = x + \int_0^t b(X_s^\varepsilon, \theta_0)\, ds + \varepsilon Q_t^\varepsilon,3

A plausible implication is that bipower-type functionals calculated on pre-averaged, filtered increments would display similar robustness and efficiency properties (Shimizu, 2015).

4. Asymptotic Theory and Robustness

Under regularity and sampling assumptions, the filtered estimators enjoy strong theoretical guarantees:

  • Consistency: If Xtε=x+0tb(Xsε,θ0)ds+εQtε,X_t^\varepsilon = x + \int_0^t b(X_s^\varepsilon, \theta_0)\, ds + \varepsilon Q_t^\varepsilon,4 and Xtε=x+0tb(Xsε,θ0)ds+εQtε,X_t^\varepsilon = x + \int_0^t b(X_s^\varepsilon, \theta_0)\, ds + \varepsilon Q_t^\varepsilon,5, then Xtε=x+0tb(Xsε,θ0)ds+εQtε,X_t^\varepsilon = x + \int_0^t b(X_s^\varepsilon, \theta_0)\, ds + \varepsilon Q_t^\varepsilon,6 in probability as Xtε=x+0tb(Xsε,θ0)ds+εQtε,X_t^\varepsilon = x + \int_0^t b(X_s^\varepsilon, \theta_0)\, ds + \varepsilon Q_t^\varepsilon,7.
  • Asymptotic normality (for continuous noise): When the noise process Xtε=x+0tb(Xsε,θ0)ds+εQtε,X_t^\varepsilon = x + \int_0^t b(X_s^\varepsilon, \theta_0)\, ds + \varepsilon Q_t^\varepsilon,8 is Brownian, the scaled estimator Xtε=x+0tb(Xsε,θ0)ds+εQtε,X_t^\varepsilon = x + \int_0^t b(X_s^\varepsilon, \theta_0)\, ds + \varepsilon Q_t^\varepsilon,9 converges in distribution to a normal with variance xRdx \in \mathbb{R}^d0.
  • Heavy-tailed/jump robustness: For xRdx \in \mathbb{R}^d1 being an xRdx \in \mathbb{R}^d2-stable Lévy process or more generally, convergence is to an xRdx \in \mathbb{R}^d3-stable law whose scale depends on the Lévy measure xRdx \in \mathbb{R}^d4.
  • Moment convergence: If xRdx \in \mathbb{R}^d5 admits moments of all orders, then all moments of the normalized estimator converge to the corresponding functionals of the limiting random variable.

It is significant that these results obtain under minimal assumptions on the laws of jump or noise processes. Only a single threshold must be specified—no detailed knowledge of the jump measure xRdx \in \mathbb{R}^d6 is required—yielding a robust, model-free estimation framework (Shimizu, 2015).

5. Practical Tuning and Implementation

Theoretical tuning for the threshold is xRdx \in \mathbb{R}^d7 for xRdx \in \mathbb{R}^d8, ensuring both xRdx \in \mathbb{R}^d9 and θ0Θ0Rp\theta_0 \in \Theta_0 \subset \mathbb{R}^p0. In practical computations, θ0Θ0Rp\theta_0 \in \Theta_0 \subset \mathbb{R}^p1 or θ0Θ0Rp\theta_0 \in \Theta_0 \subset \mathbb{R}^p2 is effective:

  • θ0Θ0Rp\theta_0 \in \Theta_0 \subset \mathbb{R}^p3 suppresses jump bias with minimal variance inflation.
  • θ0Θ0Rp\theta_0 \in \Theta_0 \subset \mathbb{R}^p4 further lowers variance but may introduce limited bias for small θ0Θ0Rp\theta_0 \in \Theta_0 \subset \mathbb{R}^p5. The retained number of increments must satisfy θ0Θ0Rp\theta_0 \in \Theta_0 \subset \mathbb{R}^p6 to avoid degeneracy.

Table: Practical Impact of Thresholding

Estimator Variant Effect of Threshold Observed Bias/Variance Impact
Usual LSE No jump-filtering Large bias, large variance
Filtered (θ0Θ0Rp\theta_0 \in \Theta_0 \subset \mathbb{R}^p7) Moderate threshold Bias suppressed, variance reduced
Filtered (θ0Θ0Rp\theta_0 \in \Theta_0 \subset \mathbb{R}^p8) Aggressive threshold Lowest variance, possible extra bias for small θ0Θ0Rp\theta_0 \in \Theta_0 \subset \mathbb{R}^p9

Simulation-based QQ plots show that the filtered estimator is nearly Gaussian, even under infinite-activity stable noise, indicating significant finite-sample robustness (Shimizu, 2015).

6. Extension: Model-Free Application and Generalizations

This threshold-based filtering is inherently nonparametric, applicable regardless of whether b:Rd×ΘRdb: \mathbb{R}^d \times \Theta \to \mathbb{R}^d0 is a compound-Poisson, infinite-activity, or variance-gamma process. There is no assumption about the nature or distribution of jumps or noise beyond those enabling uniform convergence and moment existence. A plausible implication is that pre-averaged bipower variation, as a general class, can be extended to multivariate, state-dependent, or time-inhomogeneous noise settings, as long as appropriate filtering is applied to suppress large increments.

No explicit parametric modeling of discontinuities is required: the filter acts uniformly on observed increments, removing only those "too large" relative to small-noise expectations, and the same scheme applies across process types (Shimizu, 2015).

7. Mathematical and Statistical Justification

Rigorous proof of the estimator's properties relies on:

  • Proving the negligible contribution of filtered-out terms:

b:Rd×ΘRdb: \mathbb{R}^d \times \Theta \to \mathbb{R}^d1

  • Taylor expansion of the contrast's score and Hessian shows

b:Rd×ΘRdb: \mathbb{R}^d \times \Theta \to \mathbb{R}^d2

leading to explicit characterization of the limit.

  • The filtered empirical Hessian converges to the Fisher information, and the (filtered) score function converges to a stochastic integral against the noise process, justifying normal or stable limit laws depending on the noise structure.

All these constructions confirm that pre-averaged bipower variation and related thresholded estimators provide a robust solution for inference on stochastic processes with small noise but possibly large, non-negligible jumps (Shimizu, 2015).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Pre-Averaged Bipower Variation.