Papers
Topics
Authors
Recent
Search
2000 character limit reached

Change Point Detection and Localization in High-Dimensional Time Series

Published 14 Aug 2026 in stat.ME and math.ST | (2608.14344v1)

Abstract: We present new inference tools for change point detection in high-dimensional time series. We discuss two distinct statistical applications: First, sequential change point testing in an incoming data-stream. Second, retrospective localization of multiple changes, with confidence intervals at a globally controlled error level. Test statistics are built on the maximum norm to generate power against sparse and asynchronous changes. Both problems are tackled by related multiscale statistics that search for changes in the data at many different levels of resolution. For fixed dimension, our statistical approaches can be validated using traditional Hölderian invariance principles. In this paper, we present the high-dimensional analogue: Hölder-Gauss-approximations, which can be (roughly) interpreted as the Gaussian approximation for a Hölder-norm of the high-dimensional partial sum process. Such approximations are of interest beyond change point detection and can be used for other problems such as for stationarity testing in high dimensions. We evaluate finite-sample performance in a simulation study and give an application to air contamination due to wildfires in California, which occurs asynchronously across a panel of measuring stations.

Summary

  • The paper introduces Hölder-Gauss approximations that support multiscale maximum-norm tests in dimensions growing stretched-exponentially with sample size.
  • The proposed sequential monitor detects sparse and asynchronous changes with controlled error rates, and simulations show substantially shorter delays than a leading benchmark.
  • The retrospective COMBI-SCAN procedure provides globally controlled localization intervals for multiple changes, while wildfire data illustrate earlier detection of transient PM2.5 increases.

This paper develops inference tools for detecting and localizing changes in the mean parameter of a high-dimensional time series XnRdX_n \in \mathbb{R}^d (2608.14344). Two related statistical problems are addressed: sequential change point testing on an incoming data stream relative to a stationary training sample, and retrospective localization of multiple changes with confidence intervals at a globally controlled error level. Both procedures are built on multiscale, maximum-norm statistics designed to be powerful against sparse and asynchronous changes, and both are justified by a new mathematical device the authors call a Hölder-Gauss approximation — a high-dimensional analogue of classical Hölderian invariance principles.

Hölder-Gauss approximations

The mathematical core of the paper generalizes Hölderian invariance principles, which for fixed dimension state that the partial sum process converges to Brownian motion in Hölder topology, thereby justifying weighted multiscale statistics. The high-dimensional analogue approximates the distribution of weighted maxima of increments of the dd-dimensional partial sum process by the corresponding functional of Gaussian data, uniformly in the dimension d=d(N)d = d^{(N)}.

The proof strategy combines two ingredients. On sufficiently long time scales, the statistic is approximated using high-dimensional Gaussian approximation results for max-norm functionals, in the spirit of Chernozhukov–Chetverikov–Kato-type arguments. On short scales, where sums are not approximately normal, the Hölder weighting makes the contribution negligible, which is established via concentration inequalities. This mirrors the vanishing Hölder modulus of continuity in classical fixed-dimensional results.

Under sub-Gaussian errors with uniformly bounded Ψ2\Psi_2-Orlicz norms and variance bounded away from zero and infinity, the dimension may grow at a stretched-exponential rate: dexp(NC1)d \le \exp(N^{C_1}) with C1<(12β)/(1218β)C_1 < (1-2\beta)/(12-18\beta), where β[0,1/2)\beta \in [0,1/2) is the Hölder weight parameter. The approximation holds uniformly over the open-ended monitoring horizon, which is a nontrivial point: the proof decomposes the detector into an early part (where Gaussian approximation applies) and a late part (controlled by a tail bound on the modulus of continuity on a non-compact domain), playing the role of a Hájek–Rényi inequality.

A conceptual refinement shows that, in the Gaussian case, the discrete grid of scan locations can be replaced by the supremum over all weighted increments of a dd-dimensional Brownian motion, bringing the result fully in line with classical Hölderian invariance principles. The authors note this corollary is mainly of theoretical interest, since calibration in practice uses the discrete Gaussian statistic.

Sequential monitoring

The sequential procedure compares a training sample {X1,,XN}\{X_1,\dots,X_N\}, assumed change-free, against subsequently arriving data, following the open-ended monitoring framework of Chu, Stinchcombe, and White. The detector is a CUSUM-type statistic scanning over all possible monitoring blocks <k\ell < k, aggregated via a weighted maximum with Hölder weight dd0, and taking the max-norm over components. Larger dd1 emphasizes recent data and short scales, yielding faster detection; the price is a stricter dimension constraint, though permissible growth remains stretched-exponential for any dd2.

The main theoretical guarantees are: (i) asymptotically exact level under the global null via the Hölder-Gauss approximation and a parametric bootstrap calibrated on Gaussian data with the empirical training covariance; (ii) consistency whenever dd3 for some changed component; and (iii) detection delay bounds dd4. For polynomially growing dimension, dd5 can approach dd6 and delays of order dd7 with arbitrarily small dd8 are achievable. The authors contrast this with the max-norm monitoring method of Gösmann et al., whose delays are dd9 — substantially longer — and whose finite-horizon Gaussian approximation does not by itself yield open-ended guarantees. The paper also derives a componentwise set estimator of the changed coordinates with asymptotically controlled false-inclusion probability at level d=d(N)d = d^{(N)}0; the estimator is monotone in time, so a declared change is never revoked.

Retrospective segmentation

The second application is a COMBI-SCAN procedure for offline segmentation: a symmetric scan statistic d=d(N)d = d^{(N)}1 with Hölder scaling d=d(N)d = d^{(N)}2 is computed componentwise over location-bandwidth pairs d=d(N)d = d^{(N)}3, and intervals exceeding a calibrated threshold are retained, with disjointness enforced to avoid double detection. The exhaustive grid has cardinality d=d(N)d = d^{(N)}4; a geometrically thinned grid reduces this to d=d(N)d = d^{(N)}5, an important saving given that scans run across all d=d(N)d = d^{(N)}6 components. Thresholds are obtained from a parametric bootstrap using first-difference-based covariance estimation, which requires the contribution of mean changes to the empirical covariance to be negligible.

The procedure satisfies a "weak localization property": with asymptotic probability at least d=d(N)d = d^{(N)}7, every reported interval contains a true change. Under the multiscale signal condition d=d(N)d = d^{(N)}8, a "strong localization property" holds: every sufficiently high and well-separated change is detected and isolated in a unique interval. Interval widths scale as d=d(N)d = d^{(N)}9, so for polynomial dimension and bounded jump sizes, widths grow as Ψ2\Psi_20 with arbitrarily small Ψ2\Psi_21.

The paper also develops, in the appendix, a variant with logarithmic weighting Ψ2\Psi_22, Ψ2\Psi_23, motivated by Lévy's modulus of continuity. This weakens the detectability condition to Ψ2\Psi_24, close to the essentially optimal Ψ2\Psi_25 separation for fixed dimension. The logarithmic variant is worked out under Ψ2\Psi_26-mixing dependence, demonstrating that the Hölder-Gauss machinery extends beyond independence; for the polynomial weighting under independence the theory is cleanest, and the authors remark that analogous results under physical dependence seem likely but are not established. The logarithmic weighting is only meaningful when Ψ2\Psi_27 grows at most polynomially; at stretched-exponential dimension growth, the required separation becomes polynomial anyway, so the benefit disappears.

Finite-sample performance

Simulations compare the sequential detector against Gösmann et al. under Gaussian and uniform errors, with covariance matrix Ψ2\Psi_28, Ψ2\Psi_29, dimensions up to dexp(NC1)d \le \exp(N^{C_1})0, and dexp(NC1)d \le \exp(N^{C_1})1. Null rejection rates are close to nominal across all dexp(NC1)d \le \exp(N^{C_1})2 and dimensions, supporting the choice dexp(NC1)d \le \exp(N^{C_1})3 used throughout. Power comparisons favor the proposed method in sparse settings (dexp(NC1)d \le \exp(N^{C_1})4 changed component) and for late changes; for instance, with dexp(NC1)d \le \exp(N^{C_1})5, dexp(NC1)d \le \exp(N^{C_1})6, dexp(NC1)d \le \exp(N^{C_1})7, dexp(NC1)d \le \exp(N^{C_1})8, the proposed method attains power 0.947 versus 0.335 for the benchmark. The benchmark retains a slight edge for early changes in dense settings (dexp(NC1)d \le \exp(N^{C_1})9) at small jump sizes, a concession the authors make explicit. Average detection delays are roughly halved relative to the benchmark for late changes (e.g., 28.59 versus 62.75 at C1<(12β)/(1218β)C_1 < (1-2\beta)/(12-18\beta)0).

Application to California wildfire data

The sequential procedure is applied to daily PM2.5 measurements from EPA monitoring stations in California for 2018–2020, with C1<(12β)/(1218β)C_1 < (1-2\beta)/(12-18\beta)1 training days and 69–78 stations per year. The authors are explicit that this is a "pseudo-real-time" analysis under stylized assumptions: PM2.5 series are not perfectly independent and wildfire-induced changes are transient rather than persistent, contradicting the at-most-one-permanent-change model. The proposed method and the benchmark agree on significance for roughly 91–100% of stations, but the proposed method detects earlier in the overwhelming majority of cases, often by more than a week. The largest detection gaps — up to 100 days — are interpreted not as delay differences but as a power advantage against transient changes: in 2018, three geographically proximate stations exhibit two PM2.5 spikes, with the proposed method catching the summer wildfire spike that the benchmark misses, detecting only a later winter episode.

Limitations and open questions

Several limitations are stated in the paper. The main theory assumes i.i.d. sub-Gaussian errors; extensions to C1<(12β)/(1218β)C_1 < (1-2\beta)/(12-18\beta)2-mixing are given in the appendix for the segmentation statistic, but the corresponding results for sequential monitoring under dependence are not worked out, and long-run covariance estimation under dependence — which the practical procedures would require — is not addressed. The dimension condition ties C1<(12β)/(1218β)C_1 < (1-2\beta)/(12-18\beta)3 to a stretched-exponential growth bound on C1<(12β)/(1218β)C_1 < (1-2\beta)/(12-18\beta)4; the interaction between the weight parameter and admissible dimension is a genuine constraint rather than an artifact. The strong localization guarantee depends on the separation-height condition, and the first-difference covariance estimator in segmentation is valid only when changes contribute negligibly to the empirical covariance. Finally, the logarithmic-weight variant's benefits are established only for polynomially bounded dimension, and the authors note without proof that analogous logarithmic weighting for sequential monitoring should yield near-logarithmic delays; formalizing this, and establishing optimality of the detection delays and interval widths in general, remain open.

Conclusion

The paper unifies sequential monitoring and retrospective segmentation of high-dimensional time series under a single multiscale, max-norm methodology, and supplies the theoretical foundation through new Hölder-Gauss approximations valid for dimensions growing stretched-exponentially in the sample size. The resulting procedures carry asymptotic level and false-discovery guarantees, short detection delays relative to existing max-norm monitoring, and simultaneous confidence intervals for multiple change locations — an inferential object of SMUCE/NSP type previously unavailable in high dimensions. The simulation and wildfire analyses support the practical relevance of the approach, while the dependence assumptions and the dimension-weight interplay delineate the boundaries of the current theory.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.