System-Wide Anomaly Score (SWAS)
- SWAS is a scalar anomaly signal that aggregates per-entity deviations from normality to summarize overall system behavior.
- It employs multiple aggregation methods—mean, variance, and KDE-plus-Wasserstein—to capture collective shifts and heterogeneous anomalies.
- The approach is operationalized in online quickest change-point detection, making it vital for monitoring dynamic systems in varied applications.
Searching arXiv for the cited paper and directly related anomaly-scoring work to ground the article. System-Wide Anomaly Score (SWAS) denotes a scalar anomaly signal intended to summarize the state of an entire system rather than that of an individual component. In its most explicit formulation, SWAS is introduced for quickest change point detection in multi-entity, multivariate time series, where it aggregates per-entity deviations from normality into a single time-indexed score that can be monitored for persistent or abrupt system-wide behavioral shifts (Kor et al., 22 Sep 2025). Closely related constructions appear in multi-sensor industrial monitoring, event-level anomaly detection, and calibrated nonparametric scoring, where the same underlying objective recurs: reduce distributed anomaly evidence to a single interpretable and thresholdable quantity (Alnegheimish et al., 21 Apr 2025, 0910.5461, Aarrestad et al., 2021).
1. Definition and problem class
In the formulation of "On Multi-entity, Multivariate Quickest Change Point Detection" (Kor et al., 22 Sep 2025), the observed data are a tensor
where is the number of sensors or variables, is the number of time steps, and is the number of entities. At time , entity produces an -dimensional observation vector
The changes of interest are system-wide rather than purely local. The motivating examples in the paper include panic in a crowd and regime shifts in coupled oscillators, where the salient event is not a single outlier but a collective behavioral transition. SWAS is therefore defined as a scalar time series that summarizes, at each time step, how abnormal the entire system appears after aggregating the deviations of all entities from learned normal behavior (Kor et al., 22 Sep 2025).
A recurrent misconception is that SWAS names a single fixed formula. In the cited quickest-change framework this is not the case: SWAS is a family of system-level scores obtained from the same pool of per-entity anomaly values but aggregated with different statistics, including mean, variance, and a KDE-plus-Wasserstein construction (Kor et al., 22 Sep 2025). This suggests that SWAS is better understood as a design pattern—system-level scalarization of distributed anomaly evidence—than as one canonical equation.
2. Construction from entity-level deviations
The central building block of SWAS in the multi-entity change-point framework is the Individual Deviation from Normality (IDfN) (Kor et al., 22 Sep 2025). The data are first z-score normalized per sensor over all entities and times: with
To capture temporal dependence, each entity’s trajectory is segmented into overlapping windows of length 0 with stride 1: 2 using replication padding for 3. A multivariate time series anomaly detection model is then trained only on normal data. The paper explicitly lists simple Autoencoder, TranAD, USAD, and OmniAnomaly as examples of such models (Kor et al., 22 Sep 2025).
Given reconstructed values 4, the per-entity deviation score is the mean squared reconstruction error
5
For online applicability, only the current-time score within the sliding window is retained, a strategy described as “last-point” scoring (Kor et al., 22 Sep 2025).
At each time 6, SWAS aggregates the set 7. Three aggregation strategies are explored.
The mean-based SWAS is
8
which measures the central tendency of system deviation and rises when many entities deviate simultaneously.
The variance-based SWAS is
9
which measures heterogeneity across entities and is sensitive to subsets of entities changing earlier or more strongly than others.
The distributional version estimates a kernel density over entity-level deviations,
0
with Gaussian kernel
1
and compares it with the reference density from normal training data,
2
The scalar score is then a Wasserstein-1 distance,
3
These constructions give SWAS three distinct semantics: aggregate deviation, aggregate dispersion, and distributional shift (Kor et al., 22 Sep 2025).
3. SWAS as a quickest change-point signal
In the cited framework, SWAS is not merely a descriptive statistic. It is the input to an online quickest change point detection pipeline (Kor et al., 22 Sep 2025). Baseline quantities are estimated from normal training data: 4
5
together with the reference KDE 6.
Deviation scores are then defined as
7
These are the quantities actually fed to CUSUM: 8 with 9.
Thresholding is adaptive rather than fixed. The accumulated statistic is first log-transformed,
0
then a KDE is fit to past transformed CUSUM values,
1
and a change is declared if
2
where 3 is chosen as an 4-quantile of density estimates computed from training CUSUM scores. The stopping rule is
5
The full pipeline is online: new observations arrive, windows are updated, reconstruction errors are computed, SWAS is aggregated, the deviation metric is updated, and CUSUM is checked in streaming mode. The paper also notes that mean and variance aggregation are 6 per time step, whereas KDE-plus-Wasserstein is more computationally expensive; window size 7 trades off detection stability against latency, and 8 is reported as a good compromise on crowd data (Kor et al., 22 Sep 2025).
4. Related global-scoring formulations
Functionally equivalent system-level scores appear in several neighboring literatures, although the term SWAS is not always used explicitly.
| Setting | Aggregated evidence | Scalar score |
|---|---|---|
| Multi-entity qCPD | Per-entity IDfN values | Mean, variance, or Wasserstein-based SWAS |
| M9AD | Per-sensor residual p-values | 0 |
| NN-graph anomaly detection | Neighborhood statistics of a test sample | 1 or 2 |
| Dark Machines | Event-level unsupervised detector output | Event anomaly score used to define signal regions |
In M3AD, the asset-level global anomaly score
4
aggregates sensor-wise p-values derived from residual models and is then Gamma-calibrated using training scores, with anomaly decisions made by thresholding either 5 directly or its calibrated tail probability (Alnegheimish et al., 21 Apr 2025). This is structurally close to SWAS: heterogeneous local evidence is mapped to a single scalar through calibrated aggregation.
In nearest-neighbor-graph anomaly detection, the score
6
or its 7-graph analogue estimates a multivariate p-value for the test sample and is thresholded at a desired false alarm level 8 (0910.5461). Here the global scalar is not built from multiple entities at a single time, but it serves the same role: a calibrated single-number summary of system atypicality.
In the Dark Machines challenge, each LHC event is assigned a scalar anomaly score by an unsupervised model, and model-independent signal regions are defined by thresholding this score event by event (Aarrestad et al., 2021). The paper also uses background-efficiency normalization when combining heterogeneous scores, effectively mapping them to a common operational scale before aggregation.
A plausible interpretation of MAD-EN is that a system-wide energy trace processed by a CNN can be converted into a scalar anomaly score through the model’s predicted probability of the anomalous classes, for example 9, where 0 denotes the benign class (Dipta et al., 2022). Because this formulation is an interpretation rather than the paper’s explicit terminology, it is better viewed as an SWAS-like reading of energy-based security monitoring than as a canonical definition.
5. Calibration, stability, and semantic meaning
A technically useful SWAS is not defined solely by aggregation. It also depends on calibration, statistical stability, and the meaning attached to high scores.
Calibration takes several forms in the cited literature. The quickest-change framework calibrates via baseline statistics and an adaptive KDE threshold on transformed CUSUM values (Kor et al., 22 Sep 2025). M1AD calibrates via a fitted Gamma distribution on training-time global scores (Alnegheimish et al., 21 Apr 2025). NN-graph methods produce scores that are asymptotically p-values under the nominal distribution (0910.5461). Dark Machines normalize detector outputs to background survival probabilities before combining them (Aarrestad et al., 2021). These approaches differ operationally, but all turn raw anomaly evidence into thresholdable quantities with interpretable false-alarm behavior.
Stability is a separate issue. "COGNOS: Universal Enhancement for Time Series Anomaly Detection via Constrained Gaussian-Noise Optimization and Smoothing" argues that reconstruction-based methods’ reliance on MSE leads to statistically flawed residuals and noisy, unstable anomaly scores with poor signal-to-noise ratio; it proposes Gaussian-White Noise Regularization and a Kalman Smoothing Post-processor, reporting an average F-score uplift of 2 across 3 backbone models (Shang et al., 10 Nov 2025). This suggests that any SWAS built on residuals inherits the statistical quality of those residuals.
Robustness to heterogeneous normal complexity is another recurrent concern. "Deep Generative Model using Unregularized Score for Anomaly Detection with Heterogeneous Complexity" shows that regularization terms in deep generative models can distort anomaly scores and proposes an unregularized score that is robust to inherent sample complexity (Matsubara et al., 2018). In an SWAS context, this implies that aggregation can be misleading if local scores are biased upward for complex-but-normal subsystems.
Training procedures that reshape score distributions also matter. Batch uniformization was proposed to reduce disproportionately high anomaly scores for rare-normal sounds by minimizing a density-weighted average of per-sample scores in a mini-batch (Koizumi et al., 2019). Score-guided networks, in turn, learn a scoring function that enlarges anomaly-score disparities between normal and abnormal data during training (Huang et al., 2021). These results do not define SWAS directly, but they bear on a central requirement of any global score: local scores should be comparable before they are pooled.
Finally, severity and abnormality are not identical. MAD-Bench introduces Multilevel Anomaly Detection, where the anomaly score is expected to reflect severity rather than merely binary deviation from normal, and evaluates this using AUROC, C-index, and Kendall’s Tau-b (Cao et al., 2024). A plausible implication is that a high-quality SWAS for operational systems may need to encode practical severity, not just statistical rarity.
6. Applications, limitations, and open questions
The clearest application of SWAS is online monitoring of collective dynamics. In the multi-entity qCPD setting, the method is evaluated on synthetic auto-regressive data, coupled Chen chaotic oscillators, and Unity crowd simulations, with the stated goal of detecting significant system-level changes in a scalable and privacy-preserving way (Kor et al., 22 Sep 2025). The same scalarization logic appears in industrial asset monitoring, where global scores summarize evidence across heterogeneous sensors and systems (Alnegheimish et al., 21 Apr 2025); in hardware security, where system-wide energy traces are used for anomaly detection (Dipta et al., 2022); and in collider physics, where event-level anomaly scores define model-independent signal regions (Aarrestad et al., 2021).
Several limitations are explicit. In the multi-entity quickest-change framework, mean SWAS is strongest for collective shifts, whereas isolated anomalies may only weakly perturb the aggregate mean; variance and distributional SWAS are more sensitive to subset effects and tail changes, but the framework is tuned to detect sustained system-level shifts rather than sporadic spikes (Kor et al., 22 Sep 2025). KDE-plus-Wasserstein aggregation is computationally more expensive than mean or variance. Performance depends on the quality of normal training data, and concept drift in normal behavior is an implicit challenge (Kor et al., 22 Sep 2025).
A second limitation concerns semantic interpretation. A scalar global score can be statistically well calibrated while still being poorly aligned with operational importance. MAD-Bench shows that models with strong binary anomaly detection can differ materially in severity alignment, and that area-based bias can inflate scores for visually large anomalies even when practical severity is defined differently (Cao et al., 2024). This suggests that SWAS should not automatically be equated with risk without explicit validation.
A third issue is representational mismatch. The multi-entity qCPD paper states that, to the best of the authors’ knowledge, there was no publicly available dataset explicitly designed to evaluate change point detection in complex collective and interactive systems of the required type, motivating the introduction of Unity crowd-simulation and coupled nonlinear oscillator datasets (Kor et al., 22 Sep 2025). Broader SWAS research therefore remains constrained by dataset design: the meaning of “system-wide” depends on whether the data expose interacting entities, correlated sensors, or merely pooled features.
The term SWAS is used explicitly in multi-entity quickest change point detection, but the underlying idea spans a wider technical landscape. Across these settings, the central principle is stable: a system-wide anomaly score is a scalar reduction of distributed anomaly evidence, useful only insofar as the local evidence is well modeled, the aggregation is aligned with the system’s structure, and the resulting score is properly calibrated for the operational decision it is meant to support.