---
title: Equation-Grounded Anomaly Taxonomy
url: https://www.emergentmind.com/topics/equation-grounded-anomaly-taxonomy
type: topic
---

# Equation-Grounded Anomaly Taxonomy

Equation-grounded anomaly taxonomy denotes the classification of anomalies by explicit equations, identities, or rule systems rather than by informal rarity alone. In quantum field theory and string theory, such taxonomies separate gauge, gravitational, mixed, local, global, and holomorphic anomalies through anomaly polynomials, descent equations, bordism exact sequences, and intersection-theoretic identities [0802.0634] [1111.2351] [2011.10102] [1807.05503]. In data analysis, the same expression refers to anomaly schemes whose categories are fixed by physical, kinematic, reconstruction, or symbolic relations, as in maritime AIS benchmarks, SDSS spectral outlier analysis, and symbolic-regression-based one-class detection [2606.29721] [2510.05235] [2603.17575].

## 1. Scope of the term and contrast with data-centric typologies

The term *anomaly* is not univocal across research areas. In data analysis, an anomaly is defined as a case or group of cases that is unusual and does not fit the general patterns exhibited by the majority of the data [2007.15634]. In quantum field theory, anomalies are failures of classical symmetries at the quantum level, appearing as non-invariance of the measure, non-conservation of currents, non-invariance of the effective action, and failure of Ward or Slavnov-Taylor identities [0802.0634]. Equation-grounded taxonomy therefore names a family of practices rather than a single universal formalism.

A useful contrast is provided by the data-centric review of deviations in data. That review deliberately does **not** ground anomaly types in formulas, because formulas often represent detection techniques rather than anomaly essence. Instead it organizes anomalies by five dimensions—data type, cardinality of relationship, anomaly level, data structure, and data distribution—yielding **3 broad groups**, **9 basic types**, and **63 subtypes** [2007.15634]. Its three broad groups are **atomic univariate anomalies**, **atomic multivariate anomalies**, and **aggregate anomalies**. This framework is hierarchical and domain-independent, but it is explicitly conceptual rather than equation-driven.

Equation-grounded taxonomies depart from that stance by making the defining relation itself explicit. In the maritime case, anomalies are framed not as “rare points” or “whatever experts happened to label,” but as domain-grounded behaviors that can be written down as explicit rules and equations [2606.29721]. In SDSS spectroscopy, anomaly scores are decomposed into wavelength-localized contributions and then organized by explanation vectors rather than by a scalar outlier score alone [2510.05235]. In symbolic regression, normality is encoded by learned invariants \(f(\mathbf{x}) \approx 1\), so anomaly classes can be described in terms of violated equations rather than opaque latent coordinates [2603.17575].

## 2. Gauge, gravitational, and geometric anomaly classes

In gauge theory, anomalies are organized by symmetry source and by mathematical diagnosis. The standard taxonomy distinguishes **global anomalies**, **gauge anomalies**, **mixed anomalies**, and **gravitational anomalies**. Their consistent form is captured by
\[
\delta_\alpha \Gamma[A] = \int d^dx\, \alpha(x)\,\mathcal A(x),
\]
while the relation to current non-conservation is
\[
(D_\mu \langle J^\mu(x)\rangle)_A = -\mathcal A(x).
\]
The same literature derives anomalies from measure Jacobians, triangle or higher-point diagrams, descent equations, BRST cohomology at ghost number one, and index theory in two more dimensions [0802.0634]. This establishes an equation-grounded classification in which anomaly type is tied to the symmetry being violated and to the cohomological object that detects the violation.

A particularly explicit realization appears in six-dimensional \(\mathcal N=(1,0)\) supergravity. There the anomaly polynomial must factorize as
\[
I_8=-\frac{1}{32}\,\Omega_{\alpha\beta}X_4^\alpha X_4^\beta,
\]
with anomaly coefficients \(a\), \(b_\kappa\), and \(b_{ij}\) associated respectively with gravity, non-abelian gauge factors, and abelian gauge factors. The cancellation conditions split into three classes. Pure gravitational anomalies give
\[
273=H-V+29T,\qquad a\cdot a=9-T.
\]
Mixed gauge-gravitational anomalies and pure gauge anomalies supply the remaining relations. The paper’s central result is that these field-theoretic equations are summarized by three intersection-theoretic identities on a smooth resolved elliptically fibered Calabi–Yau threefold \(\hat X\): a **quartic identity**, a **quadratic identity**, and a **gravitational identity** [1111.2351]. In that dictionary, \(a\) is the canonical class \(K\) of the base, \(b_\kappa\) is the divisor class of the discriminant component supporting \(G_\kappa\), and the abelian coefficients satisfy
\[
b_{ij}=-\pi(S_i\cdot S_j),
\]
with \(S_i\) defined by the Shioda map.

This geometric recasting changes the meaning of taxonomy. The anomaly classes are no longer only gravitational, mixed, and gauge in a low-energy effective theory; they become specific statements in intersection theory. Charged matter is likewise reinterpreted geometrically: isolated rational curves \(c_r\) and fibered rational curves \(\chi_\rho\) are precisely the curves that shrink in the F-theory limit and give the charged BPS states [1111.2351]. A plausible implication is that equation-grounded taxonomy here functions simultaneously as a classification of allowed spectra and as a dictionary between field theory and geometry.

A different but related use of anomaly equations appears in the equivariant holomorphic anomaly formalism. For local \(\mathbb P^2\), the genus-\(g\) stable quotient series satisfies
\[
\frac{1}{2}\,\frac{\partial F^{SQ}_g}{\partial A_2}
=
\frac{1}{2}\sum_{i=1}^{g-1} \frac{\partial F^{SQ}_i}{\partial T}\, \frac{\partial F^{SQ}_{g-i}}{\partial T}
+
\frac{1}{2}\,\frac{\partial^2 F^{SQ}_{g-1}}{\partial T^2},
\qquad g\ge 2,
\]
and local \(\mathbb P^3\) has an analogous equation involving a differential operator in \(A_2\), \(B_2\), and \(B_4\) [1807.05503]. Here the anomaly is the controlled failure of holomorphicity, recursively generated by lower-genus data.

## 3. Discrete anomaly classes and the algebra of local–global interplay

For non-Abelian finite symmetries, anomaly constraints can be reduced to two universal classes. The paper on non-Abelian discrete anomaly freedom states that the relevant groups are “generically subject to one of two classes of constraints” distinguished by the field equations
\[
D_{(1)}:\quad \sum_{\boldsymbol d} K^{(G,\mathcal{G})}(\phi_{\boldsymbol d}) \overset{!}{=} 2n,
\]
and
\[
D_{(2)}:\quad \sum_{\boldsymbol d_+} K^{(G,\mathcal{G})}(\phi_{\boldsymbol d_+})
\overset{!}{=}
3n + \sum_{\boldsymbol d_-} K^{(G,\mathcal{G})}(\phi_{\boldsymbol d_-}).
\]
The first class is an **evenness condition**; the second is a **multiple-of-three / mod-3 condition** with irreducible representations separated into positive- and negative-basis-logarithm sectors [1804.04237]. The same work gives additive conditions for mixed \(D\)-\(G\)-\(G\) and \(D\)-\(\mathcal G\)-\(\mathcal G\) anomalies and a multiplicative Jacobian condition
\[
\underset{f}{\prod} \det\!\left[ U_{\boldsymbol{d}^{(f)}(g)}\right]^{2\, l(\boldsymbol r^{(f)})} \overset{!}{=} 1.
\]

This produces a compact taxonomy: **Class \(D_{(1)}\)** models are governed by evenness constraints, and **Class \(D_{(2)}\)** models by multiple-of-three constraints. The paper also states an important asymmetry: “Any model subject only to \(D_{(1)}\) is free of gravitational anomalies” [1804.04237]. In this setting, anomaly taxonomy is directly operationalized as field counting. The same logic yields “pocket formulae” for \(SU(5)\) GUTs, Standard Model-like gauge groups, and multi-factor gauge sectors.

Bordism-based anomaly theory refines this picture by classifying how local and global anomalies are related. The central exact sequence is
\[
0 \to {\Ext}^1(\pi_{d+1}(MTH),\mathbb Z)
\longrightarrow
H_{IZ}^{d+2}(MTH)
\longrightarrow
{\Hom}(\pi_{d+2}(MTH),\mathbb Z)
\to 0.
\]
The left term detects **global anomalies**; the right term detects **local anomalies**. The sequence splits, but **not canonically** [2011.10102]. That non-canonical splitting is the source of anomaly interplay.

The consequence is precise. Mixed or interplay anomalies are **not** a separate cohomology group; they are a behavior of the pullback map
\[
\pi^\ast : H_{IZ}^{d+2}(MTH') \to H_{IZ}^{d+2}(MTH).
\]
Under pullback, **local anomalies** can become local and/or global anomalies, whereas **global anomalies** can pull back only to global anomalies [2011.10102]. The paper illustrates this with examples in 2, 4, and 6 dimensions, including the derivation of the Witten \(SU(2)\) global anomaly from a local anomaly in a \(U(2)\) or \(SU(3)\) theory. A common misconception is therefore corrected: mixed/interplay anomalies are not an additional anomaly species alongside local and global ones, but an effect of how the total anomaly class decomposes under symmetry reduction.

## 4. Data anomalies beyond rarity: from conceptual typology to explicit rules

In data analysis, the comprehensive review of deviations in data offers the first theoretically principled and domain-independent typology, but it does so without equations [2007.15634]. Its nine basic types range from **uncommon number anomaly** and **uncommon class anomaly** to **multidimensional numerical anomaly**, **aggregate numerical anomaly**, and **aggregate mixed data anomaly**. The review also clarifies locality in three ways: by subspace/cardinality, by density relative to neighbors, and by explicit dependence in time, space, graphs, or relational structures. Type I and Type II anomalies are global in their basic univariate reading, whereas Types IV, VI, VII, VIII, and IX often support local definitions [2007.15634].

Equation-grounded taxonomies in applied domains are built against that background. The maritime paper explicitly rejects two common alternatives: **rarity-based anomalies**, which may miss operationally meaningful events, and **expert-labeled anomalies**, which are expensive, subjective, and difficult to scale [2606.29721]. It therefore defines anomalies as domain-grounded behaviors implementable under a limited AIS observation schema. The anomaly types are separated by the number of vessels involved: **single-vessel anomalies** \(A1\) and \(A2\), and the **inter-vessel anomaly** \(A3\) [2606.29721].

That move changes the taxonomy from a distributional description to a mechanistic one. \(A1\) is **unexpected AIS activity**, where reported position jumps or spikes while motion variables, especially SOG and COG, do not change commensurately. \(A2\) is **route deviation**, where abnormal variation in SOG and COG across a consecutive window yields a physically plausible but abnormal trajectory. \(A3\) is **close approach**, an inter-vessel near-miss defined through CPA-style geometry, especially DCPA and TCPA [2606.29721]. This suggests a methodological shift from “anomaly as scarcity” to “anomaly as formally specified behavior.”

## 5. Spectroscopic and maritime implementations

In SDSS galaxy spectroscopy, the anomaly taxonomy is generated by a multi-stage equation-grounded pipeline. A VAE reconstructs rest-frame spectra on a common optical grid and computes **eight reconstruction-based anomaly scores**: plain, filtered, trimmed, and filtered-plus-trimmed versions of MSE and inverse-flux-weighted \(\chi^2\). The standard MSE is
\[
\mathrm{MSE}(X,X')=\frac{1}{N}\sum_i^N (x_i-x_i')^2,
\]
while the inverse-flux-weighted form is
\[
\chi^2(X,X')=\frac{1}{N}\sum_i^N \frac{(x_i-x_i')^2}{x_i+\delta}.
\]
The interpretation layer, **LIME-Spectra-Interpreter**, replaces image superpixels with spectral segments and fits the local surrogate
\[
g(Z')=\sum_{i=1}^{N} w_i z'_i.
\]
Its default configuration uses **flux scaling perturbation**, **scale factor 0.9**, **uniform segmentation**, and **5000 perturbed samples** [2510.05235].

The population-level taxonomy is then obtained by taking the **top 1% most anomalous spectra** under standard MSE, converting each explanation vector to its **absolute value**, **normalizing each explanation vector to unit length**, and clustering with **KMeans**. The selected solution has **seven clusters**, grouped into three interpretive categories: **artifact-driven outliers**, **hybrid physical + processing-artifact cases**, and **physically rich emission-line populations** [2510.05235]. Clusters 0 and 5 correspond to data-reduction defects and cosmic-ray-like spikes; clusters 4 and 6 are hybrid cases dominated by clipped [OIII] caused by preprocessing around [OI] 557.7 nm; clusters 1, 2, and 3 are astrophysical populations characterized respectively as **dusty, metal-rich starbursts / moderate-excitation enriched H II regions**, **chemically enriched H II regions with moderate excitation**, and **extreme emission-line galaxies / low-metallicity systems with hard ionizing fields** [2510.05235]. The paper states that these three physical clusters comprise **69%** of the top 1% anomalies.

The significance of this procedure is that the taxonomy is grounded in standard line diagnostics rather than in unsupervised labels alone. The explanation weights peak at [OIII], H\(\beta\), H\(\alpha\), [OII], [NII], and [SII], and the cluster interpretation is supported by the Balmer decrement, [NII]/H\(\alpha\), [OIII]/H\(\beta\), [OIII]/[OII], and O3N2 [2510.05235]. The resulting categories therefore separate instrumental artifacts from physically meaningful outliers in explanation space.

In maritime AIS analysis, the equation-grounded pipeline is explicit at every stage. A trajectory is written as
\[
\mathcal X=\{x_0,\ldots,x_{N-1}\},
\qquad
x^{(i)}_t = (\mathbf p^{(i)}_t,\mathbf v^{(i)}_t),
\]
with \(\mathbf p^{(i)}_t=(\mathrm{LAT}^{(i)}_t,\mathrm{LON}^{(i)}_t)\) and \(\mathbf v^{(i)}_t=(\mathrm{COG}^{(i)}_t,\mathrm{SOG}^{(i)}_t)\). The unified score–synthesize–label pipeline begins with an LLM plausibility score vector
\[
\mathbf a^{(i)} = f_{\mathrm{LLM}}(\mathcal R^{(i)},\mathbf s^{(i)},\mathcal P)\in[0,1]^T,
\]
used only to choose where to inject anomalies and to modulate severity [2606.29721].

Each anomaly class is then given by explicit synthesis equations. For **A1**, the perturbed position must satisfy a PED threshold scaled by \(\theta_p=3\). For **A2**, speed and heading are recursively perturbed with \(\theta_v=2\) and \(\theta_h=2\), and the anomaly window is split into an **injection phase** of the first \(2/3\) and a **recovery phase** of the remaining \(1/3\). For **A3**, a virtual vessel is constructed by
\[
\mathbf p^{(v)}_t=\mathbf p^{(o)}_t+d_t\,\mathbf n_t,
\]
the anchor point is chosen by \(k=\arg\max_j \mathbf a_j\), and the target separation \(\widetilde D\in[0.1,0.3]\) NM is sampled from the **DANGEROUS** DCPA range; labels are assigned only to \(t\in\{k-1,k,k+1\}\) [2606.29721]. Benchmarking is **timestamp-level**, with **no point adjustment, no event tolerance, and no window relaxation**, and reports **AUROC**, **AUPRC**, and **best F1** for \(T\in\{12,24\}\) and for each anomaly type separately [2606.29721].

## 6. Symbolic invariants, interpretability, and limitations

SYRAN generalizes the equation-grounded idea by learning the equations themselves. In the one-class setting, normality is modeled by **symbolic invariants** \(f:\mathbb R^d\to\mathbb R\) satisfying
\[
f(\mathbf x)\approx 1 \quad \text{for all } \mathbf x\in X_{\text{train}}.
\]
The optimization target is
\[
L(f) = L_1(f) + \max\bigl(0,\, \Delta - L_{\text{noise}}(f)\bigr) + \gamma\,L_c(f),
\]
where \(L_1\) is the average absolute deviation from 1 on training data, \(L_{\text{noise}}\) penalizes trivial constants using random points sampled featurewise from empirical ranges, and \(L_c\) penalizes expression complexity [2603.17575]. At test time, each invariant contributes a calibrated residual, and the final anomaly score is the ensemble average of sigmoid-transformed normalized deviations.

The paper explicitly notes that it does **not** define a formal anomaly taxonomy, but that its architecture enables one based on the violated invariants [2603.17575]. The enabled categories include point-wise residual anomalies, structural law violations, multivariate relationship breakdowns, feature-subspace anomalies, single-rule versus multi-rule anomalies, and severity-graded anomalies. Because each score is attached to a readable equation, interpretability is intrinsic rather than post hoc. The reported examples—Kepler’s third law, the breast cancer Wisconsin invariant based on cell size uniformity, the vertebral-column relation involving degree of spondylolisthesis and lumbar lordosis angle, and the wine invariant based on alcohol content—show that anomaly definition can be grounded in explicit scientific or medical relations [2603.17575].

Across domains, however, equation-grounded taxonomies remain bounded by their construction. The F-theory analysis assumes singularities are “mild,” excludes cases requiring the blow-up of a point into a four-cycle with tensionless strings, restricts the abelian discussion to massless \(U(1)\) factors, and ignores discrete quotients in the gauge group because the analysis is local anomaly cancellation [1111.2351]. The maritime benchmark covers only a subset of maritime anomalies, treats the LLM as a constrained scorer rather than a truth oracle, focuses mainly on cargo and tanker tracks, and simplifies \(A3\) to a short three-timestamp near-miss event [2606.29721]. The SDSS framework shows that explanation can separate physically meaningful spectra from artifacts, but it also identifies hybrid cases in which preprocessing masks distort real [OIII] emission [2510.05235].

A recurring misconception is that an equation-grounded taxonomy is simply a more technical form of outlier scoring. The reviewed literature points elsewhere. The data-centric review separates anomaly **type** from detection formula [2007.15634]. The bordism analysis shows that anomaly interplay is not a separate class but a property of pullback under a non-canonically split exact sequence [2011.10102]. The SDSS study requires explanation vectors and emission-line diagnostics, not only reconstruction scores, to obtain physically coherent categories [2510.05235]. The maritime benchmark argues explicitly that rare does not necessarily mean dangerous [2606.29721]. The common thread is therefore not the presence of equations as such, but the use of equations to specify what structure is being violated and why that violation constitutes an anomaly.

Source: https://www.emergentmind.com/topics/equation-grounded-anomaly-taxonomy