Papers
Topics
Authors
Recent
Search
2000 character limit reached

Area-Based Epigraph Index (ABEI)

Updated 6 July 2026
  • Area-Based Epigraph Index (ABEI) is an integrated measure that quantifies the cumulative vertical excess of sample curves above a reference, capturing both magnitude and duration of deviations.
  • It replaces proportion-based logic with an area integration approach, thereby enhancing sensitivity to subtle shape and magnitude outliers in functional data.
  • ABEI is a core component of the EHyOut methodology, producing a six-dimensional feature vector for robust multivariate outlier detection through comparisons on curves and their derivatives.

The Area-Based Epigraph Index (ABEI) is an area-based extremality functional that quantifies, for a given curve, the accumulated positive vertical distance by which other sample curves lie above it over a compact domain IR\mathcal{I}\subset\mathbb{R}. Introduced together with the Area-Based Hypograph Index (ABHI) for functional outlier detection, ABEI replaces the proportion-of-domain logic of earlier epigraph/hypograph indices with integrated distance, thereby making the resulting representation sensitive to both magnitude and shape deviations (Pulido et al., 8 Jul 2025). In the associated EHyOut methodology, ABEI and ABHI are computed for each curve and for its first and second derivatives, producing a six-dimensional feature representation to which robust multivariate outlier detection is applied (Pulido et al., 8 Jul 2025).

1. Conceptual setting and relation to earlier epigraph/hypograph indices

The underlying setting is standard functional data analysis on a compact interval I\mathcal{I} with Lebesgue measure λ()\lambda(\cdot), where curves are real-valued functions xC(I,R)x\in C(\mathcal{I},\mathbb{R}) and a random curve XX is a stochastic process X:IRX:\mathcal{I}\to\mathbb{R} with distribution PXP_X (Pulido et al., 8 Jul 2025). For a function xx, the epigraph and hypograph are

Epi(x)={(t,y)I×R:yx(t)},Hypo(x)={(t,y)I×R:yx(t)}.\mathrm{Epi}(x)=\{(t,y)\in \mathcal{I}\times\mathbb{R}: y\ge x(t)\}, \qquad \mathrm{Hypo}(x)=\{(t,y)\in \mathcal{I}\times\mathbb{R}: y\le x(t)\}.

These sets induce top-to-bottom orderings by asking whether competing curves lie above or below a reference curve.

The population epigraph and hypograph indices average pointwise above/below probabilities over the domain:

EI(x)=11λ(I)IP ⁣(X(t)x(t))dt,HI(x)=1λ(I)IP ⁣(X(t)x(t))dt.\mathrm{EI}(x)=1-\frac{1}{\lambda(\mathcal{I})}\int_{\mathcal{I}}\mathbb{P}\!\big(X(t)\ge x(t)\big)\,dt, \qquad \mathrm{HI}(x)=\frac{1}{\lambda(\mathcal{I})}\int_{\mathcal{I}}\mathbb{P}\!\big(X(t)\le x(t)\big)\,dt.

Their sample analogues, the Modified Epigraph Index and Modified Hypograph Index, are

I\mathcal{I}0

I\mathcal{I}1

These quantities lie in I\mathcal{I}2, rank curves by how often they are above or below the sample, and satisfy a sign-flip relation,

I\mathcal{I}3

A central motivation for ABEI is the limited sensitivity of MEI and MHI to magnitude outliers. If a competitor is a constant vertical shift of the reference, I\mathcal{I}4 with I\mathcal{I}5, then the set I\mathcal{I}6 has full measure regardless of whether I\mathcal{I}7 or I\mathcal{I}8. MEI and MHI therefore record the prevalence of being above or below, but not the size of the displacement. This saturation is the specific limitation that the area-based indices were designed to overcome (Pulido et al., 8 Jul 2025).

2. Formal definition of ABEI and ABHI

For a sample I\mathcal{I}9 and a curve λ()\lambda(\cdot)0, the Area-Based Epigraph Index and Area-Based Hypograph Index are defined as

λ()\lambda(\cdot)1

where λ()\lambda(\cdot)2 (Pulido et al., 8 Jul 2025). The construction is pairwise and sample-aggregated: each comparison is made against every curve in the sample, and no single central template is introduced.

The positive-part integrand imposes a directional decomposition. For ABEI, only those portions of the domain where a competing curve lies above the reference contribute to the integral; for ABHI, only the complementary below contributions are counted. The indices are not normalized by λ()\lambda(\cdot)3, λ()\lambda(\cdot)4, or amplitude range. Consequently, they are nonnegative and unbounded above, with units of amplitude times domain units. If the domain is time, the units are value-times-time (Pulido et al., 8 Jul 2025).

A simple example illustrates the distinction from MEI. On λ()\lambda(\cdot)5, let λ()\lambda(\cdot)6 and λ()\lambda(\cdot)7, and evaluate the indices for λ()\lambda(\cdot)8. Then

λ()\lambda(\cdot)9

while

xC(I,R)x\in C(\mathcal{I},\mathbb{R})0

MEI would only record that xC(I,R)x\in C(\mathcal{I},\mathbb{R})1 is above xC(I,R)x\in C(\mathcal{I},\mathbb{R})2 everywhere; ABEI records the integrated offset itself (Pulido et al., 8 Jul 2025).

3. Mathematical properties and statistical interpretation

ABEI and ABHI admit a number of basic structural properties. First, they are complementary in the sense that

xC(I,R)x\in C(\mathcal{I},\mathbb{R})3

Thus their sum is the total aggregated xC(I,R)x\in C(\mathcal{I},\mathbb{R})4 discrepancy between the target curve and the sample, while their separate values preserve directional information about whether discrepancies are predominantly above or below (Pulido et al., 8 Jul 2025).

Second, the indices satisfy a sign-flip identity,

xC(I,R)x\in C(\mathcal{I},\mathbb{R})5

Unlike MEI and MHI, ABEI and ABHI do not exhibit the sample-level linear dependence that constrains joint use of the modified indices. This makes the pair naturally usable as a bivariate feature representation (Pulido et al., 8 Jul 2025).

Third, the indices have clear invariance and monotonicity behavior. If all curves, including the reference, are shifted by the same constant xC(I,R)x\in C(\mathcal{I},\mathbb{R})6, ABEI and ABHI are unchanged because only pairwise differences matter. Under common positive scaling by xC(I,R)x\in C(\mathcal{I},\mathbb{R})7, both indices scale by xC(I,R)x\in C(\mathcal{I},\mathbb{R})8 and are therefore not scale-invariant. Under nonlinear reparameterization xC(I,R)x\in C(\mathcal{I},\mathbb{R})9, the area element XX0 changes, so the indices are generally not invariant to time warping. They are also monotone with respect to the pointwise order: if XX1 for all XX2, then XX3 and XX4 (Pulido et al., 8 Jul 2025).

These properties explain the dual sensitivity that motivated the method. Magnitude outliers generate large vertical offsets and therefore large integrated areas. Shape outliers, including localized bumps, spikes, or phase-induced discrepancies, contribute only on the subregions where the sign-constrained difference is positive, so their influence is spatially localized but still measurable. A plausible implication is that ABEI and ABHI interpolate between rank-based ordering and geometric discrepancy: they retain the directional logic of epigraph/hypograph comparisons while moving from set measure to signed area accumulation.

4. The EHyOut methodology

EHyOut reformulates functional outlier detection as a low-dimensional robust multivariate detection problem. The workflow begins with functional representation: each observed curve XX5 on a grid XX6 is represented through cubic spline interpolation, producing a twice-differentiable spline XX7 together with its first and second derivatives XX8 and XX9 (Pulido et al., 8 Jul 2025).

For each curve, ABEI and ABHI are then computed on the original function and on the first two derivatives:

X:IRX:\mathcal{I}\to\mathbb{R}0

X:IRX:\mathcal{I}\to\mathbb{R}1

X:IRX:\mathcal{I}\to\mathbb{R}2

The resulting feature vector is

X:IRX:\mathcal{I}\to\mathbb{R}3

Numerical integration is approximated by quadrature on the observation grid, and a trapezoidal rule is explicitly described:

X:IRX:\mathcal{I}\to\mathbb{R}4

Outlier detection on the six-dimensional features is performed with the Comedian method. Robust location is the componentwise median,

X:IRX:\mathcal{I}\to\mathbb{R}5

and robust marginal scale is estimated via MAD,

X:IRX:\mathcal{I}\to\mathbb{R}6

Pairwise robust scatter is based on comedian covariance, and robust Mahalanobis distances are computed as

X:IRX:\mathcal{I}\to\mathbb{R}7

A curve is flagged as an outlier when

X:IRX:\mathcal{I}\to\mathbb{R}8

The computational bottleneck is the feature construction rather than the robust multivariate step. Computing ABEI and ABHI for all curves on a grid of size X:IRX:\mathcal{I}\to\mathbb{R}9 requires PXP_X0 operations, and inclusion of first and second derivatives multiplies runtime by approximately three (Pulido et al., 8 Jul 2025). The paper reports that vectorization and reuse of precomputed differences make this practical for sample sizes up to several hundreds. Implementations are available in the R package ehymet and in the authors’ GitHub repository ehyout (Pulido et al., 8 Jul 2025).

5. Empirical performance in simulation and case studies

The empirical study benchmarked EHyOut against FASTMUOD, OG (Outliergram), MSPLOT, TVD, MBD, MDS5LOF, and BP-PWD on 19 heterogeneous data-generating processes with PXP_X1 and contamination levels PXP_X2, covering magnitude, shape, and mixed outliers (Pulido et al., 8 Jul 2025). Evaluation used Matthews Correlation Coefficient (MCC), Area Under the ROC Curve (AUC) when a single score was available, and Execution Time (ET, in seconds).

At PXP_X3, the aggregated summary across all DGPs was as follows.

Method Mean ET (s) Mean MCC
EHyOut 0.008 0.806
TVD 0.012 0.668
MSPLOT 0.054 0.633
OG 2.578 0.594
BP-PWD 0.027 0.535
FASTMUOD 0.041 0.516
MDS5LOF 0.047 0.365
MBD 0.005 0.304

The quartile summary reported for EHyOut was MCC PXP_X4 and PXP_X5; for TVD, PXP_X6 and PXP_X7, indicating higher variability despite being second best in mean MCC (Pulido et al., 8 Jul 2025). EHyOut was the second fastest method in mean execution time and the only method to achieve median MCC greater than PXP_X8 on all 19 scenarios. The paper also reports near-1 AUC in many settings where AUC was defined.

Two real-data applications illustrate the method’s interpretive use. In Spanish weather data from AEMET, involving 73 stations over 1980–2009, EHyOut flagged 11 temperature outliers and 9 precipitation outliers, with 7 stations identified in both variables: Fuerteventura, Lanzarote, Las Palmas (Gando), Hierro, Tenerife Sur, Sta. Cruz de Tenerife, and Izaña. Navacerrada appeared as a temperature outlier, while Logroño, Valencia, Madrid (Torrejón), and Colmenar Viejo were among the precipitation outliers (Pulido et al., 8 Jul 2025). The reported interpretation is meteorological: the Canary Islands exhibit distinct climatology, and the flagged peninsular stations show extreme precipitation behavior.

In the United Nations world population data for 1950–2010, filtered to 105 countries with population in PXP_X9 in 1980, EHyOut flagged 38 countries as outliers. The examples given include rapid-growth trajectories such as Saudi Arabia, Iraq, Afghanistan, Malaysia, Uganda, and Sudan; moderate or stable Eastern European countries such as Hungary and the Czech Republic; and consistently high trends such as Chile, Australia, Cuba, the Netherlands, Greece, and Portugal (Pulido et al., 8 Jul 2025). The classification is described as consistent with prior analyses while highlighting both high-growth and consistently high trajectories.

6. Scope, limitations, and relation to adjacent literature

The method’s practical defaults are explicit. The empirical study selected ABEI and ABHI on the original curves and on the first and second derivatives as the feature set, and selected the Comedian method with a boxplot-type cutoff because it had superior mean AUC and small dispersion relative to FASTMCD, OGK, and RMDsh (Pulido et al., 8 Jul 2025). For smooth functional inputs, cubic spline interpolation yields a xx0 representation; in noise-heavy settings, smoothing splines with smoothing parameter chosen by cross-validation or generalized cross-validation are recommended. Trapezoidal integration on a sufficiently dense grid, with xx1–xx2, is reported as adequate.

The limitations are equally specific. The feature construction itself is not robust to extreme contamination, since severe outliers increase pairwise areas. Robustness enters through the subsequent median/MAD/comedian stage, which has a high breakdown point in the sense described in the paper. Derivatives can amplify noise, strong phase variability may warrant alignment if phase is considered nuisance, and the xx3 pairwise cost recommends vectorization and parallelization for very large xx4 (Pulido et al., 8 Jul 2025).

A notable issue in the literature is terminology. In work on clustering multivariate functional data, the generalized modified epigraph and hypograph indices based on normalized Lebesgue measure over time have been described as “area-based” because they integrate the proportion of the domain over which all components satisfy the ordering relation (Pulido et al., 2023). In that setting, the objects denoted MEI and MHI are scale-free, invariant to one-to-one transformations xx5, and preserve interdependencies through joint componentwise orderings. By contrast, the 2025 ABEI and ABHI are unnormalized integrated positive distances between curves, measured in physical area units and generally not invariant to time warping (Pulido et al., 8 Jul 2025, Pulido et al., 2023). This suggests that “area-based” has been used in two distinct senses in the epigraph/hypograph literature: proportion-of-domain area in generalized MEI/MHI, and area-between-curves in ABEI/ABHI proper.

That distinction also clarifies a common misconception. ABEI is not merely a renaming of MEI. MEI records how often one curve is above another; ABEI records by how much and for how long, through signed area accumulation. In functional outlier detection, this difference is the mechanism by which magnitude sensitivity is restored while shape information is retained (Pulido et al., 8 Jul 2025).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Area-Based Epigraph Index (ABEI).