Area-Based Hypograph Index (ABHI)
- ABHI is an index based on the integrated area under a curve’s hypograph that quantifies how often a function lies above or below sample curves.
- It encompasses multiple formulations tailored for functional data clustering, outlier detection, and measuring joint dependence in random vectors.
- Practically, ABHI is computed via indicator functions or area-gap methods on smoothed data, with extensions to multivariate settings illustrated in weather and air quality studies.
Searching arXiv for the cited ABHI-related papers to ground the article in the current preprint record. Area-Based Hypograph Index (ABHI) denotes an area-based construction centered on the hypograph of a function or curve. In the recent functional-data literature, the term is used for indices that rank curves by integrating how often, or by how much, a reference curve lies above a sample; these formulations support clustering of multivariate functional data and functional outlier detection (Pulido et al., 2023). In a separate dependence-measure literature, the same acronym is associated with an area-under-the-Kendall-curve quantity used to measure joint dependence of a random vector (Afendras et al., 2020). This suggests that ABHI is not a single universally fixed formula, but a label attached to closely related “area-under-a-hypograph” ideas in different statistical settings.
1. Terminological scope and principal formulations
The term ABHI appears in at least three closely connected but non-identical formulations. In the multivariate functional-data clustering setting, the index is defined through indicator functions and measures the fraction of the domain on which sample curves lie at or below a reference curve, with normalization to the unit interval (Pulido et al., 2023). In the functional outlier-detection setting, the index is defined through the positive part and measures the accumulated area by which a target curve exceeds other curves (Pulido et al., 8 Jul 2025). In the AUK-based dependence setting, the relevant area is computed from a Kendall curve and an independence benchmark (Afendras et al., 2020).
| Source | Object | Defining form |
|---|---|---|
| (Pulido et al., 2023) | Functional depth / clustering | |
| (Pulido et al., 8 Jul 2025) | Functional outlier detection | |
| (Afendras et al., 2020) | Joint dependence |
A common misconception is that ABHI has a unique canonical formula. The published record instead shows that the acronym has been specialized to different inferential goals. The common element is hypograph-based area aggregation, but the operational meaning depends on whether the objective is curve ordering, outlier detection, or dependence measurement.
2. Indicator-based ABHI for functional data
For a univariate sample of continuous real-valued curves on a compact interval with Lebesgue measure , the area-based hypograph index of a reference curve is defined as
Equivalently, for each one computes the fraction of sample curves below , then integrates in 0 and normalizes by the total “time” 1 (Pulido et al., 2023).
At each time 2, dividing by 3 gives a pointwise “depth”
4
Integrating 5 over 6 measures the area of the hypograph-region “under” 7 that is occupied by the sample curves. Dividing by 8 rescales the index to 9. A curve with 0 lies above most sample curves for most 1 (very central or “high” curve), while 2 lies towards the bottom of the sample (Pulido et al., 2023).
The same source gives an illustrative example on 3 with
4
and reference
5
Using the three-point grid 6, the pointwise fractions are 7, 8, and 9. A trapezoidal or simple-average approximation yields
0
This example makes explicit that the indicator-based ABHI is an ordering device: it summarizes how often a reference curve dominates the sample in the pointwise partial order.
3. Multivariate extension, derivatives, and clustering
For multivariate functional data, each observation is a 1-variate function
2
and the reference is 3. The multivariate ABHI is
4
In words, at each 5 one requires all 6 components of 7 to lie below the corresponding coordinates of 8, then average and integrate (Pulido et al., 2023).
This formulation was introduced together with a novel formulation of the epigraph and hypograph indices, along with their generalized expressions, specifically designed for multivariate functional data. The new definitions account for interrelationships between variables, enabling effective clustering of multivariate functional data based on the original data curves and their first two derivatives. The methodology was tested on simulated datasets and further illustrated with the Canadian weather dataset and a 2023 air quality study in Madrid (Pulido et al., 2023).
The practical computation proceeds by smoothing each observed curve on 9 via splines or basis expansion if desired, optionally computing first and second derivatives, discretizing the integral on a grid 0, evaluating
1
or, in the multivariate case,
2
averaging over 3 to obtain 4, and approximating the integral by a Riemann sum (Pulido et al., 2023). The workflow is thus tightly aligned with standard FDA preprocessing, particularly spline smoothing and derivative estimation.
4. Area-gap ABHI and functional outlier detection
A later formulation extends the classical hypograph index by not only asking “on how much of the domain does a given curve lie above the others?” but also “by how much?” For curves 5 on a compact interval 6, and any target curve 7,
8
where 9 denotes the positive part (Pulido et al., 8 Jul 2025).
If one prefers an average-per-curve normalization and/or a per-unit-length normalization, one can equivalently write
0
which then takes values in 1 and interprets ABHI as the average per-curve, per-unit-time hypograph-area. The original paper works with the unnormalized version (Pulido et al., 8 Jul 2025).
This area-gap formulation is explicitly designed to capture both magnitude and shape deviations. A curve that is uniformly slightly above its peers over the entire domain still accumulates a large area, while a curve that “peaks” above the rest accumulates area in those peaks. The paper also states the decomposition
2
where 3 is the area-based epigraph index. Thus 4 measures exactly half, in a signed-area sense, of the total 5-distance between 6 and the sample (Pulido et al., 8 Jul 2025).
Several theoretical properties are given. ABHI is nonnegative, and 7 precisely if 8 for all 9 and all 0. There is no finite upper bound. ABHI under process 1 equals ABEI under process 2, that is,
3
Unlike the classical MEI/MHI, ABEI and ABHI enjoy no nontrivial linear relation, making them jointly informative. By integrating vertical deviations, ABHI is robust to small-scale fluctuations, yet sensitive to any persistent or high-amplitude deviation (Pulido et al., 8 Jul 2025).
5. Algorithms, complexity, and relation to other functional depths
For the indicator-based formulation, practical computation on a grid uses pointwise indicators, averages them across curves, and approximates the integral by a Riemann sum. Its stated complexity is 4 for univariate data or 5 for 6-variate data, where 7 is the number of grid points. The method therefore scales linearly in 8 and 9 (Pulido et al., 2023).
For the area-gap formulation, the paper gives explicit pseudocode for computing ABHI for every curve in the sample on an equally spaced grid 0 with spacing 1. For each ordered pair of curves 2 and each grid point, one computes
3
adds 4 whenever 5, and multiplies the accumulated sum by 6 to approximate the integral. The stated complexity is 7 operations, with the additional remark that one can vectorize the inner loops or use matrix operations to reduce overhead (Pulido et al., 8 Jul 2025). This suggests that the differing complexity statements across the literature reflect different computational tasks: evaluating an index for a reference curve versus computing indices for all sample curves.
The 2023 clustering paper compares ABHI with Band Depth (BD) and Random Projection Depth (RPD). BD counts the fraction of time a curve lies within random bands formed by 8 sample curves, with computational cost 9 and combinatorial selection of 0. RPD projects each curve onto random directions in the function space, applies a univariate depth on the projections, then averages, with cost 1 for 2 random projections. By contrast, ABHI is presented as directly interpretable—one sees “how often the curve is above/below” the sample in each margin, or jointly for multivariate data—and as responsive to vertical shifts and to local crossings (Pulido et al., 2023).
The 2025 outlier-detection paper sharpens that contrast by emphasizing that classical MEI and MHI rank functions by the fraction of the domain on which one curve lies above or below another, while ABEI and ABHI quantify the area between curves. A localized spike of height 3 on an interval 4 yields 5, so an anomaly occupying a small fraction of the domain can still be prominent if its area is large (Pulido et al., 8 Jul 2025).
6. EHyOut and the broader statistical role of ABHI
ABHI is a core component of EHyOut, the Area-Based Epigraph and Hypograph Outlier detection procedure. For each curve 6, EHyOut constructs the six-dimensional representation
7
where 8 and 9 are the first and second derivatives, typically estimated via spline smoothing. This feature matrix of size 0 is then passed to a robust multivariate outlier algorithm, specifically the Comedian-based robust Mahalanobis distance (Pulido et al., 8 Jul 2025).
The role of derivatives is explicit. Applying ABHI to 1 and 2 accentuates shape anomalies in slopes or curvature. A purely magnitude outlier may vanish under 3, whereas a pure shape outlier often peaks in these derivatives. The paper reports extensive simulation on 19 different data-generating processes and states that EHyOut has been shown to outperform or match specialized methods for magnitude, amplitude, and shape outliers, while remaining computationally competitive. Applications to Spanish weather data and United Nations world population data further illustrate the methodology (Pulido et al., 8 Jul 2025).
In a different statistical direction, the dependence-measure literature uses an area-under-the-Kendall-curve construction that the supplied exposition labels ABHI. Let 4 be a continuous 5-variate random vector with joint cdf 6, define 7, and let
8
With 9 denoting the product-uniform benchmark, the relevant area quantity is
00
The paper then forms a 01-vector of AUK values over all sign-rotations and defines
02
with a standardized version 03 (Afendras et al., 2020). This usage places the “hypograph” idea in a copula-based dependence framework rather than FDA.
Taken together, these strands show that ABHI serves as a bridge concept between ordering, deviation, and dependence. The precise mathematical object changes across settings, but the organizing principle remains the same: integrate a hypograph-based quantity so that geometric information accumulated over a domain becomes statistically actionable.