Papers
Topics
Authors
Recent
Search
2000 character limit reached

Area-Based Hypograph Index (ABHI)

Updated 6 July 2026
  • ABHI is an index based on the integrated area under a curve’s hypograph that quantifies how often a function lies above or below sample curves.
  • It encompasses multiple formulations tailored for functional data clustering, outlier detection, and measuring joint dependence in random vectors.
  • Practically, ABHI is computed via indicator functions or area-gap methods on smoothed data, with extensions to multivariate settings illustrated in weather and air quality studies.

Searching arXiv for the cited ABHI-related papers to ground the article in the current preprint record. Area-Based Hypograph Index (ABHI) denotes an area-based construction centered on the hypograph of a function or curve. In the recent functional-data literature, the term is used for indices that rank curves by integrating how often, or by how much, a reference curve lies above a sample; these formulations support clustering of multivariate functional data and functional outlier detection (Pulido et al., 2023). In a separate dependence-measure literature, the same acronym is associated with an area-under-the-Kendall-curve quantity used to measure joint dependence of a random vector (Afendras et al., 2020). This suggests that ABHI is not a single universally fixed formula, but a label attached to closely related “area-under-a-hypograph” ideas in different statistical settings.

1. Terminological scope and principal formulations

The term ABHI appears in at least three closely connected but non-identical formulations. In the multivariate functional-data clustering setting, the index is defined through indicator functions and measures the fraction of the domain on which sample curves lie at or below a reference curve, with normalization to the unit interval (Pulido et al., 2023). In the functional outlier-detection setting, the index is defined through the positive part (x(t)xi(t))+(x(t)-x_i(t))_+ and measures the accumulated area by which a target curve exceeds other curves (Pulido et al., 8 Jul 2025). In the AUK-based dependence setting, the relevant area is computed from a Kendall curve and an independence benchmark (Afendras et al., 2020).

Source Object Defining form
(Pulido et al., 2023) Functional depth / clustering 1nλ(T)i=1nT1{xi(t)x(t)}dt\frac{1}{n\,\lambda(T)}\sum_{i=1}^n\int_T \mathbf{1}\{x_i(t)\le x(t)\}\,dt
(Pulido et al., 8 Jul 2025) Functional outlier detection i=1nI(x(t)xi(t))+dt\sum_{i=1}^n\int_{\mathcal I}(x(t)-x_i(t))_+\,dt
(Afendras et al., 2020) Joint dependence 01Kd(t)dKΠ(t)\int_0^1 K_d(t)\,dK_\Pi(t)

A common misconception is that ABHI has a unique canonical formula. The published record instead shows that the acronym has been specialized to different inferential goals. The common element is hypograph-based area aggregation, but the operational meaning depends on whether the objective is curve ordering, outlier detection, or dependence measurement.

2. Indicator-based ABHI for functional data

For a univariate sample of continuous real-valued curves on a compact interval TRT\subset\mathbb{R} with Lebesgue measure λ(T)\lambda(T), the area-based hypograph index of a reference curve x(t)x(t) is defined as

ABHIn(x)  =  1nλ(T)i=1nT1{xi(t)x(t)}  dt.\mathrm{ABHI}_n\bigl(x\bigr) \;=\; \frac{1}{n\,\lambda(T)}\, \sum_{i=1}^n \int_{T} \mathbf{1}\bigl\{\,x_i(t)\le x(t)\bigr\} \;dt.

Equivalently, for each tt one computes the fraction of sample curves below x(t)x(t), then integrates in 1nλ(T)i=1nT1{xi(t)x(t)}dt\frac{1}{n\,\lambda(T)}\sum_{i=1}^n\int_T \mathbf{1}\{x_i(t)\le x(t)\}\,dt0 and normalizes by the total “time” 1nλ(T)i=1nT1{xi(t)x(t)}dt\frac{1}{n\,\lambda(T)}\sum_{i=1}^n\int_T \mathbf{1}\{x_i(t)\le x(t)\}\,dt1 (Pulido et al., 2023).

At each time 1nλ(T)i=1nT1{xi(t)x(t)}dt\frac{1}{n\,\lambda(T)}\sum_{i=1}^n\int_T \mathbf{1}\{x_i(t)\le x(t)\}\,dt2, dividing by 1nλ(T)i=1nT1{xi(t)x(t)}dt\frac{1}{n\,\lambda(T)}\sum_{i=1}^n\int_T \mathbf{1}\{x_i(t)\le x(t)\}\,dt3 gives a pointwise “depth”

1nλ(T)i=1nT1{xi(t)x(t)}dt\frac{1}{n\,\lambda(T)}\sum_{i=1}^n\int_T \mathbf{1}\{x_i(t)\le x(t)\}\,dt4

Integrating 1nλ(T)i=1nT1{xi(t)x(t)}dt\frac{1}{n\,\lambda(T)}\sum_{i=1}^n\int_T \mathbf{1}\{x_i(t)\le x(t)\}\,dt5 over 1nλ(T)i=1nT1{xi(t)x(t)}dt\frac{1}{n\,\lambda(T)}\sum_{i=1}^n\int_T \mathbf{1}\{x_i(t)\le x(t)\}\,dt6 measures the area of the hypograph-region “under” 1nλ(T)i=1nT1{xi(t)x(t)}dt\frac{1}{n\,\lambda(T)}\sum_{i=1}^n\int_T \mathbf{1}\{x_i(t)\le x(t)\}\,dt7 that is occupied by the sample curves. Dividing by 1nλ(T)i=1nT1{xi(t)x(t)}dt\frac{1}{n\,\lambda(T)}\sum_{i=1}^n\int_T \mathbf{1}\{x_i(t)\le x(t)\}\,dt8 rescales the index to 1nλ(T)i=1nT1{xi(t)x(t)}dt\frac{1}{n\,\lambda(T)}\sum_{i=1}^n\int_T \mathbf{1}\{x_i(t)\le x(t)\}\,dt9. A curve with i=1nI(x(t)xi(t))+dt\sum_{i=1}^n\int_{\mathcal I}(x(t)-x_i(t))_+\,dt0 lies above most sample curves for most i=1nI(x(t)xi(t))+dt\sum_{i=1}^n\int_{\mathcal I}(x(t)-x_i(t))_+\,dt1 (very central or “high” curve), while i=1nI(x(t)xi(t))+dt\sum_{i=1}^n\int_{\mathcal I}(x(t)-x_i(t))_+\,dt2 lies towards the bottom of the sample (Pulido et al., 2023).

The same source gives an illustrative example on i=1nI(x(t)xi(t))+dt\sum_{i=1}^n\int_{\mathcal I}(x(t)-x_i(t))_+\,dt3 with

i=1nI(x(t)xi(t))+dt\sum_{i=1}^n\int_{\mathcal I}(x(t)-x_i(t))_+\,dt4

and reference

i=1nI(x(t)xi(t))+dt\sum_{i=1}^n\int_{\mathcal I}(x(t)-x_i(t))_+\,dt5

Using the three-point grid i=1nI(x(t)xi(t))+dt\sum_{i=1}^n\int_{\mathcal I}(x(t)-x_i(t))_+\,dt6, the pointwise fractions are i=1nI(x(t)xi(t))+dt\sum_{i=1}^n\int_{\mathcal I}(x(t)-x_i(t))_+\,dt7, i=1nI(x(t)xi(t))+dt\sum_{i=1}^n\int_{\mathcal I}(x(t)-x_i(t))_+\,dt8, and i=1nI(x(t)xi(t))+dt\sum_{i=1}^n\int_{\mathcal I}(x(t)-x_i(t))_+\,dt9. A trapezoidal or simple-average approximation yields

01Kd(t)dKΠ(t)\int_0^1 K_d(t)\,dK_\Pi(t)0

This example makes explicit that the indicator-based ABHI is an ordering device: it summarizes how often a reference curve dominates the sample in the pointwise partial order.

3. Multivariate extension, derivatives, and clustering

For multivariate functional data, each observation is a 01Kd(t)dKΠ(t)\int_0^1 K_d(t)\,dK_\Pi(t)1-variate function

01Kd(t)dKΠ(t)\int_0^1 K_d(t)\,dK_\Pi(t)2

and the reference is 01Kd(t)dKΠ(t)\int_0^1 K_d(t)\,dK_\Pi(t)3. The multivariate ABHI is

01Kd(t)dKΠ(t)\int_0^1 K_d(t)\,dK_\Pi(t)4

In words, at each 01Kd(t)dKΠ(t)\int_0^1 K_d(t)\,dK_\Pi(t)5 one requires all 01Kd(t)dKΠ(t)\int_0^1 K_d(t)\,dK_\Pi(t)6 components of 01Kd(t)dKΠ(t)\int_0^1 K_d(t)\,dK_\Pi(t)7 to lie below the corresponding coordinates of 01Kd(t)dKΠ(t)\int_0^1 K_d(t)\,dK_\Pi(t)8, then average and integrate (Pulido et al., 2023).

This formulation was introduced together with a novel formulation of the epigraph and hypograph indices, along with their generalized expressions, specifically designed for multivariate functional data. The new definitions account for interrelationships between variables, enabling effective clustering of multivariate functional data based on the original data curves and their first two derivatives. The methodology was tested on simulated datasets and further illustrated with the Canadian weather dataset and a 2023 air quality study in Madrid (Pulido et al., 2023).

The practical computation proceeds by smoothing each observed curve on 01Kd(t)dKΠ(t)\int_0^1 K_d(t)\,dK_\Pi(t)9 via splines or basis expansion if desired, optionally computing first and second derivatives, discretizing the integral on a grid TRT\subset\mathbb{R}0, evaluating

TRT\subset\mathbb{R}1

or, in the multivariate case,

TRT\subset\mathbb{R}2

averaging over TRT\subset\mathbb{R}3 to obtain TRT\subset\mathbb{R}4, and approximating the integral by a Riemann sum (Pulido et al., 2023). The workflow is thus tightly aligned with standard FDA preprocessing, particularly spline smoothing and derivative estimation.

4. Area-gap ABHI and functional outlier detection

A later formulation extends the classical hypograph index by not only asking “on how much of the domain does a given curve lie above the others?” but also “by how much?” For curves TRT\subset\mathbb{R}5 on a compact interval TRT\subset\mathbb{R}6, and any target curve TRT\subset\mathbb{R}7,

TRT\subset\mathbb{R}8

where TRT\subset\mathbb{R}9 denotes the positive part (Pulido et al., 8 Jul 2025).

If one prefers an average-per-curve normalization and/or a per-unit-length normalization, one can equivalently write

λ(T)\lambda(T)0

which then takes values in λ(T)\lambda(T)1 and interprets ABHI as the average per-curve, per-unit-time hypograph-area. The original paper works with the unnormalized version (Pulido et al., 8 Jul 2025).

This area-gap formulation is explicitly designed to capture both magnitude and shape deviations. A curve that is uniformly slightly above its peers over the entire domain still accumulates a large area, while a curve that “peaks” above the rest accumulates area in those peaks. The paper also states the decomposition

λ(T)\lambda(T)2

where λ(T)\lambda(T)3 is the area-based epigraph index. Thus λ(T)\lambda(T)4 measures exactly half, in a signed-area sense, of the total λ(T)\lambda(T)5-distance between λ(T)\lambda(T)6 and the sample (Pulido et al., 8 Jul 2025).

Several theoretical properties are given. ABHI is nonnegative, and λ(T)\lambda(T)7 precisely if λ(T)\lambda(T)8 for all λ(T)\lambda(T)9 and all x(t)x(t)0. There is no finite upper bound. ABHI under process x(t)x(t)1 equals ABEI under process x(t)x(t)2, that is,

x(t)x(t)3

Unlike the classical MEI/MHI, ABEI and ABHI enjoy no nontrivial linear relation, making them jointly informative. By integrating vertical deviations, ABHI is robust to small-scale fluctuations, yet sensitive to any persistent or high-amplitude deviation (Pulido et al., 8 Jul 2025).

5. Algorithms, complexity, and relation to other functional depths

For the indicator-based formulation, practical computation on a grid uses pointwise indicators, averages them across curves, and approximates the integral by a Riemann sum. Its stated complexity is x(t)x(t)4 for univariate data or x(t)x(t)5 for x(t)x(t)6-variate data, where x(t)x(t)7 is the number of grid points. The method therefore scales linearly in x(t)x(t)8 and x(t)x(t)9 (Pulido et al., 2023).

For the area-gap formulation, the paper gives explicit pseudocode for computing ABHI for every curve in the sample on an equally spaced grid ABHIn(x)  =  1nλ(T)i=1nT1{xi(t)x(t)}  dt.\mathrm{ABHI}_n\bigl(x\bigr) \;=\; \frac{1}{n\,\lambda(T)}\, \sum_{i=1}^n \int_{T} \mathbf{1}\bigl\{\,x_i(t)\le x(t)\bigr\} \;dt.0 with spacing ABHIn(x)  =  1nλ(T)i=1nT1{xi(t)x(t)}  dt.\mathrm{ABHI}_n\bigl(x\bigr) \;=\; \frac{1}{n\,\lambda(T)}\, \sum_{i=1}^n \int_{T} \mathbf{1}\bigl\{\,x_i(t)\le x(t)\bigr\} \;dt.1. For each ordered pair of curves ABHIn(x)  =  1nλ(T)i=1nT1{xi(t)x(t)}  dt.\mathrm{ABHI}_n\bigl(x\bigr) \;=\; \frac{1}{n\,\lambda(T)}\, \sum_{i=1}^n \int_{T} \mathbf{1}\bigl\{\,x_i(t)\le x(t)\bigr\} \;dt.2 and each grid point, one computes

ABHIn(x)  =  1nλ(T)i=1nT1{xi(t)x(t)}  dt.\mathrm{ABHI}_n\bigl(x\bigr) \;=\; \frac{1}{n\,\lambda(T)}\, \sum_{i=1}^n \int_{T} \mathbf{1}\bigl\{\,x_i(t)\le x(t)\bigr\} \;dt.3

adds ABHIn(x)  =  1nλ(T)i=1nT1{xi(t)x(t)}  dt.\mathrm{ABHI}_n\bigl(x\bigr) \;=\; \frac{1}{n\,\lambda(T)}\, \sum_{i=1}^n \int_{T} \mathbf{1}\bigl\{\,x_i(t)\le x(t)\bigr\} \;dt.4 whenever ABHIn(x)  =  1nλ(T)i=1nT1{xi(t)x(t)}  dt.\mathrm{ABHI}_n\bigl(x\bigr) \;=\; \frac{1}{n\,\lambda(T)}\, \sum_{i=1}^n \int_{T} \mathbf{1}\bigl\{\,x_i(t)\le x(t)\bigr\} \;dt.5, and multiplies the accumulated sum by ABHIn(x)  =  1nλ(T)i=1nT1{xi(t)x(t)}  dt.\mathrm{ABHI}_n\bigl(x\bigr) \;=\; \frac{1}{n\,\lambda(T)}\, \sum_{i=1}^n \int_{T} \mathbf{1}\bigl\{\,x_i(t)\le x(t)\bigr\} \;dt.6 to approximate the integral. The stated complexity is ABHIn(x)  =  1nλ(T)i=1nT1{xi(t)x(t)}  dt.\mathrm{ABHI}_n\bigl(x\bigr) \;=\; \frac{1}{n\,\lambda(T)}\, \sum_{i=1}^n \int_{T} \mathbf{1}\bigl\{\,x_i(t)\le x(t)\bigr\} \;dt.7 operations, with the additional remark that one can vectorize the inner loops or use matrix operations to reduce overhead (Pulido et al., 8 Jul 2025). This suggests that the differing complexity statements across the literature reflect different computational tasks: evaluating an index for a reference curve versus computing indices for all sample curves.

The 2023 clustering paper compares ABHI with Band Depth (BD) and Random Projection Depth (RPD). BD counts the fraction of time a curve lies within random bands formed by ABHIn(x)  =  1nλ(T)i=1nT1{xi(t)x(t)}  dt.\mathrm{ABHI}_n\bigl(x\bigr) \;=\; \frac{1}{n\,\lambda(T)}\, \sum_{i=1}^n \int_{T} \mathbf{1}\bigl\{\,x_i(t)\le x(t)\bigr\} \;dt.8 sample curves, with computational cost ABHIn(x)  =  1nλ(T)i=1nT1{xi(t)x(t)}  dt.\mathrm{ABHI}_n\bigl(x\bigr) \;=\; \frac{1}{n\,\lambda(T)}\, \sum_{i=1}^n \int_{T} \mathbf{1}\bigl\{\,x_i(t)\le x(t)\bigr\} \;dt.9 and combinatorial selection of tt0. RPD projects each curve onto random directions in the function space, applies a univariate depth on the projections, then averages, with cost tt1 for tt2 random projections. By contrast, ABHI is presented as directly interpretable—one sees “how often the curve is above/below” the sample in each margin, or jointly for multivariate data—and as responsive to vertical shifts and to local crossings (Pulido et al., 2023).

The 2025 outlier-detection paper sharpens that contrast by emphasizing that classical MEI and MHI rank functions by the fraction of the domain on which one curve lies above or below another, while ABEI and ABHI quantify the area between curves. A localized spike of height tt3 on an interval tt4 yields tt5, so an anomaly occupying a small fraction of the domain can still be prominent if its area is large (Pulido et al., 8 Jul 2025).

6. EHyOut and the broader statistical role of ABHI

ABHI is a core component of EHyOut, the Area-Based Epigraph and Hypograph Outlier detection procedure. For each curve tt6, EHyOut constructs the six-dimensional representation

tt7

where tt8 and tt9 are the first and second derivatives, typically estimated via spline smoothing. This feature matrix of size x(t)x(t)0 is then passed to a robust multivariate outlier algorithm, specifically the Comedian-based robust Mahalanobis distance (Pulido et al., 8 Jul 2025).

The role of derivatives is explicit. Applying ABHI to x(t)x(t)1 and x(t)x(t)2 accentuates shape anomalies in slopes or curvature. A purely magnitude outlier may vanish under x(t)x(t)3, whereas a pure shape outlier often peaks in these derivatives. The paper reports extensive simulation on 19 different data-generating processes and states that EHyOut has been shown to outperform or match specialized methods for magnitude, amplitude, and shape outliers, while remaining computationally competitive. Applications to Spanish weather data and United Nations world population data further illustrate the methodology (Pulido et al., 8 Jul 2025).

In a different statistical direction, the dependence-measure literature uses an area-under-the-Kendall-curve construction that the supplied exposition labels ABHI. Let x(t)x(t)4 be a continuous x(t)x(t)5-variate random vector with joint cdf x(t)x(t)6, define x(t)x(t)7, and let

x(t)x(t)8

With x(t)x(t)9 denoting the product-uniform benchmark, the relevant area quantity is

1nλ(T)i=1nT1{xi(t)x(t)}dt\frac{1}{n\,\lambda(T)}\sum_{i=1}^n\int_T \mathbf{1}\{x_i(t)\le x(t)\}\,dt00

The paper then forms a 1nλ(T)i=1nT1{xi(t)x(t)}dt\frac{1}{n\,\lambda(T)}\sum_{i=1}^n\int_T \mathbf{1}\{x_i(t)\le x(t)\}\,dt01-vector of AUK values over all sign-rotations and defines

1nλ(T)i=1nT1{xi(t)x(t)}dt\frac{1}{n\,\lambda(T)}\sum_{i=1}^n\int_T \mathbf{1}\{x_i(t)\le x(t)\}\,dt02

with a standardized version 1nλ(T)i=1nT1{xi(t)x(t)}dt\frac{1}{n\,\lambda(T)}\sum_{i=1}^n\int_T \mathbf{1}\{x_i(t)\le x(t)\}\,dt03 (Afendras et al., 2020). This usage places the “hypograph” idea in a copula-based dependence framework rather than FDA.

Taken together, these strands show that ABHI serves as a bridge concept between ordering, deviation, and dependence. The precise mathematical object changes across settings, but the organizing principle remains the same: integrate a hypograph-based quantity so that geometric information accumulated over a domain becomes statistically actionable.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Area-Based Hypograph Index (ABHI).