---
title: Dynamic Habitat Index (DHI)
url: https://www.emergentmind.com/topics/dynamic-habitat-index-dhi
type: topic
---

# Dynamic Habitat Index (DHI)

Searching arXiv for the specified paper and any directly relevant DHI background sources.
The Dynamic Habitat Index (DHI) is a remote sensing-based measure of habitat productivity and variability that was compared with land-use/land-cover (LULC) metrics for species distribution modelling (SDM) in Île-de-France in "Modelling species distributions using remote sensing predictors: Comparing Dynamic Habitat Index and LULC" [2509.14862]. In that study, DHI is presented as a compact summary of intra-annual ecosystem functioning derived from time series of radiometric indices, classically defined through cumulative productivity, minimum cover, and variability, but operationalized there as annual mean, annual minimum, and annual maximum computed from Sentinel-2 observations for 2018 [2509.14862]. The study’s central result is that DHI-based and LULC-based SDMs had similar reliability by the Continuous Boyce Index, but diverged after binarization because DHI-based suitability surfaces were broader and more spatially neutral, whereas LULC-based models were more spatially constrained by distance effects [2509.14862].

## 1. Definition and conceptual basis

The study defines the Dynamic Habitat Index as a compact summary of intra-annual ecosystem functioning from remote sensing, classically built from a one-year time series $x_1, x_2, \ldots, x_T$ of a radiometric index [2509.14862]. In its canonical form, the three components are cumulative or annual productivity, minimum cover or greenness, and variability or seasonality. These are given as:
- $P = \sum_{t=1}^{T} x_t$
- $M = \min_t x_t$
- $S = \mathrm{CV}(x_t) = \mathrm{sd}(x_t)/\mathrm{mean}(x_t)$ [2509.14862]

In the Île-de-France study, the operationalization departs from the canonical cumulative/CV formulation and instead uses statistics described as robust and straightforward to compute with Sentinel-2 [2509.14862]. For each pixel, and separately for each index, the authors computed:
- $DHI_{\mathrm{mean}}(x) = (1/T)\sum_{t=1}^{T} x_t$
- $DHI_{\mathrm{min}}(x) = \min_t x_t$
- $DHI_{\mathrm{max}}(x) = \max_t x_t$ [2509.14862]

Here, $x_t$ denotes the index value from valid Sentinel-2 Level 2A observations in 2018, and $T$ is the number of valid acquisitions for that pixel in the year [2509.14862]. DHI is therefore treated as a three-component vector derived separately for each index rather than as a single composite score. No additional normalization of DHI components is described; standardization enters later through model selection and evaluation [2509.14862].

This formulation places DHI in the category of continuous, temporally integrated habitat predictors. A plausible implication is that its informational content differs fundamentally from categorical habitat descriptors because it summarizes ecosystem functioning directly from sensor-derived radiometric behavior rather than from human-defined land-cover classes.

## 2. Remote-sensing construction of DHI in the Île-de-France study

The DHI implementation in the study was based on Sentinel-2 Level 2A surface reflectance imagery for calendar year 2018, distributed by Theia, with atmospheric correction and cloud masking performed using the MAJA processor [2509.14862]. Indices were computed at 10 m native resolution and then upscaled to 50 m to match occurrence data accuracy [2509.14862]. For each index, all valid observations across the 2018 time series were aggregated to compute $DHI_{\mathrm{mean}}$, $DHI_{\mathrm{min}}$, and $DHI_{\mathrm{max}}$; specific compositing windows and gap-filling procedures are not detailed [2509.14862]. Orfeo ToolBox v7.2.0 was used to compute the index and DHI component layers [2509.14862].

Eighteen spectral or radiometric indices spanning vegetation, soil, and water were computed, each summarized by the three DHI statistics [2509.14862]. The indices listed include NDVI, TSAVI, GEMI, RVI, MNDWI, MSAVI, MSAVI2, NDWI/NDWI2, IPVI, SAVI, TNDVI, NDTI, BI, CI, and ISU, among others [2509.14862]. EVI, fPAR, LAI, and NPP proxies were not used [2509.14862]. The most influential remote-sensing predictors across species derived from NDVI and TSAVI, with $NDVI_{\mathrm{mean}}$ described as overwhelmingly dominant [2509.14862].

Several formulas are explicitly given for the remote-sensing indices used in the models. NDVI is defined as
$$
\mathrm{NDVI} = \frac{\mathrm{NIR} - \mathrm{Red}}{\mathrm{NIR} + \mathrm{Red}}
$$
with Sentinel-2 bands $\mathrm{NIR}=B8$ and $\mathrm{Red}=B4$ [2509.14862]. RVI is
$$
\mathrm{RVI} = \frac{\mathrm{NIR}}{\mathrm{Red}}
$$
[2509.14862]. TSAVI is
$$
\mathrm{TSAVI} = \frac{s(\mathrm{NIR} - s\,\mathrm{Red} - b)}{(s\,\mathrm{Red} + \mathrm{NIR} - b)(1 + s^{2})}
$$
where $s$ and $b$ are the slope and intercept of the soil line [2509.14862]. GEMI is
$$
\mathrm{GEMI} = \eta(1 - 0.25\,\eta) - \frac{\mathrm{Red} - 0.125}{1 - \mathrm{Red}}
$$
with
$$
\eta = \frac{2(\mathrm{NIR}^{2} - \mathrm{Red}^{2}) + 1.5\,\mathrm{NIR} + 0.5\,\mathrm{Red}}{\mathrm{NIR} + \mathrm{Red} + 0.5}
$$
[2509.14862]. MNDWI is
$$
\mathrm{MNDWI} = \frac{\mathrm{Green} - \mathrm{SWIR}}{\mathrm{Green} + \mathrm{SWIR}}
$$
with $\mathrm{Green}=B3$ and $\mathrm{SWIR}$ commonly $B11$ or $B12$ [2509.14862].

The study also notes that EVI was not computed. Its standard form is given only hypothetically and not used in the reported models [2509.14862]. This distinction matters because the paper’s conclusions concern a DHI formulation grounded in the indices actually included, not in a broader class of vegetation-function products.

## 3. Comparator framework: LULC descriptors and species distribution modelling

The comparison to DHI was performed using continuous predictors derived from a categorical LULC map rather than by using categorical land-cover labels directly [2509.14862]. LULC was obtained from a 2018 Sentinel-2 time series via a supervised Random Forest classification, specifically the OSO–CESBIO product, and aggregated into 11 classes: Urban area, Roads, Winter crops, Summer crops, Grassland, Orchards, Vineyards, Deciduous forest, Coniferous forest, Woody moorlands, and Water [2509.14862].

Two classes of LULC-derived predictors were used. The first consisted of area metrics, that is, the proportion or area of each class in local neighborhoods as provided by FRAGSTATS [2509.14862]. The second consisted of distance-to-class metrics. For each class $k$, the Euclidean distance from pixel $\mathbf{s}$ to the nearest pixel of class $k$ was defined as
$$
d_k(\mathbf{s}) = \min_{\mathbf{u} \in \Omega_k} \|\mathbf{s} - \mathbf{u}\|_2
$$
where $\Omega_k$ is the set of pixels belonging to class $k$ [2509.14862]. Distances were computed at the 50 m analysis grid and used directly as continuous predictors; no multi-scale buffers or transformations beyond the FRAGSTATS computation were reported [2509.14862].

The SDM experiment was conducted over Île-de-France, approximately $12{,}000\ \mathrm{km}^2$, described as temperate landscapes dominated by urban and intensive agriculture [2509.14862]. Species data covered eleven bird, amphibian, and mammal species from opportunistic occurrences in the Cettia ÎDF citizen database (GeoNat’ÎdF) from 2015 to 2020 [2509.14862]. The species were:
- Birds: *Athene noctua*, *Anthus pratensis*, *Pyrrhula pyrrhula*, *Sylvia curruca*
- Amphibians: *Bufo bufo*, *Hyla arborea*, *Ichthyosaura alpestris*, *Lissotriton vulgaris*, *Triturus cristatus*
- Mammals: *Eptesicus serotinus*, *Meles meles* [2509.14862]

Occurrence records were spatially thinned to 50 m, with one presence per $50 \times 50$ m pixel, to reduce spatial sorting bias and match environmental grid resolution [2509.14862]. The presence-pseudoabsence design used five pseudo-absence sets per species, randomly sampled across the background, with the number of pseudo-absences equal to the number of presences and equal weighting of presences and pseudo-absences [2509.14862].

Predictor selection differed between the remote-sensing and LULC branches. The remote-sensing set comprised 10 variables selected once and reused for all species after collinearity screening with Pearson $|r| \ge 0.7$ using `removeCollinearity` in `virtualspecies` [2509.14862]. The LULC set used species-specific variable importance from `biomod2`, retaining the 10 top predictors per species [2509.14862]. Modelling employed nine algorithms—GLM, GAM, MARS, ANN, FDA, CTA, GBM, RF, and MAXENT—with three calibrations per algorithm, a 70% train and 30% test split, and an ensemble defined as the arithmetic mean of suitability predictions from models with AUC $\ge 0.7$ [2509.14862]. The platform was `biomod2` v3.5-3 in R [2509.14862].

## 4. Evaluation framework and quantitative comparison

Model performance was assessed using the Continuous Boyce Index (CBI) and a calibrated AUC denoted AUCc [2509.14862]. CBI is defined in the study as a presence-only reliability metric based on a predicted-to-expected ratio across suitability bins. If bins $j=1,\ldots,J$ partition the suitability range, with $F_j$ the frequency of presences in bin $j$ and $E_j$ the expected frequency under random use proportional to area of bin $j$, then
$$
\mathrm{P/E}_j = F_j / E_j
$$
and CBI is computed as the Spearman rank correlation between bin mid-suitability and $\mathrm{P/E}_j$ across bins, ranging from $-1$ to $+1$ [2509.14862]. The study implemented CBI as in Hirzel et al. (2006) [2509.14862].

AUCc is described as presence-absence discrimination adjusted for spatial sorting bias by calibrating against a geographic null model [2509.14862]. Operationally,
$$
\mathrm{AUC}_\mathrm{c} = \mathrm{AUC} - \mathrm{AUC}_\mathrm{geo}
$$
where $\mathrm{AUC}_\mathrm{geo}$ arises from a null model driven by geographic distance to training presences [2509.14862]. The reported AUCc values ranged from 0.27 to 0.48, indicating low discrimination beyond geographic bias, with LULC consistently higher than remote sensing [2509.14862].

For thresholding, the study used the 10th percentile of suitability at presence locations. If the sorted suitability scores at presences are $s_{(1)} \le \cdots \le s_{(n)}$, the threshold is
$$
T_{0.10} = \mathrm{quantile}\{s_i\}_{i=1}^n(0.10)
$$
and pixels with suitability $\ge T_{0.10}$ are classified as suitable [2509.14862].

Overlap between remote-sensing and LULC niches was evaluated on continuous suitability surfaces normalized to sum to 1 across the domain, using Schoener’s $D$, Warren’s $I$, and Spearman rank correlation [2509.14862]. The formulas reported are:
$$
D = 1 - \tfrac{1}{2}\sum_i |p_i - q_i|
$$
$$
I = \sum_i \sqrt{p_i q_i}
$$
$$
\rho = 1 - \frac{6 \sum_i d_i^2}{n(n^2 - 1)}
$$
where $d_i$ is the rank difference between paired suitability values across pixels [2509.14862].

The main comparative results are concise but consequential. For CBI, both remote-sensing and LULC ensembles achieved high reliability, with values close to 1 for most species; *Eptesicus serotinus* was the main exception with weaker CBI [2509.14862]. For AUCc, LULC models consistently outperformed remote-sensing models after calibrating out spatial sorting bias, but all values remained low, indicating modest discrimination beyond geography, particularly for the remote-sensing branch [2509.14862].

Before binarization, the continuous suitability surfaces were broadly concordant. Schoener’s $D$ ranged from 0.722 to 0.827, Warren’s $I$ from 0.934 to 0.972, and Spearman $\rho$ from 0.461 to 0.757 across species [2509.14862]. After binarization, however, the niches diverged substantially. The projected overlap onto LULC niches was $O_{LU} = 73 \pm 8\%$ on average, whereas the projected overlap onto remote-sensing niches was $O_{RS} = 60 \pm 15\%$ on average [2509.14862]. The ratio of off-overlap, defined in the paper as $R$, was $2.61 \pm 1.12$ on average, with remote-sensing off-overlap generally much larger; *Hyla arborea* exhibited approximately $R \approx 5.23$ [2509.14862].

## 5. Spatial behavior, bias structure, and predictor importance

A central finding of the study is that the two predictor families impose different spatial behavior on suitability maps [2509.14862]. LULC-based models showed a pronounced distance effect: suitability decreased as distance to key LULC classes increased [2509.14862]. In contrast, remote-sensing-based models using continuous DHI predictors were not affected by distance-to-class or geographic sampling bias in the same way; their suitable areas extended beyond observed clusters and often appeared as scattered isolated pixels [2509.14862].

This contrast is visible most clearly after thresholding. The study states that remote-sensing niches were broader and more dispersed, while LULC niches were more spatially constrained and clustered around core habitats [2509.14862]. The ratio $R$ tended to decrease with mean inter-occurrence distance for most species, meaning that broader observed distributions usually yielded less disparity between remote-sensing and LULC off-overlap, although exceptions were reported for *Athene noctua*, *Lissotriton vulgaris*, and *Triturus cristatus* [2509.14862].

Feature-importance analysis identifies the specific DHI components responsible for much of the remote-sensing model behavior. Among DHI predictors, $NDVI_{\mathrm{mean}}$ was the most influential predictor across taxa, with permutation importance from 0.25 to 0.98 [2509.14862]. In several species it alone could suffice: *Bufo bufo*, *Ichthyosaura alpestris*, *Triturus cristatus*, *Lissotriton vulgaris*, and *Pyrrhula pyrrhula* [2509.14862]. For *Athene noctua*, $NDVI_{\mathrm{max}}$ was the sole sufficient variable [2509.14862]. Other contributing remote-sensing variables included $TSAVI_{\mathrm{max}}$, $GEMI_{\mathrm{max}}$, $GEMI_{\mathrm{min}}$, $MNDWI_{\mathrm{max}}$, $MNDWI_{\mathrm{min}}$, and $RVI_{\mathrm{max}}$ [2509.14862].

By contrast, the LULC branch exhibited no single dominant landscape metric, with maximum importance not exceeding 0.41 [2509.14862]. For every species, at least one distance-to-class variable ranked among the top three predictors; examples given include distance to deciduous forest, urban areas, orchards, and water [2509.14862]. Area metrics were also important, but the paper states that they contributed less to spatial constraint than the distance variables [2509.14862].

This pattern suggests that the contrast between DHI and LULC is not merely one of data source, but of predictor geometry. A plausible implication is that continuous functional summaries derived from radiometric time series can produce suitability fields less tightly tethered to mapped habitat edges than metrics defined directly by Euclidean proximity to categorical classes.

## 6. Interpretation, limitations, and methodological implications

The study identifies conditions under which DHI is preferable. It is recommended for regional-scale SDM when continuous, temporally integrated habitat proxies can mitigate biases caused by uneven sampling and avoid artificial spatial constraints imposed by categorical LULC classes and their distances [2509.14862]. It is also described as suitable when spatially neutral, sensor-derived measures of productivity and greenness are desired without reliance on human-defined classes [2509.14862].

At the same time, the paper emphasizes several limitations and caveats. The temporal window was one year only, specifically 2018; DHI sensitivity to anomalous years or interannual variability is described as real, and the authors note that incorporating multiple years and the variability term $S$ as coefficient of variation may better capture seasonality [2509.14862]. Cloud or snow contamination and variable scene counts per pixel can affect DHI robustness; although MAJA masks clouds, explicit gap-filling was not detailed [2509.14862]. The sensor and metric choice also matters: while NDVI-based DHI dominated in this study, the paper notes that other functional products such as fPAR, LAI, and GPP can outperform NDVI or EVI for biodiversity in some contexts [2509.14862].

The study also notes ecological blind spots shared, in different ways, by both predictor families. Habitat types not captured by productivity proxies or by coarse LULC typologies remain difficult to model; the paper gives the example of very small ponds critical to amphibians [2509.14862]. Remote-sensing maps tended to predict scattered suitable pixels, which the authors identify as possible overestimation, whereas LULC maps underpredicted away from observed clusters because of distance effects [2509.14862]. For that reason, combining both predictor families is presented as a way to mitigate complementary biases [2509.14862].

Generalizability is treated cautiously. The approach is described as transferable to other regions and taxa, but performance depends on species ecology, occurrence bias, the ecological fit of the LULC typology, and the temporal and spatial resolution of the remote-sensing data [2509.14862]. This suggests that DHI should not be understood as a universal replacement for LULC, but as a specific predictor family whose advantages emerge most clearly under spatial bias and categorical-distance constraints.

## 7. Position of DHI in regional SDM according to the study

The integrative conclusion of the paper characterizes the Dynamic Habitat Index as a compact, multi-component summary of intra-annual ecosystem functioning from remote sensing [2509.14862]. In the study’s implementation, it was operationalized with Sentinel-2 as three per-index statistics—annual mean, annual minimum, and annual maximum—computed for 18 vegetation, soil, and water indices and used as continuous predictors in SDMs [2509.14862]. Compared with LULC metrics based on area and Euclidean distance to classes, DHI-based models achieved similar reliability, with CBI near 1, but slightly lower calibrated discrimination after removing geographic sorting bias [2509.14862].

The decisive difference lay in the geometry of predicted niches. DHI predictors yielded suitability surfaces described as spatially neutral and not decaying with distance from observed habitats, thereby mitigating the geographic bias inherent to distance-based LULC predictors [2509.14862]. After binarization, remote-sensing or DHI niches were consistently broader and more dispersed, while LULC niches were more constrained and clustered, with quantitative overlap metrics showing that remote-sensing off-overlap was on average about 2.6 times larger than LULC off-overlap, with $R \approx 2.61 \pm 1.12$, $O_{LU} \approx 73\%$, and $O_{RS} \approx 60\%$ [2509.14862].

The dominance of $NDVI_{\mathrm{mean}}$ across species, and of $NDVI_{\mathrm{max}}$ for *Athene noctua*, underscores annual productivity and greenness as the principal functional drivers at the regional scale in this dataset [2509.14862]. The paper therefore supports DHI as a spatially neutral, temporally informed alternative to LULC for regional SDM, particularly when occurrence data are spatially biased and categorical habitat distances induce artificial constraints [2509.14862]. It simultaneously concludes that both predictor families are complementary: LULC captures structural composition and connectivity, while DHI captures functional dynamics and seasonality [2509.14862]. A plausible implication is that future regional SDM workflows may benefit most from joint formulations that preserve this complementarity rather than privileging either predictor family in isolation.

Source: https://www.emergentmind.com/topics/dynamic-habitat-index-dhi