Papers
Topics
Authors
Recent
Search
2000 character limit reached

Extended Generalized Pareto Distribution (EGPD)

Updated 14 July 2026
  • EGPD is a family of distributions that extends the Generalized Pareto Distribution to model full-range positive data while preserving its extreme tails.
  • EGPD employs flexible transformation functions to decouple bulk and tail behaviors, thereby stabilizing threshold-based estimations and improving inference.
  • EGPD methodologies have been extended to distributional regression, discrete data modeling, and multivariate settings, supporting applications from rainfall analysis to insurance claims.

Searching arXiv for recent and foundational papers on the Extended Generalized Pareto Distribution to support the encyclopedia entry. The Extended Generalized Pareto Distribution (EGPD) is a family of distributions introduced to model the full range of a positive random variable while preserving peaks-over-threshold behavior in the lower and upper tails. In the formulation reported for distributional regression, its cumulative distribution function is F(y)=G{Hξ(y/ψ)}, y>0F(y) = G\{ H_{\xi}(y/\psi) \}, \ y > 0, where HξH_{\xi} is the generalized Pareto distribution (GPD) cdf and G()G(\cdot) is a continuous cdf on [0,1][0,1] chosen so that F(y)F(y) preserves Pareto behavior in both tails (Carrer et al., 2022). In threshold-based tail estimation, related extended generalized Pareto models were proposed by incorporating an additional shape parameter while keeping the tail behaviour unaffected; this inclusion offers additional structure for the main body of the distribution, improves the stability of the modified scale, tail index and return level estimates to threshold choice, and allows a lower threshold to be selected (Papastathopoulos et al., 2011). Subsequent work has extended the EGPD framework to distributional regression for continuous data, count-data analogues, multivariate modeling, and multiscale rainfall aggregation (Carrer et al., 2022, Ahmad et al., 2022, Alotaibi et al., 7 Sep 2025, Gaetan et al., 2 Oct 2025, Ailliot et al., 13 Jan 2026).

1. Definition and core construction

A central EGPD construction for positive data is based on the GPD cdf

Hξ(z)={1(1+ξz)1/ξξ0 1exp(z)ξ=0,H_{\xi}(z) = \begin{cases} 1 - (1+\xi z)^{-1/\xi} & \xi \neq 0 \ 1 - \exp(-z) & \xi = 0 , \end{cases}

combined with a continuous cdf G()G(\cdot) on [0,1][0,1], yielding

F(y)=G{Hξ(y/ψ)},y>0F(y) = G\{ H_{\xi}(y/\psi) \}, \quad y > 0

with parameter vector θ=(ξ,ψ,κT)T\theta = (\xi, \psi, \kappa^T)^T (Carrer et al., 2022). In this parameterization, HξH_{\xi}0 and HξH_{\xi}1 correspond to the upper tail shape and scale, and HξH_{\xi}2 denotes additional shape or flexibility parameters entering through HξH_{\xi}3.

Four parameterizations of HξH_{\xi}4 are reported in the distributional-regression formulation (Carrer et al., 2022). They are: HξH_{\xi}5

HξH_{\xi}6

HξH_{\xi}7

where

HξH_{\xi}8

and

HξH_{\xi}9

A special-case relationship to the ordinary GPD is explicit: setting G()G(\cdot)0 in Model 1 reduces EGPD back to the ordinary GPD (Carrer et al., 2022). In the threshold-exceedance formulation of extended generalised Pareto models, the same idea appears through a transformation G()G(\cdot)1, where G()G(\cdot)2 is supported on G()G(\cdot)3; choosing flexible parametric distributions for G()G(\cdot)4 that include uniform as a special case yields families that generalize the GPD while retaining the original GPD tail behavior (Papastathopoulos et al., 2011).

The EGPD quantile function reported for simulation or quantile regression is

G()G(\cdot)5

This quantile representation is one of the mechanisms that makes the class suitable for simulation and regression-based applications (Carrer et al., 2022).

2. Tail behavior, lower-tail control, and relation to the GPD

A defining feature of EGPD formulations is the separation between body flexibility and tail preservation. In the threshold-estimation framework, the extra shape parameter G()G(\cdot)6 governs the body, not the tail, while G()G(\cdot)7 remains the tail index; thus, the EGPD families retain the original GPD tail behavior (Papastathopoulos et al., 2011). When G()G(\cdot)8, the reported EGP1, EGP2, and EGP3 models reduce to the standard GPD, and for G()G(\cdot)9 they reduce to variants of the exponential and gamma distributions (Papastathopoulos et al., 2011).

The same lower-tail versus upper-tail separation is emphasized in later formulations. In one univariate specification,

[0,1][0,1]0

where [0,1][0,1]1 controls the lower tail, [0,1][0,1]2 is a scale parameter, and [0,1][0,1]3 controls the upper tail (Alotaibi et al., 7 Sep 2025). In that representation, as [0,1][0,1]4, [0,1][0,1]5, and as [0,1][0,1]6, [0,1][0,1]7 (Alotaibi et al., 7 Sep 2025).

A more general formulation based on a transfer function [0,1][0,1]8 writes

[0,1][0,1]9

with density

F(y)F(y)0

where F(y)F(y)1 is the density of F(y)F(y)2 (Gaetan et al., 2 Oct 2025). In this account, F(y)F(y)3 flexibly controls the lower tail behavior, F(y)F(y)4 controls the heaviness of the upper tail, and F(y)F(y)5 allows a smooth, flexible transition between the two tails, controlling distributional behavior in the central range (Gaetan et al., 2 Oct 2025). A closely related parameterization states

F(y)F(y)6

and

F(y)F(y)7

with lower-tail expansion

F(y)F(y)8

and upper-tail expansion

F(y)F(y)9

(Ailliot et al., 13 Jan 2026).

This body-tail decomposition is the principal reason EGPD models are repeatedly described as threshold-free or threshold-bypassing. A plausible implication is that EGPD occupies an intermediate position between classical peaks-over-threshold asymptotics and whole-distribution parametric modeling: it preserves extreme-value-theoretic tail structure while making the lower tail and central mass explicitly modelable.

3. Threshold selection, stability, and tail estimation

The threshold-selection problem is a primary motivation for EGPD methodology. In threshold exceedance analysis with the standard GPD, selecting a high threshold can leave almost no exceedances when few observations are recorded, reducing power and stability of inference (Papastathopoulos et al., 2011). Extended generalised Pareto models were proposed specifically to improve model fit for lower thresholds, increase the usable sample size, and stabilize inference for the modified scale, tail index, and return level estimates across a range of thresholds (Papastathopoulos et al., 2011).

In that setting, all three parameters Hξ(z)={1(1+ξz)1/ξξ0 1exp(z)ξ=0,H_{\xi}(z) = \begin{cases} 1 - (1+\xi z)^{-1/\xi} & \xi \neq 0 \ 1 - \exp(-z) & \xi = 0 , \end{cases}0 are estimated jointly by maximum likelihood, and standard errors or confidence intervals are calculated via standard likelihood theory or the delta method (Papastathopoulos et al., 2011). The fitted shape parameter Hξ(z)={1(1+ξz)1/ξξ0 1exp(z)ξ=0,H_{\xi}(z) = \begin{cases} 1 - (1+\xi z)^{-1/\xi} & \xi \neq 0 \ 1 - \exp(-z) & \xi = 0 , \end{cases}1 can also be tested for departure from Hξ(z)={1(1+ξz)1/ξξ0 1exp(z)ξ=0,H_{\xi}(z) = \begin{cases} 1 - (1+\xi z)^{-1/\xi} & \xi \neq 0 \ 1 - \exp(-z) & \xi = 0 , \end{cases}2 using a likelihood ratio test, which provides a formal check of GPD adequacy (Papastathopoulos et al., 2011). Plotting Hξ(z)={1(1+ξz)1/ξξ0 1exp(z)ξ=0,H_{\xi}(z) = \begin{cases} 1 - (1+\xi z)^{-1/\xi} & \xi \neq 0 \ 1 - \exp(-z) & \xi = 0 , \end{cases}3 against threshold is reported as a direct diagnostic: flatness near Hξ(z)={1(1+ξz)1/ξξ0 1exp(z)ξ=0,H_{\xi}(z) = \begin{cases} 1 - (1+\xi z)^{-1/\xi} & \xi \neq 0 \ 1 - \exp(-z) & \xi = 0 , \end{cases}4 suggests the GPD is reasonable, while deviation suggests inadequacy (Papastathopoulos et al., 2011).

The simulation study summarized for this threshold-based framework states that, for simulations based on normal data with Hξ(z)={1(1+ξz)1/ξξ0 1exp(z)ξ=0,H_{\xi}(z) = \begin{cases} 1 - (1+\xi z)^{-1/\xi} & \xi \neq 0 \ 1 - \exp(-z) & \xi = 0 , \end{cases}5, Hξ(z)={1(1+ξz)1/ξξ0 1exp(z)ξ=0,H_{\xi}(z) = \begin{cases} 1 - (1+\xi z)^{-1/\xi} & \xi \neq 0 \ 1 - \exp(-z) & \xi = 0 , \end{cases}6, and Hξ(z)={1(1+ξz)1/ξξ0 1exp(z)ξ=0,H_{\xi}(z) = \begin{cases} 1 - (1+\xi z)^{-1/\xi} & \xi \neq 0 \ 1 - \exp(-z) & \xi = 0 , \end{cases}7, EGP1 and EGP2 achieved lower Root Mean Square Error (RMSE) in return level estimation than GPD, especially for small samples and at low thresholds (Papastathopoulos et al., 2011). The same summary reports that the minimum RMSE for EGPD estimates was found at very low thresholds, while for GPD the optimal thresholds were higher; as sample size increased, the performance difference between GP and EGP diminished (Papastathopoulos et al., 2011). In the River Nidd flood data, EGP1 parameter estimates and return level estimates were reported as much more stable across thresholds than their GP counterparts (Papastathopoulos et al., 2011).

These results concern threshold exceedance models rather than the later full-range formulations, but they establish a recurring EGPD theme: the extra body-shape parameter is introduced not to alter the tail index but to reduce sensitivity to threshold choice. This suggests that, even when a threshold-based analysis is retained, EGPD can function as a robustness device against threshold misspecification.

4. Distributional regression for continuous positive data

A major development is the extension of EGPD to distributional regression, in which the parameters of the distribution depend on covariates through additive predictors (Carrer et al., 2022). The conditional specification is

Hξ(z)={1(1+ξz)1/ξξ0 1exp(z)ξ=0,H_{\xi}(z) = \begin{cases} 1 - (1+\xi z)^{-1/\xi} & \xi \neq 0 \ 1 - \exp(-z) & \xi = 0 , \end{cases}8

where Hξ(z)={1(1+ξz)1/ξξ0 1exp(z)ξ=0,H_{\xi}(z) = \begin{cases} 1 - (1+\xi z)^{-1/\xi} & \xi \neq 0 \ 1 - \exp(-z) & \xi = 0 , \end{cases}9 may include covariates such as space or time, and each parameter G()G(\cdot)0 is modeled through

G()G(\cdot)1

For Model 1, the reported links are

G()G(\cdot)2

Estimation is performed via penalized likelihood with smoothness penalties on the functions G()G(\cdot)3, using basis expansions (Carrer et al., 2022).

The framework is presented within Generalized Additive Models for Location, Scale and Shape (GAMLSS), where all parameters can vary with covariates and each has its own link function (Carrer et al., 2022). Example additive specifications include a time-only predictor,

G()G(\cdot)4

and a space-time predictor,

G()G(\cdot)5

with thin-plate splines in spatial coordinates and cyclic splines in time (Carrer et al., 2022). Estimation and inference utilize standard tools for penalized regression and are implemented in the R package gamlss (Carrer et al., 2022).

The hourly rainfall case study uses positive rainfall from the MeteoNet dataset over the North-West region of France, with longitude, latitude, and cyclic time as covariates; small values below 0.5 mm are censored to improve upper tail fits (Carrer et al., 2022). Models where EGPD parameters are constant or vary in time and space are compared using information criteria (AIC, BIC, global deviance) and validation deviance (Carrer et al., 2022). The reported findings are that models allowing distributions to vary smoothly in time performed much better than constant-parameter models, spatial variation gave relatively small extra improvement for these data, and EGPD4 was the most accurate for modeling the full range, particularly extremes, outperforming both the Gamma and lower-parameter EGPD variants (Carrer et al., 2022). P-P plots are reported to confirm the superiority of EGPD4, especially for upper tails (Carrer et al., 2022).

This regression extension places EGPD within modern distributional modeling rather than only within classical EVT. It also changes the interpretation of the parameters: they are no longer merely static shape and scale terms, but structured latent functions over covariate domains such as seasonality and geography.

5. Count-data extensions and zero inflation

The EGPD idea has also been extended to discrete data through the Discrete Extended Generalized Pareto Distribution (DEGPD), intended to model threshold exceedances of integer random variables without separating bulk and extremes (Ahmad et al., 2022). The reported probability mass function is

G()G(\cdot)6

with cdf

G()G(\cdot)7

and quantile function

G()G(\cdot)8

Several parametric families for G()G(\cdot)9 are proposed: power law, beta family, power of beta, and mixture of powers (Ahmad et al., 2022).

To handle excess zeros, a Zero-Inflated DEGPD (ZIDEGPD) is defined by mixing a point mass at zero with the DEGPD: [0,1][0,1]0 The count-data regression extension again relates parameters such as [0,1][0,1]1 to additive smoothed predictors via appropriate links, with penalized maximum likelihood estimation implemented via extensions to the evgam R package (Ahmad et al., 2022).

A later count-data paper develops new flexible versions of extended generalized Pareto models for counts and states three modeling scenarios: the entire distribution of the data, including both bulk and tail and bypassing the threshold selection step; the entire distribution along with zero inflation; and the tail of the distribution for low threshold exceedances (Ahmad et al., 2024). In that formulation, the transformation [0,1][0,1]2 can be a power model, a truncated Normal model, or a truncated Beta model (Ahmad et al., 2024). For the Automobile Insurance Company Complaint Rankings data, all three EGPD models are reported to offer better tail fits and overall performance than standard count models and DGPD, with Bayesian Information Criterion favoring the truncated Normal version ([0,1][0,1]3) (Ahmad et al., 2024). The same summary reports that a zero-inflated version successfully modeled hospital doctor visits with a heavy tail and 42% zeros (Ahmad et al., 2024).

These discrete formulations preserve the structural logic of the continuous EGPD: a GPD-based tail mechanism is retained, while a separate transformation controls the lower tail, bulk, and zero inflation. A plausible implication is that EGPD is better viewed as a modeling principle—smooth tail-preserving extension of generalized Pareto behavior—than as a single fixed parametric family.

6. Multivariate and multiscale developments

Recent work extends EGPD beyond univariate margins. One multivariate proposal models low, moderate, and large intensities without threshold selection steps by decomposing multivariate data into radial and angular components, with the radial component modeled using a semi-parametric EGPD and the angular distribution permitted to vary conditionally (Gaetan et al., 2 Oct 2025). In the univariate case used there,

[0,1][0,1]4

and the upper and lower tails satisfy

[0,1][0,1]5

(Gaetan et al., 2 Oct 2025). The multivariate construction decomposes [0,1][0,1]6 as [0,1][0,1]7, where [0,1][0,1]8 is a radial component and [0,1][0,1]9 is angular, and proposes estimating the radial transfer-function density using Bernstein polynomials together with maximum likelihood for F(y)=G{Hξ(y/ψ)},y>0F(y) = G\{ H_{\xi}(y/\psi) \}, \quad y > 00 (Gaetan et al., 2 Oct 2025).

Another multivariate formulation defines

F(y)=G{Hξ(y/ψ)},y>0F(y) = G\{ H_{\xi}(y/\psi) \}, \quad y > 01

where F(y)=G{Hξ(y/ψ)},y>0F(y) = G\{ H_{\xi}(y/\psi) \}, \quad y > 02 has a univariate EGPD cdf, F(y)=G{Hξ(y/ψ)},y>0F(y) = G\{ H_{\xi}(y/\psi) \}, \quad y > 03 and F(y)=G{Hξ(y/ψ)},y>0F(y) = G\{ H_{\xi}(y/\psi) \}, \quad y > 04 are simplex-valued random vectors controlling lower-tail and upper-tail dependence, and F(y)=G{Hξ(y/ψ)},y>0F(y) = G\{ H_{\xi}(y/\psi) \}, \quad y > 05 is a continuous cdf used as a dynamic weight (Alotaibi et al., 7 Sep 2025). This model is reported to have the following properties: its marginal distributions behave like univariate eGPDs; its lower and upper joint tails comply with multivariate extreme-value theory, with key parameters separately controlling dependence in each joint tail; and the model allows for fast simulation and is thus amenable to simulation-based inference (Alotaibi et al., 7 Sep 2025). Estimation is proposed through modern neural approaches, including neural Bayes estimators and neural posterior estimators, so that a neural network, once trained, can provide point estimates, credible intervals, or full posterior approximations in a fraction of a second (Alotaibi et al., 7 Sep 2025).

A different line of development studies aggregated rainfall across scales from sub-hourly to weekly using EGPD (Ailliot et al., 13 Jan 2026). The reported framework models all rainfall intensities across scales and establishes a general result on the behavior of EGPD variables under various aggregation procedures (Ailliot et al., 13 Jan 2026). If F(y)=G{Hξ(y/ψ)},y>0F(y) = G\{ H_{\xi}(y/\psi) \}, \quad y > 06 and F(y)=G{Hξ(y/ψ)},y>0F(y) = G\{ H_{\xi}(y/\psi) \}, \quad y > 07 is a sum or more generally a transformation of iid EGPDs, then under broad conditions

F(y)=G{Hξ(y/ψ)},y>0F(y) = G\{ H_{\xi}(y/\psi) \}, \quad y > 08

where F(y)=G{Hξ(y/ψ)},y>0F(y) = G\{ H_{\xi}(y/\psi) \}, \quad y > 09 is determined by the aggregation structure and θ=(ξ,ψ,κT)T\theta = (\xi, \psi, \kappa^T)^T0 is a new bulk cdf (Ailliot et al., 13 Jan 2026). To address the absence of a closed-form likelihood for aggregated rainfall, the paper links the EGPD class to Poisson compound sums and uses the Panjer algorithm for efficient composite likelihood evaluation (Ailliot et al., 13 Jan 2026). The empirical claim reported there is that only eight parameters are needed per station to capture scales from six minutes to three days, and that the resulting return levels do not cross across scales (Ailliot et al., 13 Jan 2026).

Taken together, these multivariate and multiscale developments place EGPD within a broader class of models that seek coherence across the entire support, across dependence structures, and across aggregation levels, rather than only within upper-tail asymptotics.

7. Software, implementation, and modeling practice

The continuous distributional-regression paper provides an add-on script for the R package gamlss, available at https://github.com/noemielc/egpd4gamlss, to implement the EGPD family in a generic way (Carrer et al., 2022). The script can assemble new EGPD family distributions and supports density, cdf, quantile, and random generation, with symbolic differentiation via the Deriv R package for likelihood optimization (Carrer et al., 2022). The supplied examples include creation of an EGPD1 family and fitting baseline or time-varying models with gamlss (Carrer et al., 2022).

For count data, the methodology is implemented via extensions to the evgam R package, and a repository is provided at github.com/touqeerahmadunipd/degpd-and-zidegpd (Ahmad et al., 2022). In both settings, the estimation strategy is penalized maximum likelihood, with smoothing penalties on nonlinear additive predictors (Carrer et al., 2022, Ahmad et al., 2022).

Several recurring modeling strategies appear across the literature summarized here. First, EGPD is repeatedly used to bypass or relax threshold selection, either by replacing a threshold-based analysis entirely or by stabilizing threshold-based inference (Papastathopoulos et al., 2011, Carrer et al., 2022, Ahmad et al., 2022). Second, the extra shape mechanism—variously denoted θ=(ξ,ψ,κT)T\theta = (\xi, \psi, \kappa^T)^T1, θ=(ξ,ψ,κT)T\theta = (\xi, \psi, \kappa^T)^T2, θ=(ξ,ψ,κT)T\theta = (\xi, \psi, \kappa^T)^T3, or θ=(ξ,ψ,κT)T\theta = (\xi, \psi, \kappa^T)^T4—is used to control the lower tail and bulk while preserving upper-tail interpretation through θ=(ξ,ψ,κT)T\theta = (\xi, \psi, \kappa^T)^T5 (Carrer et al., 2022, Ahmad et al., 2024, Gaetan et al., 2 Oct 2025). Third, empirical comparisons are commonly made against Gamma models for rainfall, and Poisson, Negative Binomial, or zero-inflated variants for counts (Carrer et al., 2022, Ahmad et al., 2022, Ahmad et al., 2024).

A common misconception is that EGPD is merely a heavier-tailed GPD. The reported formulations do not support that characterization. Rather, they preserve the tail index interpretation of the GPD and introduce additional structure for the main body or lower tail, with the stated goal of modeling the entire positive support or stabilizing threshold analyses (Papastathopoulos et al., 2011, Carrer et al., 2022). Another potential misconception is that all EGPD formulations are identical. The data instead show multiple constructions: threshold-exceedance EGP1–EGP3 models, full-range positive-data EGPDs using θ=(ξ,ψ,κT)T\theta = (\xi, \psi, \kappa^T)^T6 or θ=(ξ,ψ,κT)T\theta = (\xi, \psi, \kappa^T)^T7, discrete DEGPD and ZIDEGPD analogues, and multivariate radial-angular or stochastic-mixture formulations (Papastathopoulos et al., 2011, Carrer et al., 2022, Ahmad et al., 2022, Alotaibi et al., 7 Sep 2025, Gaetan et al., 2 Oct 2025).

The overall trajectory of the literature suggests a unifying theme rather than a single canonical parameterization. That theme is a tail-compliant extension of generalized Pareto structure to the full distribution, with separate mechanisms for lower-tail or bulk flexibility, covariate dependence, discreteness, multivariate dependence, or aggregation.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Extended Generalized Pareto Distribution (EGPD).