---
title: Extended Generalized Pareto Distribution (EGPD)
url: https://www.emergentmind.com/topics/extended-generalized-pareto-distribution-egpd
type: topic
---

# Extended Generalized Pareto Distribution (EGPD)

Searching arXiv for recent and foundational papers on the Extended Generalized Pareto Distribution to support the encyclopedia entry.
The Extended Generalized Pareto Distribution (EGPD) is a family of distributions introduced to model the full range of a positive random variable while preserving peaks-over-threshold behavior in the lower and upper tails. In the formulation reported for distributional regression, its cumulative distribution function is \( F(y) = G\{ H_{\xi}(y/\psi) \}, \ y > 0 \), where \(H_{\xi}\) is the generalized Pareto distribution (GPD) cdf and \(G(\cdot)\) is a continuous cdf on \([0,1]\) chosen so that \(F(y)\) preserves Pareto behavior in both tails [2209.04660]. In threshold-based tail estimation, related extended generalized Pareto models were proposed by incorporating an additional shape parameter while keeping the tail behaviour unaffected; this inclusion offers additional structure for the main body of the distribution, improves the stability of the modified scale, tail index and return level estimates to threshold choice, and allows a lower threshold to be selected [1111.6899]. Subsequent work has extended the EGPD framework to distributional regression for continuous data, count-data analogues, multivariate modeling, and multiscale rainfall aggregation [2209.04660] [2210.15253] [2509.05982] [2510.02152] [2601.08350].

## 1. Definition and core construction

A central EGPD construction for positive data is based on the GPD cdf
\[
H_{\xi}(z) =
\begin{cases}
1 - (1+\xi z)^{-1/\xi} & \xi \neq 0 \\
1 - \exp(-z) & \xi = 0 ,
\end{cases}
\]
combined with a continuous cdf \(G(\cdot)\) on \([0,1]\), yielding
\[
F(y) = G\{ H_{\xi}(y/\psi) \}, \quad y > 0
\]
with parameter vector \(\theta = (\xi, \psi, \kappa^T)^T\) [2209.04660]. In this parameterization, \(\xi\) and \(\psi\) correspond to the upper tail shape and scale, and \(\kappa\) denotes additional shape or flexibility parameters entering through \(G\).

Four parameterizations of \(G\) are reported in the distributional-regression formulation [2209.04660]. They are:
\[
G(u; \kappa) = u^{\kappa_1}, \ \kappa_1 > 0,
\]
\[
G(u; \kappa) = \kappa_3 u^{\kappa_1} + (1-\kappa_3)u^{\kappa_2}, \ \kappa_2 \geq \kappa_1 > 0,\ \kappa_3 \in [0,1],
\]
\[
G(u; \kappa) = 1 - Q_{\kappa_1}[(1-u)^{\kappa_1}],
\]
where
\[
Q_{\kappa_1}(u) = \frac{1+\kappa_1}{\kappa_1} u^{1/\kappa_1} \left( 1 - \frac{u}{1+\kappa_1} \right),
\]
and
\[
G(u;\kappa) = [1-Q_{\kappa_1}\{(1-u)^{\kappa_1}\}]^{\kappa_2/2}, \ \kappa_1, \kappa_2 > 0.
\]

A special-case relationship to the ordinary GPD is explicit: setting \(\kappa_1=1\) in Model 1 reduces EGPD back to the ordinary GPD [2209.04660]. In the threshold-exceedance formulation of extended generalised Pareto models, the same idea appears through a transformation \(W = F^{-1}(V; \sigma, \xi)\), where \(V\) is supported on \([0,1]\); choosing flexible parametric distributions for \(V\) that include uniform as a special case yields families that generalize the GPD while retaining the original GPD tail behavior [1111.6899].

The EGPD quantile function reported for simulation or quantile regression is
\[
y_p = F^{-1}(p) =
\begin{cases}
\frac{\psi}{\xi} \left[ \{ 1 - G^{-1}(p) \}^{-\xi} - 1 \right], & \xi \neq 0 \\
-\psi \log{ \{ 1 - G^{-1}(p) \} }, & \xi = 0 .
\end{cases}
\]
This quantile representation is one of the mechanisms that makes the class suitable for simulation and regression-based applications [2209.04660].

## 2. Tail behavior, lower-tail control, and relation to the GPD

A defining feature of EGPD formulations is the separation between body flexibility and tail preservation. In the threshold-estimation framework, the extra shape parameter \(\kappa\) governs the body, not the tail, while \(\xi\) remains the tail index; thus, the EGPD families retain the original GPD tail behavior [1111.6899]. When \(\kappa = 1\), the reported EGP1, EGP2, and EGP3 models reduce to the standard GPD, and for \(\xi = 0\) they reduce to variants of the exponential and gamma distributions [1111.6899].

The same lower-tail versus upper-tail separation is emphasized in later formulations. In one univariate specification,
\[
F_Y(y) =
\begin{cases}
\left[ 1 - \left( 1 + \xi y/\sigma \right)^{-1/\xi} \right]^\kappa, & \xi \neq 0 \\
\left[ 1 - \exp(-y/\sigma) \right]^\kappa, & \xi = 0 ,
\end{cases}
\]
where \(\kappa > 0\) controls the lower tail, \(\sigma > 0\) is a scale parameter, and \(\xi\) controls the upper tail [2509.05982]. In that representation, as \( y \to 0 \), \( F_Y(y) \sim (y/\sigma)^\kappa \), and as \( y \to \infty \), \( 1 - F_Y(y) \sim \kappa (\xi y/\sigma)^{-1/\xi} \) [2509.05982].

A more general formulation based on a transfer function \(B(\cdot)\) writes
\[
F(x) = B\left\{ H_\xi(x)^\kappa \right\},
\]
with density
\[
f(x) = \kappa \, h_\xi(x) \, H_\xi(x)^{\kappa - 1} \, b\left\{ H_\xi(x)^\kappa \right\},
\]
where \(b(u)\) is the density of \(B(\cdot)\) [2510.02152]. In this account, \(\kappa\) flexibly controls the lower tail behavior, \(\xi\) controls the heaviness of the upper tail, and \(B(\cdot)\) allows a smooth, flexible transition between the two tails, controlling distributional behavior in the central range [2510.02152]. A closely related parameterization states
\[
F(y) = B\left( H_\xi^\kappa(y/\sigma) \right), \quad y \ge 0,
\]
and
\[
f(y) = \frac{\kappa}{\sigma} h_\xi(y/\sigma) H_\xi^{\kappa-1}(y/\sigma) b\left( H_\xi^\kappa(y/\sigma) \right), \quad y>0,
\]
with lower-tail expansion
\[
P(Y \leq y) = B(0) + b(0^+)\left(\frac{y}{\sigma}\right)^\kappa + o(y^\kappa)
\]
and upper-tail expansion
\[
P(Y > y) = \kappa b(1)\overline{H}_\xi(y/\sigma) + o(y^{-1/\xi})
\]
[2601.08350].

This body-tail decomposition is the principal reason EGPD models are repeatedly described as threshold-free or threshold-bypassing. A plausible implication is that EGPD occupies an intermediate position between classical peaks-over-threshold asymptotics and whole-distribution parametric modeling: it preserves extreme-value-theoretic tail structure while making the lower tail and central mass explicitly modelable.

## 3. Threshold selection, stability, and tail estimation

The threshold-selection problem is a primary motivation for EGPD methodology. In threshold exceedance analysis with the standard GPD, selecting a high threshold can leave almost no exceedances when few observations are recorded, reducing power and stability of inference [1111.6899]. Extended generalised Pareto models were proposed specifically to improve model fit for lower thresholds, increase the usable sample size, and stabilize inference for the modified scale, tail index, and return level estimates across a range of thresholds [1111.6899].

In that setting, all three parameters \((\kappa,\sigma,\xi)\) are estimated jointly by maximum likelihood, and standard errors or confidence intervals are calculated via standard likelihood theory or the delta method [1111.6899]. The fitted shape parameter \(\kappa\) can also be tested for departure from \(1\) using a likelihood ratio test, which provides a formal check of GPD adequacy [1111.6899]. Plotting \(\hat{\kappa}\) against threshold is reported as a direct diagnostic: flatness near \(\kappa=1\) suggests the GPD is reasonable, while deviation suggests inadequacy [1111.6899].

The simulation study summarized for this threshold-based framework states that, for simulations based on normal data with \(n=100\), \(1000\), and \(10000\), EGP1 and EGP2 achieved lower Root Mean Square Error (RMSE) in return level estimation than GPD, especially for small samples and at low thresholds [1111.6899]. The same summary reports that the minimum RMSE for EGPD estimates was found at very low thresholds, while for GPD the optimal thresholds were higher; as sample size increased, the performance difference between GP and EGP diminished [1111.6899]. In the River Nidd flood data, EGP1 parameter estimates and return level estimates were reported as much more stable across thresholds than their GP counterparts [1111.6899].

These results concern threshold exceedance models rather than the later full-range formulations, but they establish a recurring EGPD theme: the extra body-shape parameter is introduced not to alter the tail index but to reduce sensitivity to threshold choice. This suggests that, even when a threshold-based analysis is retained, EGPD can function as a robustness device against threshold misspecification.

## 4. Distributional regression for continuous positive data

A major development is the extension of EGPD to distributional regression, in which the parameters of the distribution depend on covariates through additive predictors [2209.04660]. The conditional specification is
\[
Y|\boldsymbol{x} \sim \mathsf{EGPD}(\cdot; \theta(\boldsymbol{x})),
\]
where \(\boldsymbol{x}\) may include covariates such as space or time, and each parameter \(\theta_j(\boldsymbol{x})\) is modeled through
\[
\eta_j(\boldsymbol{x}) = f_{j1}(\boldsymbol{x}) + \cdots + f_{jJ_j}(\boldsymbol{x}),
\qquad
\theta_j(\boldsymbol{x}) = h_j(\eta_j(\boldsymbol{x})).
\]
For Model 1, the reported links are
\[
\xi(\boldsymbol{x}) = \eta_\xi(\boldsymbol{x}), \qquad
\psi(\boldsymbol{x}) = \exp(\eta_\psi(\boldsymbol{x})), \qquad
\kappa_1(\boldsymbol{x}) = \exp(\eta_{\kappa_1}(\boldsymbol{x})).
\]
Estimation is performed via penalized likelihood with smoothness penalties on the functions \(f_{jk}\), using basis expansions [2209.04660].

The framework is presented within Generalized Additive Models for Location, Scale and Shape (GAMLSS), where all parameters can vary with covariates and each has its own link function [2209.04660]. Example additive specifications include a time-only predictor,
\[
\eta_j(\boldsymbol{x}) = \beta_0 + s_t(x_t),
\]
and a space-time predictor,
\[
\eta_j(\boldsymbol{x}) = \beta_0 + TP(x_l, x_L) + s_t(x_t),
\]
with thin-plate splines in spatial coordinates and cyclic splines in time [2209.04660]. Estimation and inference utilize standard tools for penalized regression and are implemented in the R package `gamlss` [2209.04660].

The hourly rainfall case study uses positive rainfall from the MeteoNet dataset over the North-West region of France, with longitude, latitude, and cyclic time as covariates; small values below 0.5 mm are censored to improve upper tail fits [2209.04660]. Models where EGPD parameters are constant or vary in time and space are compared using information criteria (AIC, BIC, global deviance) and validation deviance [2209.04660]. The reported findings are that models allowing distributions to vary smoothly in time performed much better than constant-parameter models, spatial variation gave relatively small extra improvement for these data, and EGPD4 was the most accurate for modeling the full range, particularly extremes, outperforming both the Gamma and lower-parameter EGPD variants [2209.04660]. P-P plots are reported to confirm the superiority of EGPD4, especially for upper tails [2209.04660].

This regression extension places EGPD within modern distributional modeling rather than only within classical EVT. It also changes the interpretation of the parameters: they are no longer merely static shape and scale terms, but structured latent functions over covariate domains such as seasonality and geography.

## 5. Count-data extensions and zero inflation

The EGPD idea has also been extended to discrete data through the Discrete Extended Generalized Pareto Distribution (DEGPD), intended to model threshold exceedances of integer random variables without separating bulk and extremes [2210.15253]. The reported probability mass function is
\[
\Pr(Y = y) = G\left( F(y+1;\sigma, \xi) \right) - G\left( F(y;\sigma, \xi) \right),\qquad y \in \mathbb{N}_0,
\]
with cdf
\[
\Pr(Y \le y) = G\left( F(y+1; \sigma, \xi) \right),
\]
and quantile function
\[
q_p =
\begin{cases}
\left\lceil \frac{\sigma}{\xi} \left[ \left( 1 - G^{-1}(p) \right)^{-\xi} - 1 \right] \right\rceil - 1, & \xi > 0 \\
\left\lceil -\sigma \log \left(1 - G^{-1}(p)\right) \right\rceil - 1, & \xi = 0 .
\end{cases}
\]
Several parametric families for \(G\) are proposed: power law, beta family, power of beta, and mixture of powers [2210.15253].

To handle excess zeros, a Zero-Inflated DEGPD (ZIDEGPD) is defined by mixing a point mass at zero with the DEGPD:
\[
\Pr(Y = y) =
\begin{cases}
\pi + (1 - \pi) G(F(1; \sigma, \xi); \psi), & y = 0 \\
(1-\pi) \left[ G(F(y+1; \sigma, \xi); \psi) - G(F(y; \sigma, \xi); \psi) \right], & y \geq 1 .
\end{cases}
\]
The count-data regression extension again relates parameters such as \(\sigma, \xi, \kappa, \pi\) to additive smoothed predictors via appropriate links, with penalized maximum likelihood estimation implemented via extensions to the `evgam` R package [2210.15253].

A later count-data paper develops new flexible versions of extended generalized Pareto models for counts and states three modeling scenarios: the entire distribution of the data, including both bulk and tail and bypassing the threshold selection step; the entire distribution along with zero inflation; and the tail of the distribution for low threshold exceedances [2409.18719]. In that formulation, the transformation \(\mathcal{G}\) can be a power model, a truncated Normal model, or a truncated Beta model [2409.18719]. For the Automobile Insurance Company Complaint Rankings data, all three EGPD models are reported to offer better tail fits and overall performance than standard count models and DGPD, with Bayesian Information Criterion favoring the truncated Normal version (\(M_2\)) [2409.18719]. The same summary reports that a zero-inflated version successfully modeled hospital doctor visits with a heavy tail and 42% zeros [2409.18719].

These discrete formulations preserve the structural logic of the continuous EGPD: a GPD-based tail mechanism is retained, while a separate transformation controls the lower tail, bulk, and zero inflation. A plausible implication is that EGPD is better viewed as a modeling principle—smooth tail-preserving extension of generalized Pareto behavior—than as a single fixed parametric family.

## 6. Multivariate and multiscale developments

Recent work extends EGPD beyond univariate margins. One multivariate proposal models low, moderate, and large intensities without threshold selection steps by decomposing multivariate data into radial and angular components, with the radial component modeled using a semi-parametric EGPD and the angular distribution permitted to vary conditionally [2510.02152]. In the univariate case used there,
\[
F(x) = B\left\{ H_\xi(x)^\kappa \right\},
\qquad
F^{-1}(u) = H_\xi^{-1}\left((B^{-1}(u))^{1/\kappa}\right),
\]
and the upper and lower tails satisfy
\[
\lim_{x \to \infty} \frac{\overline{F}(x)}{\kappa \overline{H}_\xi(x)} = b(1),
\qquad
\lim_{x \to 0} \frac{F(x)}{x^\kappa} = b(0)
\]
[2510.02152]. The multivariate construction decomposes \(\boldsymbol{X}\) as \( \boldsymbol{X} = R \cdot \boldsymbol{U} \), where \(R\) is a radial component and \(\boldsymbol{U}\) is angular, and proposes estimating the radial transfer-function density using Bernstein polynomials together with maximum likelihood for \((\kappa,\xi)\) [2510.02152].

Another multivariate formulation defines
\[
\bm{Y} = R \left( [1 - \omega\{F_R(R)\}] \bm{L} + \omega\{F_R(R)\} \bm{U} \right),
\]
where \(R\) has a univariate EGPD cdf, \(\bm{L}\) and \(\bm{U}\) are simplex-valued random vectors controlling lower-tail and upper-tail dependence, and \(\omega\) is a continuous cdf used as a dynamic weight [2509.05982]. This model is reported to have the following properties: its marginal distributions behave like univariate eGPDs; its lower and upper joint tails comply with multivariate extreme-value theory, with key parameters separately controlling dependence in each joint tail; and the model allows for fast simulation and is thus amenable to simulation-based inference [2509.05982]. Estimation is proposed through modern neural approaches, including neural Bayes estimators and neural posterior estimators, so that a neural network, once trained, can provide point estimates, credible intervals, or full posterior approximations in a fraction of a second [2509.05982].

A different line of development studies aggregated rainfall across scales from sub-hourly to weekly using EGPD [2601.08350]. The reported framework models all rainfall intensities across scales and establishes a general result on the behavior of EGPD variables under various aggregation procedures [2601.08350]. If \(Y \sim EGPD(\sigma, \kappa, \xi, B)\) and \(T\) is a sum or more generally a transformation of iid EGPDs, then under broad conditions
\[
T \sim EGPD(\sigma, \gamma \kappa, \xi, B_T),
\]
where \(\gamma\) is determined by the aggregation structure and \(B_T\) is a new bulk cdf [2601.08350]. To address the absence of a closed-form likelihood for aggregated rainfall, the paper links the EGPD class to Poisson compound sums and uses the Panjer algorithm for efficient composite likelihood evaluation [2601.08350]. The empirical claim reported there is that only eight parameters are needed per station to capture scales from six minutes to three days, and that the resulting return levels do not cross across scales [2601.08350].

Taken together, these multivariate and multiscale developments place EGPD within a broader class of models that seek coherence across the entire support, across dependence structures, and across aggregation levels, rather than only within upper-tail asymptotics.

## 7. Software, implementation, and modeling practice

The continuous distributional-regression paper provides an add-on script for the R package `gamlss`, available at `https://github.com/noemielc/egpd4gamlss`, to implement the EGPD family in a generic way [2209.04660]. The script can assemble new EGPD family distributions and supports density, cdf, quantile, and random generation, with symbolic differentiation via the `Deriv` R package for likelihood optimization [2209.04660]. The supplied examples include creation of an EGPD1 family and fitting baseline or time-varying models with `gamlss` [2209.04660].

For count data, the methodology is implemented via extensions to the `evgam` R package, and a repository is provided at `github.com/touqeerahmadunipd/degpd-and-zidegpd` [2210.15253]. In both settings, the estimation strategy is penalized maximum likelihood, with smoothing penalties on nonlinear additive predictors [2209.04660] [2210.15253].

Several recurring modeling strategies appear across the literature summarized here. First, EGPD is repeatedly used to bypass or relax threshold selection, either by replacing a threshold-based analysis entirely or by stabilizing threshold-based inference [1111.6899] [2209.04660] [2210.15253]. Second, the extra shape mechanism—variously denoted \(\kappa\), \(G\), \(\mathcal{G}\), or \(B\)—is used to control the lower tail and bulk while preserving upper-tail interpretation through \(\xi\) [2209.04660] [2409.18719] [2510.02152]. Third, empirical comparisons are commonly made against Gamma models for rainfall, and Poisson, Negative Binomial, or zero-inflated variants for counts [2209.04660] [2210.15253] [2409.18719].

A common misconception is that EGPD is merely a heavier-tailed GPD. The reported formulations do not support that characterization. Rather, they preserve the tail index interpretation of the GPD and introduce additional structure for the main body or lower tail, with the stated goal of modeling the entire positive support or stabilizing threshold analyses [1111.6899] [2209.04660]. Another potential misconception is that all EGPD formulations are identical. The data instead show multiple constructions: threshold-exceedance EGP1–EGP3 models, full-range positive-data EGPDs using \(G\) or \(B\), discrete DEGPD and ZIDEGPD analogues, and multivariate radial-angular or stochastic-mixture formulations [1111.6899] [2209.04660] [2210.15253] [2509.05982] [2510.02152].

The overall trajectory of the literature suggests a unifying theme rather than a single canonical parameterization. That theme is a tail-compliant extension of generalized Pareto structure to the full distribution, with separate mechanisms for lower-tail or bulk flexibility, covariate dependence, discreteness, multivariate dependence, or aggregation.

Source: https://www.emergentmind.com/topics/extended-generalized-pareto-distribution-egpd