Papers
Topics
Authors
Recent
Search
2000 character limit reached

Double Pareto Distribution Overview

Updated 14 July 2026
  • Double Pareto distribution is a two-sided power-law model defined by a log-Laplace transformation that yields distinct power-law behaviors near zero and at infinity.
  • It emerges from stochastic processes like geometric Brownian motion with random observation times, with parameters explicitly controlling tail indices and moment existence.
  • Extensions such as the Beta Rank Function (BRF) and double Pareto lognormal (dPlN) address modeling refinements by smoothing the central peak while preserving heavy-tail properties.

The double Pareto distribution is a two-sided power-law distribution on positive support. In its canonical form, it is obtained by exponentiating a skewed Laplace variable, so that Z=logXZ=\log X has a piecewise exponential density with a cusp at a central location μ\mu, while XX has Pareto-type behavior both near zero and in the upper tail. This construction makes the distribution simultaneously a log-Laplace law on the log scale and a piecewise power law on the original scale. The family appears in multiplicative-growth models with random observation time, in rank-size modeling, and as a baseline against which smoother or more flexible variants such as the Beta Rank Function (BRF) and the double Pareto lognormal (dPlN) are compared (Fontanelli et al., 2019, Yamamoto et al., 2024).

1. Canonical distribution on the log scale and the original scale

A standard parameterization starts from the skewed log-Laplace density for Z=logXZ=\log X: fZDP(z)=1a+b{exp((zμ)/b),zμ, exp((zμ)/a),zμ.f_Z^{DP}(z) = \frac{1}{a+b} \begin{cases} \exp((z-\mu)/b), & z \le \mu,\ \exp(-(z-\mu)/a), & z \ge \mu. \end{cases} It is continuous at μ\mu with a sharp angle, and the probability mass splits as b/(a+b)b/(a+b) to the left of μ\mu and a/(a+b)a/(a+b) to the right (Fontanelli et al., 2019).

Transforming back to X=eZX=e^Z with μ\mu0 gives the canonical double Pareto density

μ\mu1

and the cdf

μ\mu2

In this form, the left side is a power law near zero and the right side is a Pareto tail for large μ\mu3 (Fontanelli et al., 2019).

An equivalent parameterization, common in stochastic-process derivations, uses positive tail indices μ\mu4 and μ\mu5 with matching point μ\mu6: μ\mu7 The cdf is then

μ\mu8

Both branches meet continuously at μ\mu9, and the tail exponents are explicit in the density and survival asymptotics (Yamamoto et al., 29 Sep 2025).

The canonical model is therefore best viewed as a positive-support distribution whose defining regularity is exact piecewise power-law behavior with a non-smooth mode on the log scale. That cusp is central to later generalizations: some preserve the tails but smooth the peak, while others preserve the idea of two tail indices but replace the core altogether.

2. Stochastic-generation mechanisms

A particularly important derivation starts from geometric Brownian motion,

XX0

with XX1. If the observation time XX2 is independent and exponentially distributed, XX3, then the log-return

XX4

has mgf

XX5

Factoring the denominator yields

XX6

so XX7 is asymmetric Laplace on the log scale and XX8 is double Pareto on the original scale (Yamamoto et al., 2024).

This observation-time mechanism generalizes. For a nonnegative random time XX9 with mgf Z=logXZ=\log X0, the exact moment formula is

Z=logXZ=\log X1

whenever the mgf is finite at that argument. In this sense, moment existence for the observed GBM is reduced to the domain of Z=logXZ=\log X2, and the double Pareto case arises when the observation-time mgf has the exponential form Z=logXZ=\log X3 (Yamamoto et al., 29 Sep 2025).

The same framework clarifies when the double Pareto law ceases to hold. Replacing the exponential stopping time by a uniform observation time Z=logXZ=\log X4 produces a uniform mixture of lognormals rather than a log-Laplace mixture. In that model, the pdf and ccdf admit exact formulas in terms of Z=logXZ=\log X5, all polynomial moments are finite, and the tails are lognormal-like rather than power-law; the paper describes this as a deformation of the double Pareto distribution (Yamamoto et al., 2024).

A further generalization uses generalized inverse Gaussian observation times. In that setting the pdf of the observed GBM is obtained exactly in terms of modified Bessel functions, and the gamma, inverse gamma, and inverse Gaussian cases appear as specializations. The exponential-observation-time double Pareto law is recovered as a limiting case of the gamma subclass with shape Z=logXZ=\log X6 (Yamamoto et al., 29 Sep 2025).

3. Tail indices, moments, and parameter interpretation

The double Pareto distribution is characterized by two tail indices, one governing the lower tail near zero and the other governing the upper tail at infinity. In the Z=logXZ=\log X7 parameterization,

Z=logXZ=\log X8

so Z=logXZ=\log X9 controls the right tail and fZDP(z)=1a+b{exp((zμ)/b),zμ, exp((zμ)/a),zμ.f_Z^{DP}(z) = \frac{1}{a+b} \begin{cases} \exp((z-\mu)/b), & z \le \mu,\ \exp(-(z-\mu)/a), & z \ge \mu. \end{cases}0 the left tail (Yamamoto et al., 29 Sep 2025).

Moment existence is correspondingly asymmetric. For the exponential-observation-time double Pareto derived from GBM,

fZDP(z)=1a+b{exp((zμ)/b),zμ, exp((zμ)/a),zμ.f_Z^{DP}(z) = \frac{1}{a+b} \begin{cases} \exp((z-\mu)/b), & z \le \mu,\ \exp(-(z-\mu)/a), & z \ge \mu. \end{cases}1

and the moment is finite if and only if fZDP(z)=1a+b{exp((zμ)/b),zμ, exp((zμ)/a),zμ.f_Z^{DP}(z) = \frac{1}{a+b} \begin{cases} \exp((z-\mu)/b), & z \le \mu,\ \exp(-(z-\mu)/a), & z \ge \mu. \end{cases}2. The poles at fZDP(z)=1a+b{exp((zμ)/b),zμ, exp((zμ)/a),zμ.f_Z^{DP}(z) = \frac{1}{a+b} \begin{cases} \exp((z-\mu)/b), & z \le \mu,\ \exp(-(z-\mu)/a), & z \ge \mu. \end{cases}3 and fZDP(z)=1a+b{exp((zμ)/b),zμ, exp((zμ)/a),zμ.f_Z^{DP}(z) = \frac{1}{a+b} \begin{cases} \exp((z-\mu)/b), & z \le \mu,\ \exp(-(z-\mu)/a), & z \ge \mu. \end{cases}4 encode the tail indices directly (Yamamoto et al., 29 Sep 2025).

In the BRF-induced double-Pareto-like family, the analogous statement is expressed with parameters fZDP(z)=1a+b{exp((zμ)/b),zμ, exp((zμ)/a),zμ.f_Z^{DP}(z) = \frac{1}{a+b} \begin{cases} \exp((z-\mu)/b), & z \le \mu,\ \exp(-(z-\mu)/a), & z \ge \mu. \end{cases}5 and fZDP(z)=1a+b{exp((zμ)/b),zμ, exp((zμ)/a),zμ.f_Z^{DP}(z) = \frac{1}{a+b} \begin{cases} \exp((z-\mu)/b), & z \le \mu,\ \exp(-(z-\mu)/a), & z \ge \mu. \end{cases}6. The positive moments satisfy

fZDP(z)=1a+b{exp((zμ)/b),zμ, exp((zμ)/a),zμ.f_Z^{DP}(z) = \frac{1}{a+b} \begin{cases} \exp((z-\mu)/b), & z \le \mu,\ \exp(-(z-\mu)/a), & z \ge \mu. \end{cases}7

and are finite iff fZDP(z)=1a+b{exp((zμ)/b),zμ, exp((zμ)/a),zμ.f_Z^{DP}(z) = \frac{1}{a+b} \begin{cases} \exp((z-\mu)/b), & z \le \mu,\ \exp(-(z-\mu)/a), & z \ge \mu. \end{cases}8. Hence the mean exists if fZDP(z)=1a+b{exp((zμ)/b),zμ, exp((zμ)/a),zμ.f_Z^{DP}(z) = \frac{1}{a+b} \begin{cases} \exp((z-\mu)/b), & z \le \mu,\ \exp(-(z-\mu)/a), & z \ge \mu. \end{cases}9, the variance if μ\mu0, and the right tail controls the existence of moments, while the left tail near zero is benign for μ\mu1 (Fontanelli et al., 2019).

The parameter mappings in the GBM representation are explicit: μ\mu2 with inverse relations

μ\mu3

Larger observation-time scale implies heavier tails by reducing the effective tail indices, whereas bounded observation windows remove the power-law regime altogether (Yamamoto et al., 29 Sep 2025, Yamamoto et al., 2024).

A common misconception is that any distribution with two apparent heavy tails is necessarily canonical double Pareto. The supplied literature shows that this is not generally correct. A plausible implication is that the decisive diagnostic is not merely two-sided decay, but the combination of exact power-law tails, the behavior near the matching point, and the observation-time or mixture mechanism that generates the model.

4. Smooth and composite extensions

The BRF provides a smooth double-Pareto-like alternative derived from rank-size structure rather than a piecewise-defined density. It is defined by

μ\mu4

where μ\mu5 is normalized continuous rank. On the log scale,

μ\mu6

and

μ\mu7

Its tails are approximately exponential on the log scale,

μ\mu8

which implies the original-scale approximations

μ\mu9

Thus the BRF matches the canonical double Pareto in the tails but differs at the mode: its peak is smooth and has no cusp (Fontanelli et al., 2019).

Several closed-form features of the BRF on the log scale are exact. The mode is

b/(a+b)b/(a+b)0

and the probability split at the peak is

b/(a+b)b/(a+b)1

Near the peak,

b/(a+b)b/(a+b)2

with the cubic term vanishing when b/(a+b)b/(a+b)3. In the symmetric case b/(a+b)b/(a+b)4, the exact log-scale pdf is

b/(a+b)b/(a+b)5

which is quadratic near the mode and exponential in the tails; it is therefore log-normal-like near the peak but not identical to a lognormal (Fontanelli et al., 2019).

The dPlN distribution modifies the canonical model in a different way by introducing a lognormal core. If b/(a+b)b/(a+b)6 with b/(a+b)b/(a+b)7 and b/(a+b)b/(a+b)8 skewed Laplace, then b/(a+b)b/(a+b)9 is dPlN with pdf

μ\mu0

Its upper tail is Pareto-type with index μ\mu1, its lower tail is Pareto-type near zero with index μ\mu2, and moments exist for orders μ\mu3: μ\mu4 Unlike the canonical double Pareto, the dPlN does not possess a Laplace transform in closed form (Ramirez-Cobo et al., 2010).

These two extensions answer different modeling deficiencies. The BRF preserves the idea of two-sided power-law tails while smoothing the cusp intrinsically. The dPlN preserves Pareto tails but replaces the rigid center by a lognormal core. This suggests that empirical departures from the canonical distribution often concern the central body or modal regularity rather than the tail exponents alone.

5. Estimation, diagnostics, and computation

For BRF and double-Pareto-like data, a practical diagnostic pipeline works on μ\mu5. The recommended steps are to plot a histogram with y-axis in log scale and inspect the tails: linear decay only on the right suggests a one-sided power law; linear decay on both sides with similar slopes suggests an approximately log-normal pattern; linear decay on both sides with different slopes suggests BRF or double Pareto-like behavior; and poor data on one side should not be over-interpreted (Fontanelli et al., 2019).

Tail parameters can be estimated directly from the semi-log histogram of μ\mu6, where the left slope is approximately μ\mu7 and the right slope approximately μ\mu8. BRF parameters can also be estimated by rank-size regression: μ\mu9 or, with discrete ranks,

a/(a+b)a/(a+b)0

The paper further proposes linking a/(a+b)a/(a+b)1, a/(a+b)a/(a+b)2, and a/(a+b)a/(a+b)3 to sample moments through

a/(a+b)a/(a+b)4

and suggests Jackknife resampling to reduce estimator bias (Fontanelli et al., 2019).

For random-time GBM models, the exact pdfs in the exponential, gamma, inverse Gaussian, and GIG cases support maximum-likelihood estimation by direct numerical optimization. The exact moment formula

a/(a+b)a/(a+b)5

also yields a method-of-moments route, while tail-slope estimation can be followed by back-mapping from the estimated tail indices to a/(a+b)a/(a+b)6 or a/(a+b)a/(a+b)7 (Yamamoto et al., 29 Sep 2025).

For dPlN, Bayesian inference is developed on the log scale using the Normal-Laplace decomposition a/(a+b)a/(a+b)8. The Gibbs sampler augments the model with latent Gaussian and exponential variables, updates a/(a+b)a/(a+b)9, X=eZX=e^Z0, X=eZX=e^Z1, and X=eZX=e^Z2 from their conditional posteriors, and requires proper priors for X=eZX=e^Z3 and X=eZX=e^Z4, since improper priors yield improper posteriors. Because the dPlN Laplace transform is not available in closed form, queueing calculations use the Transform Approximation Method (TAM), based on quantiles and weighted exponentials (Ramirez-Cobo et al., 2010).

The estimation literature therefore reflects the structural diversity of the topic: piecewise power laws admit direct tail fitting, rank-derived smooth variants admit regression on rank-size equations, and composite core-tail models often require latent-variable Bayesian computation.

6. Applications, adjacent uses, and conceptual distinctions

The double Pareto viewpoint has been applied to urban populations and financial log-returns. In the BRF study, India 2011 cities showed one-sided power-law structure, whereas China 2010 urban units showed two-sided power-law structure with different left and right tail slopes. Daily log-returns for four US indices over 30 years were described as having tails heavier than normal, with log-BRF fitting the near-peak shape comparably while matching the heavier tails through the exponential decay rates X=eZX=e^Z5 and X=eZX=e^Z6 on the log scale (Fontanelli et al., 2019).

The dPlN extension has been used in internet traffic analysis and insurance. For 50,000 interarrival times from BC-pAug89 and for Bellcore bytes/sec data, the dPlN was reported to fit both body and tails better than a single Pareto, while in insurance portfolios the model accommodated very heavy right tails. Queueing applications include X=eZX=e^Z7 and X=eZX=e^Z8 systems, where stability and waiting-time summaries depend on whether moments such as X=eZX=e^Z9 and μ\mu00 exist; in heavy-tailed settings, the paper recommends medians and quantiles rather than means for equilibrium summaries (Ramirez-Cobo et al., 2010).

A separate line of work uses the term “generalized double Pareto” for a Bayesian shrinkage prior in linear models. That prior is not the classic positive-support double Pareto used in economics or multiplicative-growth modeling. It is symmetric about zero,

μ\mu01

has a finite mass at zero μ\mu02, and has polynomial tails of order μ\mu03. The paper emphasizes the contrast explicitly: the classic “double Pareto” typically models strictly positive quantities and uses different parameterizations and support, whereas the GDP prior is designed for shrinkage in regression and is constructed by reflecting a one-sided generalized Pareto density about zero (Armagan et al., 2011).

This distinction is important because the phrase “double Pareto distribution” refers to a family resemblance rather than a single universal object. On positive support, the canonical model is log-Laplace and piecewise power law. In smooth rank-based variants, the tails remain double-Pareto-like but the cusp disappears. In dPlN models, a lognormal core is inserted between the tails. In shrinkage priors, symmetry about zero and sparsity at the origin become the dominant design features. The shared motif is two-sided polynomial behavior, but the support, generative mechanism, and inferential role differ substantially across these constructions.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Double Pareto Distribution.