Papers
Topics
Authors
Recent
Search
2000 character limit reached

Ewens–Pitman Partitions

Updated 30 November 2025
  • Ewens–Pitman partitions are a two-parameter family of exchangeable random partitions that generalize the classical Ewens sampling formula and the Pitman–Yor process.
  • They are constructed via predictive models like the Chinese Restaurant Process and stick-breaking representations, providing clear probabilistic and combinatorial insights.
  • These partitions underpin practical applications in population genetics, Bayesian nonparametrics, and machine learning through their well-defined asymptotic and large deviation behaviors.

Ewens–Pitman partitions constitute a two-parameter family of exchangeable random partitions over [n]={1,,n}[n] = \{1,\dots, n\}, determined by (α,θ)(\alpha,\theta) with either 0α<10 \leq \alpha < 1 and θ>α\theta > -\alpha, or α<0\alpha < 0 and θ=mα\theta = -m\alpha for mNm \in \mathbb{N}. They interpolate between the classical Ewens sampling formula (α=0\alpha=0) and the two-parameter Poisson–Dirichlet (Pitman–Yor) distribution (0<α<10 < \alpha < 1), and admit deep connections with Gibbs partitions, stable subordinators, generalized Stirling numbers, compound Poisson representations, and the combinatorics of symmetric groups. Rich probabilistic, asymptotic, and algebraic structures underlie these partitions, yielding both practical statistical tools and theoretical insight into fragmentation, random trees, and Bayesian nonparametrics.

1. Formal Definition and Exchangeable Partition Probability Function

A random partition of [n][n] with block sizes (α,θ)(\alpha,\theta)0 ((α,θ)(\alpha,\theta)1) is assigned the Ewens–Pitman probability

(α,θ)(\alpha,\theta)2

where (α,θ)(\alpha,\theta)3 is the Pochhammer symbol, and (α,θ)(\alpha,\theta)4 encodes multiplicative block-size weights (Greve, 6 Mar 2025, Dolera et al., 2021).

In combinatorial terms, the total probability of an unordered set partition with sizes (α,θ)(\alpha,\theta)5 is

(α,θ)(\alpha,\theta)6

The parameters must satisfy (α,θ)(\alpha,\theta)7, (α,θ)(\alpha,\theta)8 or (α,θ)(\alpha,\theta)9, 0α<10 \leq \alpha < 10.

This generalizes the Ewens sampling formula (0α<10 \leq \alpha < 11), for which

0α<10 \leq \alpha < 12

and the Poisson–Dirichlet process for 0α<10 \leq \alpha < 13.

2. Probabilistic Constructions and Predictive Structure

Ewens–Pitman partitions are equivalently described via the Chinese Restaurant Process (CRP). Given a partial partition of 0α<10 \leq \alpha < 14 into 0α<10 \leq \alpha < 15 blocks with sizes 0α<10 \leq \alpha < 16, the (n+1)-st element joins block 0α<10 \leq \alpha < 17 with probability

0α<10 \leq \alpha < 18

or initiates a new block with probability

0α<10 \leq \alpha < 19

This sequential construction yields exchangeable distributions on partitions and underpins their representation as de Finetti mixtures over random discrete measures, notably the Pitman–Yor process (Greve, 6 Mar 2025, Dolera et al., 2021, Favaro et al., 2014).

In the stick-breaking representation, for θ>α\theta > -\alpha0, θ>α\theta > -\alpha1, the mass partition has weights

θ>α\theta > -\alpha2

and drawing i.i.d. samples from the resulting random measure induces the Ewens–Pitman random partition (Favaro et al., 2016, Ho et al., 2018).

3. Compound Poisson Interpretations

Ewens–Pitman partitions admit an interpretation as mixtures of compound Poisson sampling models (Dolera et al., 2021):

  • For θ>α\theta > -\alpha3 (Ewens), block counts correspond to conditioning the total size of a log-series compound Poisson sample (LS-CPSM).
  • For general θ>α\theta > -\alpha4, block counts arise as mixtures (in θ>α\theta > -\alpha5) over negative-Binomial compound Poisson samples (NB-CPSM), with the mixing variable θ>α\theta > -\alpha6 a product of a Gamma and a scaled Mittag–Leffler (generalized stable) variable.

Specifically, setting θ>α\theta > -\alpha7, with θ>α\theta > -\alpha8 independent Gamma and θ>α\theta > -\alpha9 a random variable with density α<0\alpha < 00 (where α<0\alpha < 01 is the positive α<0\alpha < 02-stable density), the EPα<0\alpha < 03 partition law coincides with the NB-CPSMα<0\alpha < 04 marginal, and the number of blocks α<0\alpha < 05 concentrates (almost surely) as α<0\alpha < 06 (Dolera et al., 2021).

This compound Poisson approach seamlessly yields asymptotic results, closed-form formulas, and generalizations to Poisson–Kingman partitions.

4. Asymptotics: Laws of Large Numbers, Fluctuations, and Limit Theorems

For fixed α<0\alpha < 07, the key scaling regimes are as follows (Contardi et al., 2024, Bercu et al., 2024, Tsukuda, 2020):

  • For α<0\alpha < 08 (Ewens): α<0\alpha < 09 with Gaussian central limit fluctuations.
  • For θ=mα\theta = -m\alpha0: θ=mα\theta = -m\alpha1 almost surely, where θ=mα\theta = -m\alpha2 is θ=mα\theta = -m\alpha3-Mittag–Leffler distributed; θ=mα\theta = -m\alpha4.
  • CLT: θ=mα\theta = -m\alpha5.
  • LIL: Law of the iterated logarithm applies to the centered, scaled block counts (Bercu et al., 2024).
  • Higher moments: θ=mα\theta = -m\alpha6 (Tsukuda, 2020).

For microclustering applications and scalable settings, scaling θ=mα\theta = -m\alpha7 linearly with θ=mα\theta = -m\alpha8 (i.e., θ=mα\theta = -m\alpha9) yields a "microclustering" regime where the number of blocks and the counts of blocks of any fixed size both grow linearly with mNm \in \mathbb{N}0 while the maximal cluster size remains mNm \in \mathbb{N}1 (Beraha et al., 24 Jul 2025, Contardi et al., 2024).

5. Large Deviations, Moderate Deviations, and Concentration

mNm \in \mathbb{N}3

where mNm \in \mathbb{N}4 is a logarithmic transform involving the Mittag–Leffler function (Bercu et al., 9 Mar 2025, Favaro et al., 2014). An explicit sharp concentration inequality describes the probability of mNm \in \mathbb{N}5 deviating from its mean.

  • Moderate deviations: Intermediate scaling regimes, for sequences mNm \in \mathbb{N}6 with mNm \in \mathbb{N}7, yield corresponding rate functions mNm \in \mathbb{N}8 providing precise transition descriptors between CLT and LDP scales (Favaro et al., 2016).
  • Block frequencies: Analogous large and moderate deviation principles hold for counts mNm \in \mathbb{N}9 of blocks of fixed size α=0\alpha=00 (Favaro et al., 2014, Favaro et al., 2016).
  • Conditional LDP/MDP: Conditioning on partially observed partitions, the deviation rate functions remain unchanged – the initial sample's impact is negligible at large α=0\alpha=01 or sample-augmentation settings (Favaro et al., 2014, Favaro et al., 2016).

6. Representation Theory and Algebraic Structures

Ewens–Pitman partitions are characterized as non-extreme harmonic functions on the Kingman branching graph (infinite Young lattice) and are tightly linked to the combinatorics of symmetric group characters and interpolation polynomials (Greve, 6 Mar 2025). The partition probabilities admit explicit expansion in terms of Sheffer polynomial sequences and Riordan array sums, yielding effective computational methods for summary statistics, moments, and marginals. For example, the marginal probability of α=0\alpha=02 or joint factorial moments of block counts can be written as closed-form coefficients in generalized Stirling number expansions obtainable via generating function and Riordan array technology.

This algebraic approach both encapsulates the full system of sampling-consistent marginals and facilitates symbolic computations (Greve, 6 Mar 2025).

7. Applications, Biological and Statistical Significance

  • Population genetics: Ewens–Pitman partitions generalize the Ewens sampling formula (ESF) for modeling allelic diversity and mutation structures in finite populations (Giordano et al., 2019).
  • Bayesian nonparametrics: The Pitman–Yor process induced partitions serve as priors for clustering in Dirichlet and stable process mixture models—central in Bayesian statistics and machine learning.
  • Entity resolution: Microclustering variants (scaling α=0\alpha=03 with α=0\alpha=04) underpin scalable clustering and de-duplication/identity resolution with provable guarantees on block size and count growth rates (Beraha et al., 24 Jul 2025).
  • Species sampling and discovery probabilities: Tail asymptotics and conditional LDPs enable calculation of discovery probabilities, facilitating design and inference in ecological, genomic, and risk-assessment contexts (Favaro et al., 2014).
  • Random trees and fragmentation: Fragmentation and coagulation operations on Ewens–Pitman partitions generate Markov chains and random trees (e.g., continuum random trees), with the scaled block-size limits governed by Mittag–Leffler and stable laws (Ho et al., 2018, Mano, 2013).

8. Summary Table: Core Properties of Ewens–Pitman Partitions

Property α=0\alpha=05 (Ewens) α=0\alpha=06 (Pitman–Yor)
Block count growth α=0\alpha=07 α=0\alpha=08 a.s.
Block size distribution weak Dirichlet/multinomial Power law; Sibuya law
Compound Poisson repr. Log-series mixing (LS-CPSM) NB-CPSM, mixed by ML law
Large deviation rate Explicit, convex analytic Given by α=0\alpha=09
Integrable structure Stirling/Riordan (binomial) Generalized Stirling, Riordan
Microclustering regime 0<α<10 < \alpha < 10 iff 0<α<10 < \alpha < 11 0<α<10 < \alpha < 12 iff 0<α<10 < \alpha < 13

These properties summarize both the classical and non-standard regimes and their implications for stochastic modeling and asymptotic analysis.


References:

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Ewens–Pitman Partitions.