Papers
Topics
Authors
Recent
Search
2000 character limit reached

Poisson Hierarchical Indian Buffet Process

Updated 9 July 2026
  • PHIBP is a Bayesian nonparametric latent feature model that uses Poisson restrictions to control feature counts in hierarchical, group-specific settings.
  • It combines restricted IBP and hierarchical CRM formulations to generate sparse, count-valued data representations while sharing global atoms.
  • Computational methods such as exact compound-Poisson factorization and collapsed samplers improve inference for high-dimensional count data.

Searching arXiv for recent and foundational papers on PHIBP and closely related formulations. arXiv query: "Poisson Hierarchical Indian Buffet Process" The Poisson Hierarchical Indian Buffet Process (PHIBP) denotes a class of hierarchical Bayesian nonparametric latent feature models in which Poisson structure is introduced into Indian-buffet-style feature allocation. In one construction, PHIBP is obtained from the Restricted Indian Buffet Process by choosing group-specific restricting distributions fg=Poisson(λg)f_g=\mathrm{Poisson}(\lambda_g), so that the number of active features in each observation is Poisson within group while feature-sharing remains global through a common directing measure μ\mu (Doshi-Velez et al., 2015). In another construction, PHIBP refers to a hierarchical completely random measure (CRM) model for sparse multivariate count data, with a global discrete CRM B0B_0, group-specific CRMs BjB_j, and Poisson point-process observations (James et al., 4 Feb 2025). This suggests that the literature uses the name for a closely related family of Poisson-hierarchical extensions of Indian buffet ideas rather than for a single universally fixed specification.

1. Genealogy within Indian buffet research

A central starting point is the standard Indian Buffet Process (IBP), which can be written as a beta–Bernoulli process. In the formulation

μ=iπiδθiBP(c,α,H),Zn=izniδθii.i.d.BeP(μ),\mu=\sum_i \pi_i \delta_{\theta_i} \sim BP(c,\alpha,H), \qquad Z_n=\sum_i z_{ni}\delta_{\theta_i} \stackrel{i.i.d.}{\sim} BeP(\mu),

integrating out μ\mu yields an exchangeable binary matrix, and the number of active features in a row,

mn=k=1znk,m_n=\sum_{k=1}^{\infty} z_{nk},

is marginally Poisson(α)\mathrm{Poisson}(\alpha) (Doshi-Velez et al., 2015). The Poisson row-count is therefore not an added option in the ordinary IBP; it is an implicit consequence of the CRM construction.

A second precursor is the Hierarchical Indian Buffet Process (HIBP), which extends latent feature allocation across groups. In the Bernoulli case one may write

Zj(1),,Zj(Mj)μj iidBeP(μj),μjB0BP(θj,B0),B0BP(θ0,G0),Z^{(1)}_{j}, \dots, Z^{(M_j)}_{j}\mid \mu_j \ \text{iid} \sim BeP(\mu_j),\qquad \mu_j\mid B_0 \sim BP(\theta_j,B_0),\qquad B_0\sim BP(\theta_0,G_0),

and more generally replace the beta process by a general CRM and the Bernoulli slab by a spike-and-slab distribution (James et al., 2021). The Poisson specialization appears when

AjsPoisson(rjs),A_j\mid s \sim \mathrm{Poisson}(r_j s),

yielding count-valued feature allocations rather than binary ones (James et al., 2021, James et al., 2023).

These two lines of development motivate two distinct PHIBP interpretations. One preserves binary latent features but makes the row-count law explicitly Poisson and group-dependent through restriction (Doshi-Velez et al., 2015). The other treats observations themselves as count-valued point measures generated from hierarchical CRMs and Poisson likelihoods (James et al., 4 Feb 2025, James et al., 2023).

2. PHIBP as a Poisson-restricted hierarchical IBP

The Restricted Indian Buffet Process (R-IBP) was introduced to decouple the row-count distribution from the IBP’s implicit μ\mu0 law. Instead of drawing each row from a Bernoulli process, one uses a restricted Bernoulli process that conditions on the row sum. For a fixed row count μ\mu1,

μ\mu2

with normalization given by the Poisson–binomial probability of exactly μ\mu3 successes (Doshi-Velez et al., 2015).

In this framework, PHIBP is the partially exchangeable specialization obtained by assigning group-specific Poisson restrictions: μ\mu4 Equivalently, one may specify group-specific restricting distributions μ\mu5 directly and choose μ\mu6 (Doshi-Velez et al., 2015). Global feature-sharing is retained through the common directing measure μ\mu7, whose atoms and weights μ\mu8 are shared across all rows, while the number of active features per observation varies by group through μ\mu9.

This construction preserves exchangeability only in the appropriate sense. Rows are exchangeable in the fully exchangeable model with common B0B_00, and with group-specific B0B_01 the matrix is partially exchangeable, invariant to permutations within groups (Doshi-Velez et al., 2015). At the same time, the restriction breaks complete randomness, because conditioning on the row sum induces dependence among entries in the same row. This is precisely what permits nonstandard row-count laws. The paper also emphasizes a caveat: although inference for B0B_02 and B0B_03 under restrictions is developed, it does not specify conjugate updates for the group-level B0B_04 (Doshi-Velez et al., 2015).

A notable point is that choosing B0B_05 does not recover the ordinary IBP unless B0B_06 is embedded in the original completely random construction. Setting B0B_07 matches the IBP’s marginal Poisson rate, but the resulting process is not the standard IBP because the restriction removes complete randomness while keeping exchangeable rows (Doshi-Velez et al., 2015).

3. PHIBP as a hierarchical CRM model for count-valued features

A different and later usage of PHIBP treats the model as a hierarchical CRM construction for sparse multivariate count data. In this formulation, the global feature pool is a discrete CRM

B0B_08

where B0B_09 are feature labels and BjB_j0 are global mean abundance rates. For each group BjB_j1,

BjB_j2

so each observation is a Poisson point process with mean intensity measure BjB_j3 (James et al., 4 Feb 2025).

Conditioning on

BjB_j4

the group-specific CRM can be written

BjB_j5

where BjB_j6 is the group-specific mean abundance rate of feature BjB_j7 in group BjB_j8. Sparsity arises through Poisson thinning, and over-dispersion arises naturally from CRM mixing (James et al., 4 Feb 2025).

The species-allocation representation is particularly explicit. The number of observed species is

BjB_j9

and for each observed species μ=iπiδθiBP(c,α,H),Zn=izniδθii.i.d.BeP(μ),\mu=\sum_i \pi_i \delta_{\theta_i} \sim BP(c,\alpha,H), \qquad Z_n=\sum_i z_{ni}\delta_{\theta_i} \stackrel{i.i.d.}{\sim} BeP(\mu),0, the total number of latent OTU clusters across groups satisfies

μ=iπiδθiBP(c,α,H),Zn=izniδθii.i.d.BeP(μ),\mu=\sum_i \pi_i \delta_{\theta_i} \sim BP(c,\alpha,H), \qquad Z_n=\sum_i z_{ni}\delta_{\theta_i} \stackrel{i.i.d.}{\sim} BeP(\mu),1

while group allocations are multinomial with probabilities

μ=iπiδθiBP(c,α,H),Zn=izniδθii.i.d.BeP(μ),\mu=\sum_i \pi_i \delta_{\theta_i} \sim BP(c,\alpha,H), \qquad Z_n=\sum_i z_{ni}\delta_{\theta_i} \stackrel{i.i.d.}{\sim} BeP(\mu),2

Within each group and species, latent jump sizes are drawn from

μ=iπiδθiBP(c,α,H),Zn=izniδθii.i.d.BeP(μ),\mu=\sum_i \pi_i \delta_{\theta_i} \sim BP(c,\alpha,H), \qquad Z_n=\sum_i z_{ni}\delta_{\theta_i} \stackrel{i.i.d.}{\sim} BeP(\mu),3

and their total counts are zero-truncated Poisson (James et al., 4 Feb 2025).

This count-valued PHIBP is directly connected to the generalized hierarchical spike-and-slab IBP. In that broader framework, taking

μ=iπiδθiBP(c,α,H),Zn=izniδθii.i.d.BeP(μ),\mu=\sum_i \pi_i \delta_{\theta_i} \sim BP(c,\alpha,H), \qquad Z_n=\sum_i z_{ni}\delta_{\theta_i} \stackrel{i.i.d.}{\sim} BeP(\mu),4

produces a Poisson HIBP with mixed-Poisson and compound-Poisson structure, and gamma-process choices yield Gamma–Negative Binomial random count matrices and Gamma–Poisson or negative-binomial Poisson factorization models (James et al., 2021, James et al., 2023).

4. Posterior structure and computational methods

The R-IBP formulation and the count-valued CRM formulation lead to different inference machinery.

For the R-IBP, finite truncations yield normalization constants

μ=iπiδθiBP(c,α,H),Zn=izniδθii.i.d.BeP(μ),\mu=\sum_i \pi_i \delta_{\theta_i} \sim BP(c,\alpha,H), \qquad Z_n=\sum_i z_{ni}\delta_{\theta_i} \stackrel{i.i.d.}{\sim} BeP(\mu),5

which satisfy the dynamic programming recursion

μ=iπiδθiBP(c,α,H),Zn=izniδθii.i.d.BeP(μ),\mu=\sum_i \pi_i \delta_{\theta_i} \sim BP(c,\alpha,H), \qquad Z_n=\sum_i z_{ni}\delta_{\theta_i} \stackrel{i.i.d.}{\sim} BeP(\mu),6

This gives an μ=iπiδθiBP(c,α,H),Zn=izniδθii.i.d.BeP(μ),\mu=\sum_i \pi_i \delta_{\theta_i} \sim BP(c,\alpha,H), \qquad Z_n=\sum_i z_{ni}\delta_{\theta_i} \stackrel{i.i.d.}{\sim} BeP(\mu),7 computation of the row-sum normalizers, while cached inclusion probabilities for draw-by-draw sampling require μ=iπiδθiBP(c,α,H),Zn=izniδθii.i.d.BeP(μ),\mu=\sum_i \pi_i \delta_{\theta_i} \sim BP(c,\alpha,H), \qquad Z_n=\sum_i z_{ni}\delta_{\theta_i} \stackrel{i.i.d.}{\sim} BeP(\mu),8 time (Doshi-Velez et al., 2015). The paper develops a collapsed sampler via IBP subset selection with auxiliary discarded rows, uncollapsed samplers based on weak-limit beta or truncated stick-breaking approximations, exact retrospective sampling for degenerate restrictions, Metropolis–Hastings updates for μ=iπiδθiBP(c,α,H),Zn=izniδθii.i.d.BeP(μ),\mu=\sum_i \pi_i \delta_{\theta_i} \sim BP(c,\alpha,H), \qquad Z_n=\sum_i z_{ni}\delta_{\theta_i} \stackrel{i.i.d.}{\sim} BeP(\mu),9, and a hybrid variational method for linear-Gaussian likelihoods that alternates coordinate-ascent updates for μ\mu0, μ\mu1, and μ\mu2 with Metropolis–Hastings resampling of μ\mu3 (Doshi-Velez et al., 2015).

For the count-valued PHIBP, the key computational device is an exact compound-Poisson factorization. Exact marginal simulation first samples the number of observed species μ\mu4, then the observed parent jumps μ\mu5, then the across-group allocations μ\mu6, and finally the within-group latent counts μ\mu7 and their customer-level multinomial allocations (James et al., 4 Feb 2025). Posterior analysis has a similarly explicit decomposition: μ\mu8 and

μ\mu9

Within each observed species and group, the latent OTU partition conditional on the total count mn=k=1znk,m_n=\sum_{k=1}^{\infty} z_{nk},0 follows a finite Gibbs EPPF with block-size weights mn=k=1znk,m_n=\sum_{k=1}^{\infty} z_{nk},1 (James et al., 4 Feb 2025).

Special cases substantially simplify computation. Under gamma-process priors,

mn=k=1znk,m_n=\sum_{k=1}^{\infty} z_{nk},2

posterior species-level rates become gamma, latent OTU jumps become gamma, and normalized posterior species probabilities are Dirichlet (James et al., 4 Feb 2025). Under generalized gamma processes,

mn=k=1znk,m_n=\sum_{k=1}^{\infty} z_{nk},3

the residual unseen component is generalized gamma, observed OTU jumps are gamma with shape shifted by mn=k=1znk,m_n=\sum_{k=1}^{\infty} z_{nk},4, and normalized posterior rates are no longer Dirichlet but involve sums of mn=k=1znk,m_n=\sum_{k=1}^{\infty} z_{nk},5 and gamma variables (James et al., 4 Feb 2025).

5. Structural properties and relation to adjacent models

Several structural features recur across PHIBP formulations. In the R-IBP lineage, the decisive operation is the explicit decoupling of the row-count distribution from the CRM-induced mn=k=1znk,m_n=\sum_{k=1}^{\infty} z_{nk},6 law. This preserves de Finetti exchangeability but breaks complete randomness, and it makes the total number of active features per observation directly modelable through mn=k=1znk,m_n=\sum_{k=1}^{\infty} z_{nk},7 or mn=k=1znk,m_n=\sum_{k=1}^{\infty} z_{nk},8 (Doshi-Velez et al., 2015). In the count-valued lineage, the decisive operation is the use of hierarchical CRMs and Poisson point processes, so that sharing occurs through common global atoms and group-specific abundance transformations (James et al., 4 Feb 2025).

The term is therefore not synonymous with the ordinary HIBP. The Bernoulli HIBP produces binary random matrices through beta–Bernoulli machinery, whereas Poisson HIBP and PHIBP replace Bernoulli observations by count-valued Poisson slabs or Poisson point processes (James et al., 2021, James et al., 2023). The distinction is substantive rather than cosmetic: the Poisson versions correspond to random multisets or sparse count matrices, admit mixed-Poisson and compound-Poisson representations, and connect naturally to random count matrix priors and Poisson factorization (James et al., 2021).

The model family also differs from hierarchical clustering priors such as HDP, HPY, or Poisson–Kingman models. HIBP is the latent feature analogue of HDP in the sense that it supports sharing across groups, but PHIBP targets latent features rather than partitions (James et al., 2021). In the 2025 duality formulation, PHIBP sits in the CRM family and produces Poisson–Kingman laws and Gibbs-type EPPFs, while remaining distinct from pure PD/Poisson–Kingman special cases or HPY; stable-subordinator specializations recover PDmn=k=1znk,m_n=\sum_{k=1}^{\infty} z_{nk},9 mechanisms and Pitman duality (James, 26 Aug 2025).

A plausible implication is that “PHIBP” should be read as a modeling pattern: shared global atoms, group-specific Poisson structure, and unbounded latent dimensionality. What changes across papers is whether the Poisson aspect governs row counts, feature counts, within-feature multiplicities, or all three.

6. Duality theory, applications, and later developments

PHIBP acquired a broader mathematical role in work on coagulation–fragmentation duality. A 2025 paper builds a four-component coupled process upon PHIBP for microbiome species sampling across multiple groups. The construction simultaneously defines the fine-grained partition, its coagulation operator, a forward-in-time system of coupled time-homogeneous fragmentation processes, and a dual backward-in-time structured coalescent, all governed by the same underlying compositional structure and admitting exact compound-Poisson representations (James, 26 Aug 2025).

At fixed Poisson(α)\mathrm{Poisson}(\alpha)0, the coarse and fine partition laws are linked by the Poissonized duality identity

Poisson(α)\mathrm{Poisson}(\alpha)1

Marginalizing over Poisson(α)\mathrm{Poisson}(\alpha)2 yields a joint EPPF duality for nested partitions, and the result holds for arbitrary subordinators and for any Poisson(α)\mathrm{Poisson}(\alpha)3, thereby generalizing Pitman’s duality beyond the stable setting and into the multi-group regime (James, 26 Aug 2025).

Application domains reflect the same count-valued, shared-feature emphasis. The 2023 generalized HIBP analysis states that Poisson HIBP corresponds to generalizations of mixed Poisson random count models arising in genetics, imaging, topic modeling, random occupancy, and species sampling models (James et al., 2023). The 2025 PHIBP analysis develops the model for microbiome species sampling, emphasizing global mean abundance rates, group-specific mean abundances, unseen species prediction, and diversity indices derived from posterior mean rates (James et al., 4 Feb 2025). A later chapter applies PHIBP to infectious disease prediction in data-sparse environments, using a hierarchical CRM model with exposure offsets to borrow strength across regions that share diseases through the global pool (Fong et al., 24 Dec 2025).

Across these uses, the unifying theme is not merely “Poisson observations,” but hierarchical sharing of an unbounded feature set under explicit Poisson or mixed-Poisson laws. In binary latent feature models, the main question is whether a feature is present. In PHIBP, the main question is often how much of a globally shared feature is expressed in each group, observation, or species block, and how that expression should be regularized by hierarchical nonparametric structure.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Poisson Hierarchical Indian Buffet Process (PHIBP).