---
title: Poissonization in Probability and Geometry
url: https://www.emergentmind.com/topics/poissonization
type: topic
---

# Poissonization in Probability and Geometry

Searching arXiv for recent and foundational uses of Poissonization across probability, statistics, geometry, and operator algebras.
arxiv_search(query="Poissonization", max_results=10, sort_by="submittedDate")

Poissonization is a context-dependent transformation that replaces a fixed counting mechanism by a Poisson one, or replaces a non-Poisson structure by an associated Poisson structure. In probability, statistics, combinatorics, and algorithms, it typically means randomizing a deterministic sample size, number of updates, or number of jumps by a Poisson random variable or Poisson process, often to obtain independence, tractable generating functions, or continuous-time embeddings [1109.2990, 1802.06638, 1407.0505]. In Jacobi and contact geometry it denotes the passage from a Jacobi manifold to a Poisson manifold on a one-dimensional extension [1111.4705, 1801.06681], while in nonholonomic dynamics it refers to embedding a system in a larger space and reparametrizing time so that a Poisson operator emerges [1806.05003]. In operator-algebraic settings, Poissonization is a noncommutative generalization of the Poisson random measure and a functor from von Neumann algebras with weights to von Neumann algebras with states [2303.14580]. The common theme is structural simplification by adjoining a Poisson clock, Poisson random measure, or Poisson geometric lift.

## 1. Definitions and recurring mechanism

In its classical probabilistic form, Poissonization randomizes a fixed sample size. In the urn-model setting, the number of draws is taken to be \(N \sim \mathrm{Poisson}(\lambda)\) instead of a fixed \(n\), so that for class \(i\) with frequency \(\pi_i\), the probability of never being sampled is \(e^{-\lambda \pi_i}\), and the expected unseen mass is
\[
\mathbb{E}[\text{fraction unobserved}] = \sum_{i=1}^{K} \pi_i e^{-\lambda \pi_i}.
\]
This is the basis of conditionally unbiased predictors and exact prediction intervals for the fraction of the environment in unsampled classes in an urn model [1109.2990].

A closely related formulation appears in rare-event approximation. For independent observations with distributions \(F_i\), one replaces each deterministic single observation by a Poisson(\(1\)) number of i.i.d. copies, obtaining the accompanying compound Poisson law
\[
H_2 = \bigast_{i=1}^{n} e(F_i), \qquad
e(F_i)=e^{-1}\sum_{m=0}^{\infty}\frac{F_i^{*m}}{m!},
\]
which is compared to the original law
\[
H_1 = \bigast_{i=1}^{n} F_i.
\]
Under \(f(x)=0\) on the common part of the sample, the paper gives the bound
\[
p(H_1,H_2)\le c(d)\,p,
\]
with \(p=\max_i p_i\) the rare-event probability [1802.06638].

A third standard form is continuous-time embedding. The continuous-time simple symmetric random walk is defined as a Poissonization of the discrete-time random walk: jumps occur at exponential waiting times of rate \(1\), and the characteristic function becomes
\[
\mathbb{E}[e^{izV(t)}]=\exp\{t(\cos z-1)\}.
\]
Its transition probability is
\[
p(t,y\mid x)=e^{-t}I_{|y-x|}(t),
\]
with \(I_\nu\) the modified Bessel function [1407.0505].

These constructions suggest that Poissonization is not a single operation but a family of transformations whose shared purpose is to trade rigid counting for Poisson structure. The same term is therefore used both for probabilistic randomization and for geometric or operator-algebraic lifts.

## 2. Statistical asymptotics and inference

In nonparametric testing, Poissonization is used to analyze one-sided \(L_p\)-type functionals of kernel estimators under inequality constraints. For regression functions \(m_j(x)=\mathbb{E}[Y_{ji}\mid X_i=x]\), the hypothesis is
\[
H_0: m_j(x)\le 0 \text{ for all } (x,j)
\quad\text{vs.}\quad
H_1: m_j(x)>0 \text{ for some } (x,j),
\]
and the functional is
\[
\Gamma_j(\varphi)=\int_{\mathcal X}\Lambda_p(\varphi(x))w_j(x)\,dx,
\qquad
\Lambda_p(v)=\max\{v,0\}^p.
\]
With
\[
\hat g_{jn}(x)=\frac{1}{nh^d}\sum_{i=1}^n Y_{ji}K\!\left(\frac{x-X_i}{h}\right),
\]
the studentized statistic is
\[
T_n \equiv \frac{1}{\sigma_n}\sum_{j=1}^J
\left\{n^{p/2}h^{(p-1)d/2}\Gamma_j(\hat g_{jn})-a_{jn}\right\}.
\]
The Poissonized estimator replaces \(n\) by an independent \(N\sim \mathrm{Poisson}(n)\),
\[
\hat g_{jN}(x)=\frac{1}{nh^d}\sum_{i=1}^{N}Y_{ji}K\!\left(\frac{x-X_i}{h}\right),
\]
and the resulting process can be partitioned into nearly independent blocks. This allows the use of Shergin’s CLT for finitely dependent random fields and the Lindeberg-Feller CLT, after which de-Poissonization transfers the limit theorem back to the original statistic. Under the least favorable null \(m_j(x)=0\) a.e., the paper obtains
\[
T_n \xrightarrow{d} N(0,1),
\]
so the test is asymptotically distribution free and standard normal critical values have asymptotically correct size [1208.2733].

A different statistical use appears in the analysis of generalization error for discrete-time Markov algorithms. For a time-homogeneous Markov process \((X_k^S)_{k\ge 0}\), Poissonization is defined by
\[
Y_t^S := X_{N_t}^S,
\]
where \(N_t\) is a unit-rate Poisson process. The density \(u_t\) of the Poissonized process satisfies
\[
\frac{\partial u_t}{\partial t}=(P^*-I)u_t,
\]
and the paper derives the entropy flow identity
\[
\frac{d}{dt} D(\rho_t^S \,\|\, \pi_t)
=
\Delta_{P,P_S}(v_t)
-
\mathbb{E}_{x\sim \pi_t,\; y\sim \delta_x P}
\!\left[D_\Phi(v_t(x),v_t(y))\right].
\]
This yields PAC-Bayesian bounds and, via modified logarithmic Sobolev inequalities, time-uniform generalization bounds for both noisy and non-noisy Markov algorithms. A depoissonization theorem then relates expectations under the Poissonized chain to those of the original chain under ergodicity assumptions [2502.07584].

A common misconception is that Poissonization in inference is merely a proof trick. In both papers, it is more specific: it changes the dependence structure in a way that produces either a CLT-compatible block decomposition or an entropy flow unavailable in the original discrete-time formulation.

## 3. Stochastic processes and continuous-time embeddings

For random walks, Poissonization is the standard passage from discrete to continuous time. In the noncolliding system of continuous-time simple symmetric random walks on \(\mathbb{Z}\), the underlying single-particle walk is a compound Poisson process. This Poissonized formulation yields transition probabilities in terms of modified Bessel functions and supports a determinantal description of the noncolliding system for any finite initial configuration without multiple points; the spatio-temporal correlation kernel is expressed using the modified Bessel functions, and the infinite-particle extension with equidistant initial spacing exhibits relaxation to the equilibrium determinantal point process with the sine kernel [1407.0505].

The same continuous-time embedding is used to analyze collisions of two independent simple random walkers on \(\mathbb{Z}^d\). Each walker jumps at rate \(1\) according to a Poisson process, and the collision problem is reduced to the return probability of the difference walk \(D(t)=X(t)-Y(t)\). Coordinatewise analysis gives
\[
\mathbb{P}\big(D_j(t)=0\big)
=
I_0\!\left(\frac{2t}{d}\right)e^{-2t/d},
\]
hence
\[
\mathbb{P}\big(D(t)=0\big)
=
\left[
I_0\!\left(\frac{2t}{d}\right)e^{-2t/d}
\right]^d.
\]
From the asymptotic decay of \(I_0\), the expected number of collisions is finite if and only if \(d\ge 3\) [2505.02973].

Poissonization also appears in branching systems as a representation theorem. For branching Markov processes with spatially varying birth and death rates, particles carry both locations and levels. In the finite-particle model, conditioned on the state at time \(t\), levels are independent and uniformly distributed on \([0,r]\). In the measure-valued limit, conditioned on the measure-valued state \(K(t)\), the joint distribution of locations and levels is conditionally Poisson with mean measure \(K(t)\times \Lambda\), where \(\Lambda\) is Lebesgue measure. The corresponding Laplace functional is
\[
E\!\left[\exp\!\left\{-\int f\,d\xi\right\}\middle|K\right]
=
\exp\!\left\{-\int (1-e^{-f})\, dK\times du\right\},
\]
which simplifies calculations for extinction, nonextinction, Harris-type limits, and diffusion approximations in random environments [1104.1496].

Across these examples, Poissonization converts stepwise or combinatorially constrained models into continuous-time or point-process models with explicit semigroup, kernel, or Laplace-functional structure.

## 4. Combinatorics, occupancy, and analytic de-Poissonization

In infinite occupancy schemes, Poissonization traditionally means replacing a fixed number \(n\) of balls by a Poisson(\(n\)) number. This turns box counts independent and facilitates renewal-theoretic or generating-function arguments. In the Bernoulli sieve, earlier work on \(K_n^*(1)\) used Poissonization-de-Poissonization, but the 2016 paper proves functional limit theorems without that step. Its key approximation is
\[
K_n^*(t) \approx \rho^*(n)-\rho^*(n^{(1-t)-}),
\qquad t\in[0,1],
\]
where \(\rho^*(x)\) counts “large boxes,” and for the Bernoulli sieve
\[
\rho^*(n^t)=\#\{k:T_k^*\le t\log n\}
\]
for a perturbed random walk \(T_k^*\). The paper emphasizes that this direct approximation allows pathwise functional limit theorems in Skorohod space and dispenses with a Poissonization-de-Poissonization step that had been essential in previous studies [1601.04274].

The longest increasing subsequence problem provides a complementary picture: Poissonization remains central, but de-Poissonization becomes analytic. With \(N_r\sim \mathrm{Poisson}(r)\), the Poissonized distribution is
\[
P(r;\ell)=e^{-r}\sum_{n=0}^{\infty}\Pr(L_n\le \ell)\frac{r^n}{n!}.
\]
The paper replaces Johansson’s monotonicity-based de-Poissonization by analytic de-Poissonization of Jacquet and Szpankowski, based on complex-plane growth conditions and a tameness hypothesis on zeros of the analytically continued Poissonized length distribution. This yields an asymptotic expansion in powers of \(n^{-1/3}\) for \(\Pr(L_n\le \ell)\) near edge scaling [2301.02022].

Urn extrapolation offers a further use of Poissonization at finite sample sizes. In the microbial-unknown problem, modeling the number of draws by a Poisson variable leads to exact formulas for the unsampled mass and to an “Embedding algorithm” for prediction intervals. The paper states that conditionally unbiased predictors and exact prediction intervals of constant length in logarithmic scale are possible for the fraction of the environment that belongs to unsampled classes [1109.2990].

These works clarify that de-Poissonization is neither automatic nor always necessary. In some settings it is indispensable and highly nontrivial; in others, new arguments bypass it entirely.

## 5. Coding, online selection, and latent-variable learning

For large-alphabet compression, Poissonization replaces multinomially constrained counts by independent Poisson counts. If
\[
N_j \sim \operatorname{Poisson}(\lambda_j)\quad\text{independently},\qquad
N=\sum_{j=1}^m N_j \sim \operatorname{Poisson}(\lambda_{\mathrm{sum}}),
\]
then conditioning on \(N=n\) recovers the multinomial model with \(\theta_j=\lambda_j/\lambda_{\mathrm{sum}}\). The paper combines this with exponential tilting:
\[
P_a(N_j)=\frac{M(N_j)e^{-aN_j}}{C_a},
\qquad
M(k)=\frac{k^k e^{-k}}{k!},
\]
so that the joint coding distribution factorizes as \(Q_a(\underline N)=\prod_{j=1}^m P_a(N_j)\). The stated result is that the strategy is optimal within the class of distributions satisfying a moment condition and close to optimal for the class of all i.i.d. distributions on strings of a given length [1401.3760].

In prophet inequalities, Poissonization is combined with sharding. Each random variable is split into many independent shards, and for threshold \(\tau\) the number of shards exceeding \(\tau\) approaches a Poisson random variable with parameter
\[
\lambda_\tau = -\log\!\left(\prod_{i=1}^n F_i(\tau)\right).
\]
This unified framework improves analyses for the \(\textsc{Top-1-of-}k\) prophet inequality, prophet secretary inequality, and semi-online prophet inequality, while also simplifying several known proofs [2307.00971].

For Gaussian-mixture learning, Poissonization is used as a structural reduction. If \(R\sim \mathrm{Poisson}(\lambda)\), one draws \(R\) independent samples from the mixture and sums them. The number of times component \(j\) is selected becomes
\[
S_j \sim \mathrm{Poisson}(w_j\lambda),
\]
and the \(S_j\) are independent. The transformed sample has the form
\[
Y = AS + \eta(R),
\]
or, after noise standardization,
\[
X = AS + \eta(\tau),
\]
which is an underdetermined ICA model. The paper’s main theorem uses this Poissonization-based technique to transform a mixture of Gaussians to a linear map of a product distribution, enabling recovery via tensor decomposition and ICA under a nondegeneracy condition on the means [1311.2891].

A plausible implication is that Poissonization is especially effective when the obstruction is a global counting constraint—multinomial dependence in coding, threshold exceedance counts in online selection, or categorical latent variables in mixture models.

## 6. Geometric Poissonization in Jacobi and nonholonomic systems

In Jacobi geometry, Poissonization associates a Jacobi manifold \((M,\pi,E)\) to a Poisson manifold on \(M\times \mathbb{R}\). The standard bivector is
\[
e^{-t}(\pi+\partial_t\wedge E),
\]
and the 2018 paper shows that gauge transformations of Jacobi structures are related to gauge transformations of Poisson structures via this Poissonization. The same work proves that Poissonization commutes with Jacobi gauge transformation after lifting the \(1\)-form \(B\) to the corresponding closed \(2\)-form on \(M\times \mathbb{R}\) [1801.06681].

The degree-\(1\) supergeometric version identifies Poissonization with symplectization. For a Jacobi manifold \((M,\Lambda,R)\), the associated Poisson structure is
\[
\Lambda' = e^{-t}(\Lambda + \theta R),
\]
and the corresponding graded symplectic form is
\[
\tilde\omega = d(e^t\alpha)=e^t(dt\wedge \alpha + d\alpha).
\]
The paper establishes a one-to-one correspondence between degree-\(1\) contact \(NQ\)-manifolds with fixed contact form and Jacobi manifolds, making Poissonization the supergeometric counterpart of symplectization [1111.4705].

This passage is used directly in field theory. For Jacobi sigma models, the Poissonized target is \(\widetilde M=\mathbb{R}\times M\) with
\[
\alpha = e^{-t}(\Lambda+\partial_t\wedge R).
\]
The fields are then expressed as perturbative expansions in terms of reduced boundary phase-space variables, and the procedure is carried out explicitly up to second order, including the contact-manifold case [2012.02756].

A distinct but related use occurs in nonholonomic dynamics. Starting from
\[
\dot{\mathbf x}=\mathbf w\times \nabla H,
\]
one extends the antisymmetric operator from \(3\) to \(4\) dimensions by introducing a variable \(s\),
\[
\mathfrak J = \mathcal J + a\,\partial_x\wedge\partial_s + b\,\partial_y\wedge\partial_s + c\,\partial_z\wedge\partial_s,
\]
with conformal factor
\[
r=\mathbf w\cdot \mathbf D + sh.
\]
After the time reparametrization
\[
\frac{d\tau}{dt}=r,
\]
the rescaled operator satisfies the Jacobi identity, so the system becomes Hamiltonian in extended coordinates. The paper emphasizes that this Poissonization does not rely on the specific form of the Hamiltonian [1806.05003].

More recently, Jacobi Hamiltonian integrators use Poissonization to construct geometric integrators. The Jacobi manifold \((J,\Lambda,E)\) is lifted to the homogeneous Poisson manifold \(P=J\times \mathbb{R}^\times\) with
\[
\Pi=\frac{1}{t}\Lambda + \frac{\partial}{\partial t}\wedge E,
\qquad
\hat H(x,t)=tH(x).
\]
This enables the application of homogeneous symplectic bi-realizations and Magnus-type constructions to obtain structure-preserving Jacobi Hamiltonian integrators [2601.20799].

The geometric literature therefore uses “Poissonization” in a stronger sense than mere randomization: it denotes a functorial or constructive passage from Jacobi or non-Poisson data to genuinely Poisson data.

## 7. Noncommutative Poissonization and many-body or universe field theory

In operator algebra, Poissonization is a noncommutative Poisson random measure. Given a von Neumann algebra \(N\) with a normal semifinite faithful weight \(\omega\), the construction produces a new von Neumann algebra \(\mathbb{P}_\omega N\) with a canonical normal faithful state \(\phi_\omega\). The paper states this as a functor
\[
Poiss: vNa_w \to vNa_s,
\qquad
(N,\omega)\mapsto (\mathbb{P}_\omega N,\phi_\omega).
\]
For selfadjoint \(x\), the Poisson moment formula is
\[
\phi_\omega(\Gamma(e^{ix})) = \exp(\omega(e^{ix}-1)),
\]
and the Haagerup \(L_2\)-space satisfies
\[
L_2(\mathbb{P}_\omega N,\phi_\omega)\cong \mathcal F_s(L_2(N,\omega)).
\]
The construction is compatible with normal weight-preserving homomorphisms and with unital normal completely positive weight-preserving maps [2303.14580].

A 2025 development extends this perspective to “third quantization” and topology change in gravity. There, Poissonization takes as input the observable algebra and an unnormalized state of a quantum system and outputs a von Neumann algebra of a many-body theory represented on its symmetric Fock space. The resulting correlators are governed by a set-partition formula,
\[
\varphi_\omega(\lambda(O_1)\cdots\lambda(O_p))
=
\sum_{\sigma\in P_p}\prod_{A\in \sigma}
\omega\!\left(\overrightarrow{\prod_{j\in A} O_j}\right),
\]
with connected correlators
\[
\varphi_\omega(\lambda(O_1)\cdots\lambda(O_p))_{\rm conn}
=
\omega(O_1\cdots O_p).
\]
The paper argues that rare topology change events are universally described by a Poisson process at exponentially late times, that the statistics of the total number of baby universes are captured by a coherent state, and that the multi-boundary correlators of the Marolf-Maxfield model, closed-open \(2\)D TQFT, and late-time JT gravity are entirely captured by Poissonization [2509.02293].

The literature therefore shows that Poissonization is not confined to classical probability. It extends from Poisson sample-size randomization and Poisson-process embeddings to geometric symplectization, Hamiltonization of nonholonomic systems, and noncommutative many-body constructions. The word keeps the same formal center—a passage to Poisson structure—while the surrounding mathematics changes radically with the domain.

Source: https://www.emergentmind.com/topics/poissonization