---
title: Optimal Sub-Poisson Variance Proxy
url: https://www.emergentmind.com/topics/optimal-sub-poisson-variance-proxy
type: topic
---

# Optimal Sub-Poisson Variance Proxy

Searching arXiv for the key papers on sub-Poisson variance proxies and related proxy-variance frameworks.
{"query":"id:2508.12103 OR \"Sub-Poisson distributions: Concentration inequalities, optimal variance proxies, and closure properties\"","max_results":5}
{"query":"id:1106.1892 OR \"Photon Antibunching, Sub-Poisson Statistics and Cauchy-Bunyakovsky and Bell's Inequalities\"","max_results":5}
{"query":"id:1906.11019 OR \"Observation of scalable sub-Poissonian-field lasing in a microlaser\"","max_results":5}
{"query":"id:1705.00048 OR id:1901.09188 OR id:2403.08628 OR id:2606.06384","max_results":10}
{"query":"id:1701.07263 OR \"Likelihood ratio Haar variance stabilization and normalization for Poisson and other non-Gaussian noise removal\"","max_results":5}
Optimal sub-Poisson variance proxy is the smallest variance-like parameter that makes the centered moment generating function of a random variable no larger than that of a centered Poisson benchmark. In the concentration-theoretic formulation, the benchmark function is
\[
\phi(\lambda)\coloneqq e^\lambda-1-\lambda,
\]
the logarithmic moment generating function of a centered Poisson variable with unit variance, and the resulting proxy plays the Poisson/Bennett analogue of the optimal sub-Gaussian variance proxy [2508.12103]. The term also has an older quantum-optical lineage, where sub-Poisson behavior is identified by photon-number variance falling below the mean, and variance-based scalars such as \(K=\langle \Delta n^2\rangle-\langle n\rangle\) serve as nonclassicality witnesses rather than as mgf-optimal parameters [1106.1892].

## 1. Formal definition in the Poisson-mgf framework

The modern concentration-theoretic definition starts from an integrable real-valued random variable \(X\), centered in the exponent by subtracting \(\mathbb{E}X\). The basic one-sided notion is **upper sub-Poisson**: there exists \(\sigma^2\ge 0\) such that
\[
\mathbb{E}e^{\lambda(X-\mathbb{E}X)}\le e^{\sigma^2\phi(\lambda)}\qquad \text{for all }\lambda\ge 0.
\]
Lower sub-Poissonity is defined by applying the same condition to \(-X\), and two-sided sub-Poissonity requires both sides; equivalently,
\[
\mathbb{E}e^{\lambda(X-\mathbb{E}X)}\le e^{\sigma^2\phi(|\lambda|)}\qquad \text{for all }\lambda\in\mathbb{R}.
\]
The optimal two-sided sub-Poisson variance proxy is then
\[
\tau_{\mathrm{SP}}^2(X)\coloneqq \sup_{\lambda\ne 0}\frac{\log \mathbb{E}e^{\lambda(X-\mathbb{E}X)}}{\phi(|\lambda|)},
\]
while the optimal upper proxy is
\[
\tau_{\mathrm{SP},+}^2(X)\coloneqq \sup_{\lambda>0}\frac{\log \mathbb{E}e^{\lambda(X-\mathbb{E}X)}}{\phi(\lambda)},
\]
and the lower proxy satisfies
\[
\tau_{\mathrm{SP},-}^2(X)=\tau_{\mathrm{SP},+}^2(-X).
\]
The two-sided proxy is the larger of the upper and lower proxies. For integrable \(X\), finiteness of these quantities is equivalent to the corresponding sub-Poisson property, and the optimal proxy is exactly the smallest admissible \(\sigma^2\) in the defining mgf inequality [2508.12103].

This formulation can also be expressed through mgf order. If \(\tilde X=X-\mathbb{E}X\) and \(\tilde Y=Y-\mathbb{E}Y\) with \(Y\sim \mathrm{Poisson}(\sigma^2)\), then upper sub-Poissonity means \(\tilde X\) is dominated by \(\tilde Y\) in moment generating function order for all nonnegative tilts. The parameter \(\sigma^2\) is therefore not merely a convenient constant; it is the intensity and variance of the benchmark centered Poisson variable itself [2508.12103].

## 2. Quantum-optical antecedents and variance-shortfall diagnostics

Before the concentration-theoretic notion of optimality was introduced, sub-Poisson language was already standard in quantum optics. In the single-mode radiation setting with annihilation and creation operators \(a\) and \(a^*\), satisfying \([a,a^*]=1\), the photon number operator is \(n=a^*a\), and the quantity
\[
K=\langle \Delta n^2\rangle-\langle n\rangle
\]
measures the difference between the photon-number variance and mean. Since Poisson statistics are characterized by equality of variance and mean, \(K<0\) is exactly the sub-Poisson condition, \(K=0\) is the Poisson boundary, and \(K>0\) corresponds to non-sub-Poisson behavior by that criterion [1106.1892].

Using the commutation relation, the same quantity can be written in normally ordered form,
\[
K=\langle a^{*2}a^2\rangle-\langle a^*a\rangle^2,
\]
and for a pure state \(\psi\),
\[
K=\|a^2\psi\|^2-\|a\psi\|^4.
\]
A central result is that any classical hidden-variable representation of the relevant moments would imply
\[
K=E(|\alpha|^4)-\bigl(E(|\alpha|^2)\bigr)^2\ge 0,
\]
by the Cauchy-Bunyakovsky inequality. Thus \(K<0\) rules out such a classical representation. For \(n\)-particle Fock states \(\psi_n\), the paper computes
\[
K=-n<0,\qquad n=1,2,\dots,
\]
so number states are explicit sub-Poisson and nonclassical examples [1106.1892].

This older use of “sub-Poisson variance proxy” differs from the later optimal-proxy literature. The 2011 analysis singles out \(K=\operatorname{Var}(n)-\langle n\rangle\) as the natural variance-based witness of sub-Poisson nonclassicality, but it does **not** prove any extremality or optimization theorem among all possible witnesses. The paper explicitly does **not** establish that \(K\) is uniquely best, maximally sensitive, or optimal in a universal sense [1106.1892].

## 3. Variance lower bounds, degeneracy, and exact examples

The concentration-theoretic optimal proxy is genuinely variance-like. For any integrable random variable,
\[
\operatorname{Var}(X)\le \tau_{\mathrm{SP},+}^2(X)\le \tau_{\mathrm{SP}}^2(X),
\]
so every admissible one-sided or two-sided sub-Poisson proxy dominates the true variance. Moreover, the optimal proxy vanishes if and only if the variable is degenerate:
\[
X=\mathbb{E}X\ \text{a.s.}\quad \Longleftrightarrow\quad \tau_{\mathrm{SP}}^2(X)=0.
\]
These facts place the proxy on the same conceptual footing as the optimal sub-Gaussian variance proxy, while preserving Poisson-tail asymmetry [2508.12103].

Several canonical families admit exact formulas, and in many of them the optimal sub-Poisson proxy equals the ordinary variance.

| Distribution | Optimal sub-Poisson proxy | Notes |
|---|---:|---|
| \(\mathrm{Bernoulli}(p)\) | \(p(1-p)\) | Equals variance |
| \(\mathrm{Binomial}(n,p)\) | \(np(1-p)\) | Equals variance |
| Rademacher | \(1\) | Equals variance |
| Scaled Rademacher \(\pm a\) | \(a^2\) | Equals variance |
| \(\mathrm{Poisson}(a)\) | \(a\) | Equals variance/intensity |
| Skellam \(\mathrm{Poisson}(a_1)-\mathrm{Poisson}(a_2)\) | \(a_1+a_2\) | Equals variance |
| Gaussian with variance \(\sigma^2\) | \(\sigma^2\) | Also sub-Gaussian |
| Exponential with rate \(a\) | lower proxy \(=1/a^2\) | Not upper sub-Poisson |

The Bernoulli case is especially revealing. The optimal sub-Poisson proxy is \(p(1-p)\), whereas the optimal sub-Gaussian proxy is
\[
\frac{1/2-p}{\log(1/p)+\log(1-p)}\qquad (p\neq 1/2).
\]
In the sparse regime \(p\to 0\), the sub-Poisson proxy behaves like \(p\), while the sub-Gaussian proxy behaves like \(1/\log(1/p)\). This is a sharp illustration of why a Poisson benchmark is more faithful than a quadratic Gaussian benchmark for sparse Bernoulli or count-like tails [2508.12103].

The examples also show that one-sided and two-sided notions need not coincide. The exponential distribution is not upper sub-Poisson because its mgf blows up at a finite positive tilt, but it is lower sub-Poisson with optimal lower proxy \(1/a^2\). This asymmetry is intrinsic to the framework rather than an artifact of proof technique [2508.12103].

## 4. Bennett-type concentration and structural properties

Once an upper sub-Poisson proxy \(\sigma^2\) is available, the standard Chernoff argument yields a Bennett-type tail bound:
\[
\Pr(X\ge t)\le e^{-\sigma^2}\left(\frac{e\sigma^2}{\sigma^2+t}\right)^{\sigma^2+t},\qquad t\ge 0.
\]
Equivalently,
\[
\Pr(X\ge t)\le \exp\!\left(-\sigma^2 h\!\left(\frac{t}{\sigma^2}\right)\right),
\qquad
h(u)=(1+u)\log(1+u)-u.
\]
The same proposition also gives the Bernstein-style relaxations
\[
\Pr(X\ge t)\le \exp\left(-\frac{t^2/2}{\sigma^2+t/3}\right)
\]
and
\[
\Pr(X\ge t)\le \exp\left(-\left(\frac{t^2}{4\sigma^2}\wedge \frac{3t}{4}\right)\right).
\]
Using the optimal upper proxy \(\tau_{\mathrm{SP},+}^2(X)\) gives the sharpest inequality available within this mgf class [2508.12103].

The optimal proxy also has robust closure properties. For independent integrable \(X_1,\dots,X_n\),
\[
\tau_{\mathrm{SP},+}^2\!\left(\sum_i X_i\right)\le \sum_i \tau_{\mathrm{SP},+}^2(X_i),
\qquad
\tau_{\mathrm{SP}}^2\!\left(\sum_i X_i\right)\le \sum_i \tau_{\mathrm{SP}}^2(X_i).
\]
For convex combinations, the functional \(X\mapsto \tau_{\mathrm{SP}}^2(X)\) is convex and balanced:
\[
\tau_{\mathrm{SP}}^2((1-a)X+aY)\le (1-a)\tau_{\mathrm{SP}}^2(X)+a\tau_{\mathrm{SP}}^2(Y),
\qquad 0\le a\le 1,
\]
and for \(|a|\le 1\),
\[
\tau_{\mathrm{SP}}^2(aX)\le a^2\tau_{\mathrm{SP}}^2(X)\le \tau_{\mathrm{SP}}^2(X).
\]
If \(X\) is sub-Poisson, then \(|X-\mathbb{E}X|\) is upper sub-Poisson with
\[
\tau_{\mathrm{SP},+}^2(|X-\mathbb{E}X|)\le 2\tau_{\mathrm{SP}}^2(X).
\]
These properties make the proxy stable under operations that routinely appear in nonasymptotic probability [2508.12103].

A crucial limitation is failure under arbitrary scalar multiplication. For a symmetric Skellam variable \(X_1\) with parameters \(a_1=a_2=1\), the scaled variable \(X_a=aX_1\) is sub-Poisson only for \(a\le 1\), with
\[
\tau_{\mathrm{SP}}^2(X_a)=2a^2,
\]
whereas for \(a>1\),
\[
\tau_{\mathrm{SP}}^2(X_a)=\infty.
\]
Thus the centered sub-Poisson class is not a linear space, and the proxy is not a norm [2508.12103].

## 5. Relation to optimal proxy variance more broadly

The phrase “optimal variance proxy” comes from a broader mgf-based literature on sub-Gaussianity. In that setting, a random variable is sub-Gaussian if
\[
\mathbb{E}e^{\lambda(X-\mu)}\le \exp\!\left(\frac{\lambda^2\sigma^2}{2}\right)\qquad \forall \lambda\in\mathbb{R},
\]
and the optimal proxy is the smallest such \(\sigma^2\), equivalently the supremum of a normalized centered log-mgf. Exact optimal sub-Gaussian proxies have been derived for Beta and Dirichlet laws [1705.00048], for bounded random variables through the variational maximization of
\[
h(\lambda)=\frac{2}{\lambda^2}\mathcal K(\lambda)
\]
[1901.09188], and for truncated Gaussian and truncated exponential laws [2403.08628]. An estimation theory for the sub-Gaussian parameter now treats the proxy as a distributional functional
\[
\xi_*^2=\sup_{\lambda\in\mathbb{R}}L(\lambda),
\qquad
L(\lambda)=\frac{2}{\lambda^2}\log \mathbb{E}e^{\lambda X},
\]
and shows that truncated empirical maximization can be consistent, with rates governed by whether the maximizer lies in a bounded \(\lambda\)-region [2606.06384].

The optimal sub-Poisson variance proxy follows the same variational pattern, but with the quadratic Gaussian benchmark replaced by the Poisson benchmark \(\phi(\lambda)=e^\lambda-1-\lambda\). This replacement preserves the “smallest global envelope” interpretation while adapting the tail geometry from Gaussian to Poisson/Bennett form [2508.12103]. A plausible implication is that empirical estimation of the sub-Poisson proxy should proceed by maximizing a normalized empirical cgf over a suitably controlled set of tilts, although that estimation theory is not developed in the cited sub-Poisson paper.

Not every variance-stabilizing construction belongs to this optimal-proxy lineage. In Poisson denoising, for example, the likelihood ratio Haar methodology replaces ordinary Haar coefficients by signed square-root generalized likelihood ratio statistics \(g_{j,k}\) that are approximately \(N(0,1)\) under the local null. Its contribution is a local deviance-based normalization rather than a single scalar optimal variance proxy [1701.07263].

## 6. Operational proxies in quantum-optical experiments

In experimental quantum optics, sub-Poissonianity is often reported through normalized variance diagnostics rather than through mgf-optimal proxies. A prominent example is the Mandel parameter
\[
Q=\frac{\Delta n^2}{\langle n\rangle}-1,
\]
which is the main operational indicator in microlaser experiments. It is directly tied to the zero-delay second-order correlation by
\[
g^{(2)}(0)=1+\frac{Q}{\langle n\rangle},
\qquad
Q=\langle n\rangle\bigl(g^{(2)}(0)-1\bigr).
\]
Thus \(Q<0\), \(g^{(2)}(0)<1\), and \(\Delta n^2<\langle n\rangle\) are equivalent signatures of sub-Poissonianity for the single-mode field considered in that work [1906.11019].

The microlaser study reports deadtime-free Mandel \(Q\) less than \(-0.6\) at
\[
\langle n\rangle=592\pm 5,
\]
corresponding to photon-number variance \(4\,\mathrm{dB}\) below the standard quantum limit and to normalized variance
\[
\frac{\Delta n^2}{\langle n\rangle}=1+Q_0=0.38(5)
\]
for the best observed point \(Q_0=-0.62(5)\). The paper emphasizes that \(Q\) is preferable to raw \(g^{(2)}(0)\) as a magnitude proxy when \(\langle n\rangle\) is large, because \(g^{(2)}(0)-1=Q/\langle n\rangle\) becomes numerically small even when variance suppression remains strong [1906.11019].

These experimental proxies are operationally central but conceptually distinct from the optimal sub-Poisson variance proxy of concentration theory. The former normalize photon-number variance relative to the Poisson baseline and are tailored to inference from \(g^{(2)}\) and \(\langle n\rangle\); the latter is the smallest mgf parameter yielding a global Poisson-envelope inequality. Both quantify deviation from Poisson behavior, but they do so at different mathematical levels and for different inferential purposes [1906.11019].

## 7. Conceptual synthesis

The modern optimal sub-Poisson variance proxy is best understood as the Poisson analogue of the optimal sub-Gaussian proxy: a sharp mgf-based parameter obtained by comparing the centered log-mgf of \(X\) to the centered Poisson log-mgf \(\phi(\lambda)=e^\lambda-1-\lambda\). It is finite exactly for sub-Poisson variables, always dominates the variance, equals the variance for many canonical families, and yields Bennett-type concentration without boundedness assumptions [2508.12103].

Historically, the same adjective “sub-Poisson” entered through quantum optics, where the decisive scalar was the variance shortfall \(K=\operatorname{Var}(n)-\langle n\rangle\), negative for number states and classically forced to be nonnegative under a hidden-variable representation [1106.1892]. Experimental practice then favored normalized observables such as Mandel \(Q\), which remain directly measurable in cavity-QED and microlaser settings [1906.11019]. The current concentration-theoretic notion of optimality does not replace those earlier diagnostics; rather, it provides a different and more general answer to the question of what the sharpest Poisson-tail variance parameter should be.

Source: https://www.emergentmind.com/topics/optimal-sub-poisson-variance-proxy