---
title: Probability Distribution Reconstruction
url: https://www.emergentmind.com/topics/probability-distribution-reconstruction
type: topic
---

# Probability Distribution Reconstruction

Probability distribution reconstruction denotes the recovery of a probability measure, a probability density, or a conditional law from partial information such as finitely many moments, gauge-invariant cumulants, halfspace depth, geometric distribution functions, weighted edge integrals, indirect experimental observables, or an observed pushforward law \(\rho_y=G_\#\rho_x\). In the cited literature, the reconstructed object ranges from a Wannier-center distribution in the Rice–Mele model to a discrete law on \(\mathbb Z^+\), a multivariate density on \([0,1]^d\), a conditional density \(p(y\mid x)\), or an input measure \(\rho_x\) consistent with an output distribution. This suggests that probability distribution reconstruction is best understood as a family of inverse problems whose structure is determined by the observation operator, the admissible class of distributions, and the criterion used to resolve nonuniqueness or instability [1705.07955] [2208.11551] [2504.18999].

## 1. Problem classes and reconstructed objects

Across the literature, the forward information supplied to the reconstruction procedure varies substantially. In some settings, one observes low-order statistics such as moments or cumulants. In others, one observes geometric summaries such as halfspace depth or the geometric cdf, or indirect experimental data generated by a physical forward model. There are also supervised settings in which one learns a map from one distribution to another from paired training densities, and conditional settings in which the target is a full law \(p(y\mid x)\) rather than a point prediction [1408.2653] [2208.03959] [2206.06194] [1811.08793].

| Setting | Observed object | Reconstructed object |
|---|---|---|
| Moment/cumulant inversion | finitely many moments or gauge-invariant cumulants | a density or mass function |
| Geometric inversion | halfspace depth or geometric cdf | support, atoms, or the full measure |
| Pushforward inverse problem | \(\rho_y=G_\#\rho_x\) | an input measure \(\rho_x\) |
| Indirect physical observation | ensemble-averaged or transformed measurements | a latent distribution over states or parameters |

The same distinction appears in tree-indexed Markov random fields, where the object of interest is the residual distribution at the root: the law of the posterior vector \(\eta_T(L(n))\) induced by a random boundary configuration. There, exact recursive computation is available in principle, but the support of the posterior law grows exponentially with depth, so reconstruction becomes a problem of controlled approximation rather than direct enumeration [0903.4812].

## 2. Moment, cumulant, and coefficient-based reconstruction

A large class of methods reconstructs distributions from finitely many statistics. In the Rice–Mele model, the starting point is the hierarchy of gauge-invariant cumulants \(C_n\) associated with the Zak phase. These are converted into gauge-invariant moments \(\mu_C^{(n)}\), and the distribution is reconstructed by maximizing Shannon entropy under moment constraints. The resulting density has the standard form
\[
P(x)=C\exp\left(-\sum_k A_k x^k\right),
\]
with the coefficients obtained numerically by minimizing
\[
\chi^2=\sum_k(\mu_P^{(k)}-\mu_C^{(k)})^2.
\]
The reconstructions use the first six gauge-invariant moments, start from the Gaussian defined by the first two cumulants, and use simulated annealing. When the Wannier functions are localized within one unit cell, the reconstructed probability corresponds to the squared modulus of the Wannier function [1705.07955].

For one-dimensional discrete distributions on \(\mathbb Z^+\), the same maximum-entropy principle leads to the exponential-family form
\[
q(x)=\frac1Z\exp\left(-\sum_{k=1}^{M}\lambda_k x^k\right),
\]
subject to finitely many non-central moments \(\mu_k=E[X^k]\). Because the support is unbounded, the main technical issue is not the formal entropy maximization itself but the adaptive support approximation used to localize computation where the main part of the probability mass is located. The dual objective
\[
\Psi(\lambda)=\ln Z+\sum_{k=1}^{M}\lambda_k\mu_k
\]
is minimized by a Levenberg–Marquardt iteration, with gradient and Hessian computed from truncated exponential sums [1408.2653].

A different route replaces moment inversion by orthonormal expansions. For yield-curve parameters, the joint density is expanded on \([0,1]^d\) as
\[
\rho(x)=\sum_{j_1\ldots j_d=0}^m a_j\,f_{j_1}(x_1)\cdots f_{j_d}(x_d),
\]
where the \(f_j\) are orthonormal polynomials and the coefficients are estimated independently by sample means,
\[
a_j=\frac1n\sum_{t=1}^n f_j(x^t).
\]
After marginal normalization, the baseline is \(\rho\approx 1\), so nonconstant coefficients act as cumulant-like perturbations from uniformity. The same logic underlies Hierarchical Correlation Reconstruction for conditional densities, where
\[
\tilde\rho(Y=y\mid X=x)=1+\sum_{i=1}^m f_i(y)a_i(x)
\]
and each coefficient \(a_i(x)\) is predicted separately by MSE regression, optionally with CCA and lasso [1807.11743] [2206.06194].

These constructions differ in algebraic form, but they share a common structure: a low-order statistical summary is first converted into a function-space representation, and reconstruction is then obtained by a variational principle, orthogonal projection, or coefficient-wise regression.

## 3. Geometric and measure-theoretic inversion

A second major paradigm reconstructs distributions from geometric summaries. For Tukey’s halfspace depth,
\[
D(x;\mu)=\inf_{H\in\mathcal H(x)}\mu(H),
\]
full reconstruction fails in general: distinct measures can have the same halfspace depth function everywhere in \(\mathbb R^d\). The paper nevertheless proves three kinds of recoverable information. First, the support is constrained by
\[
\operatorname{supp}\mu \subseteq \operatorname{cl}\Big(\bigcup_{\alpha\ge 0}\operatorname{bd}(D_\alpha(\mu))\Big).
\]
Second, an atom \(x\) with \(D(x;\mu)=\alpha\) must be an extreme point of every \(D_\beta(\mu)\) for \(\beta\in(\alpha-\mu(\{x\}),\alpha]\). Third, atoms create directional jumps in depth along lines passing through them. The framework relies crucially on minimizing flag halfspaces, which always exist even when minimizing ordinary halfspaces do not [2208.03959].

By contrast, the geometric or spatial distribution function
\[
F_P^\#(x)=E\!\left[\frac{x-Z}{\|x-Z\|}I[Z\neq x]\right]
\]
admits an explicit inversion. The key operator is
\[
L_d=\gamma_d\,(-\Delta)^{\frac{d-1}{2}}(\nabla\cdot\,),
\qquad
\gamma_d^{-1}=2^d\pi^{(d-1)/2}\Gamma\!\left(\frac{d+1}{2}\right),
\]
and the reconstruction theorem states that
\[
P=L_d(F_P^\#)
\]
in \(\mathcal S'(\mathbb R^d)\) for arbitrary Borel probability measures. For sufficiently regular densities, the same identity holds pointwise. The inversion is local in odd dimension and nonlocal in even dimension, because \((d-1)/2\) is an integer in the first case and a half-integer in the second [2208.11551].

A third formulation places the inverse problem directly on probability measure space. Given \(\rho_y=G_\#\rho_x\), the overdetermined problem is posed as
\[
\min_{\rho_x} D(G_\#\rho_x,\rho_y),
\]
while the underdetermined problem is
\[
\min_{\{G_\#\rho_x=\rho_y\}} E[\rho_x].
\]
The paper shows that the choice of \(D\) or \(E\) changes the meaning of “reconstruction.” With a \(\phi\)-divergence, the output reconstruction is the conditional distribution of \(\rho_y\) on the attainable range \(R\); with Wasserstein distance, it is the marginal distribution obtained by metric projection onto \(R\). In the underdetermined case, entropy minimization yields a piecewise constant, fiber-uniform reconstruction, whereas minimizing the second moment yields the least-norm solution [2504.18999].

## 4. Indirect observables, controls, and reconstructed latent laws

Many reconstruction problems do not observe the target distribution directly at all. Instead, the distribution is latent in a dynamical or experimental forward model. In the manifold-based formulation of inverse problems over the probability simplex, the observables are \(Y(p)\), the data are \(Y^{\mathrm{EXP}}\), and reconstruction minimizes a mismatch functional \(\Phi(Y(p)\|Y^{\mathrm{EXP}})\). Because the feasible set is the interior of the simplex rather than a Euclidean space, the paper derives the manifold-based gradient descent rule
\[
\Delta \log p = -\tau \nabla_p \Phi(p) + C,
\]
equivalently
\[
p_i^{k+1}=\frac{p_i^k\exp\!\left(-\tau\frac{\partial\Phi}{\partial p_i}(p^k)\right)}
{\sum_j p_j^k\exp\!\left(-\tau\frac{\partial\Phi}{\partial p_j}(p^k)\right)}.
\]
In the protein-ensemble reconstruction example, this manifold-aware update recovers both the pseudo-experimental SAXS data and the pseudo-true distribution when the step size is sufficiently small [2410.01499].

In NMR spin ensembles, the unknown object is a discrete approximation of the joint distribution of \((\alpha,\Delta)\). The measurement model is linear in the probabilities,
\[
Y^{\mathrm{exp}}_{u}(t_f)=\sum_{\ell=1}^{K}P_\star(\ell)\,Y_{u,(\Delta,\alpha)_\ell}(t_f),
\]
and identifiability is controlled by a control-dependent Gram matrix \(W\). The greedy reconstruction algorithm chooses controls so that \(W\) is positive definite, hence the constrained least-squares reconstruction is uniquely solvable. The later SPIRED implementation extends this program to inhomogeneities in all spatial directions and provides the corresponding code for GRA and OGRA [2108.11745] [2402.12379].

Probability distribution reconstruction by circuit cutting targets yet another latent law: the bitstring distribution of a variational quantum circuit after the circuit has been decomposed into fragments. For one cut, the paper reconstructs the output probability \(p(x)\) from weighted sums of empirical fragment probabilities \(q_i(x)\); for a cut CZ gate, six circuits are run and post-selection produces ten cases. The sampling guarantee is of Hoeffding type, and the method is compared under “train-then-cut” and “cut-then-train.” In the 4-qubit classifier example with 4 cut CNOT gates and 24 trainable parameters, a full parameter-shift gradient in the cut-then-fit regime requires
\[
2\times 24\times 6^4 = 62208
\]
circuit evaluations [2510.03077].

Two further examples sharpen the same pattern. In pulse-stream reconstruction, the distribution of short sample trains along a curve in \(\mathbb R^{d+1}\) determines the underlying pulse uniquely when the curve is regularly parametrized, so the latent object is a signal recovered from the probability law of local observations [2302.02664]. In stochastic gravitational-wave backgrounds, the target is the full 1-point PDF \(P(h)\) of strain fluctuations. There the reconstruction is performed by inverting a compound-Poisson characteristic function, not by estimating a mean or a power spectrum alone [1910.04587].

## 5. Distribution-to-distribution regression and local probabilistic representations

A different branch of the literature reconstructs one distribution from another by learning a map between distributions. In structural health monitoring, the observed input is a density \(g_0\), the missing target is \(f_0\), and the training data are paired PDFs \(\{(g_i,f_i)\}_{i=1}^n\). The method first applies the log-quantile-density transform
\[
\psi(t)=\log q(t)=-\log\{f(Q(t))\},
\]
then represents the transformed response by FPCA scores, and finally fits a function-to-vector RKHS regression
\[
\boldsymbol\xi^{f^*}=F_{\mathrm{reg}}(\psi^{g^*})+\varepsilon.
\]
The predicted score vector reconstructs \(\widehat\psi_0^{f^*}\), which is mapped back to density space by the inverse LQD transform. In general cases this LQD-RKHS method outperforms the conventional distribution-to-distribution regression and the distribution-to-warping function regression, but in extrapolation cases it is inferior to the latter [1811.08793].

Local probabilistic reconstruction also appears in hyperspectral anomaly detection. There, each pixel is represented by a latent multivariate Gaussian distribution via a VAE, and a local expected distribution is estimated from a Chebyshev neighborhood. The anomaly score is the modified Wasserstein distance between the pixel’s own latent Gaussian and the neighborhood average expectation. This is not full density reconstruction in the nonparametric sense; it is a local probabilistic representation and neighborhood-based distribution comparison [2105.06775].

The same local viewpoint is extended to weighted histopolation on triangular meshes. Edge measurements are interpreted as weighted moments under probability densities \(\omega\) on \([-1,1]\), and orthogonal polynomials associated with \(\omega\) generate unisolvent triples for reconstruction in \(\mathbb P_k(T)\). The paper proves a general framework for arbitrary polynomial order \(k\), with Jacobi-type and Gegenbauer densities as explicit examples, and uses an adaptive parameter selection algorithm to minimize the global reconstruction error [2511.07972].

## 6. Identifiability, regularity, and recurring limitations

The literature repeatedly emphasizes that reconstruction is rarely exact without structural assumptions. Finitely many moments do not uniquely determine a distribution in general, which is why maximum entropy is used as a selection principle in both the Rice–Mele and discrete \(\mathbb Z^+\) settings [1705.07955] [1408.2653]. Halfspace depth does not determine a multivariate measure completely; at best it yields support restrictions, persistent extreme points, and atom-detection criteria [2208.03959]. In measure-space inversion, even the meaning of the inverse depends on whether the problem is overdetermined or underdetermined and on whether one chooses a \(\phi\)-divergence, Wasserstein distance, entropy, or second moment [2504.18999].

Regularity is equally decisive. The PDE inversion of the geometric cdf is exact in \(\mathcal S'(\mathbb R^d)\), but a continuous density need not generate a geometric cdf with enough regularity to reconstruct the density pointwise. The paper’s odd/even dichotomy makes this sharper: the inverse operator is local in odd dimension and completely nonlocal in even dimension [2208.11551]. In circuit cutting, reconstructing the full output distribution becomes infeasible when most probabilities are exponentially small, so the method is practical only when a small number of states carry significant mass [2510.03077].

Several methods also trade exact probabilistic validity for tractability. Polynomial density expansions on \([0,1]^d\) are not guaranteed nonnegative, and the yield-curve paper explicitly discusses negative densities as artifacts of the method [1807.11743]. Hierarchical Correlation Reconstruction likewise produces a raw density-like function that must be passed through a positive calibration map and renormalized [2206.06194]. On trees, exact recursive propagation of the residual distribution is infeasible because the support grows exponentially with depth, so the survey formalism keeps only a bounded-support compressed representation that still upper-bounds expected convex functionals of the true posterior law [0903.4812].

Taken together, these limitations show that probability distribution reconstruction is governed by three persistent issues: identifiability of the target law, regularity of the forward summary used for inversion, and numerical control of approximation error when the exact inverse is high-dimensional, nonlocal, or combinatorial.

Source: https://www.emergentmind.com/topics/probability-distribution-reconstruction