Papers
Topics
Authors
Recent
Search
2000 character limit reached

Probability Distribution Reconstruction

Updated 14 July 2026
  • Probability distribution reconstruction is a family of inverse methods that recover probability measures from partial data using moment, cumulant, and geometric summaries.
  • Techniques like maximum entropy, orthonormal projection, and variational principles address challenges of nonuniqueness, instability, and high-dimensional approximation.
  • These methods are applied across fields—from quantum circuits to physical systems—to overcome issues of identifiability and regularity in recovering latent distributions.

Probability distribution reconstruction denotes the recovery of a probability measure, a probability density, or a conditional law from partial information such as finitely many moments, gauge-invariant cumulants, halfspace depth, geometric distribution functions, weighted edge integrals, indirect experimental observables, or an observed pushforward law ρy=G#ρx\rho_y=G_\#\rho_x. In the cited literature, the reconstructed object ranges from a Wannier-center distribution in the Rice–Mele model to a discrete law on Z+\mathbb Z^+, a multivariate density on [0,1]d[0,1]^d, a conditional density p(yx)p(y\mid x), or an input measure ρx\rho_x consistent with an output distribution. This suggests that probability distribution reconstruction is best understood as a family of inverse problems whose structure is determined by the observation operator, the admissible class of distributions, and the criterion used to resolve nonuniqueness or instability (Yahyavi et al., 2017, Konen, 2022, Li et al., 26 Apr 2025).

1. Problem classes and reconstructed objects

Across the literature, the forward information supplied to the reconstruction procedure varies substantially. In some settings, one observes low-order statistics such as moments or cumulants. In others, one observes geometric summaries such as halfspace depth or the geometric cdf, or indirect experimental data generated by a physical forward model. There are also supervised settings in which one learns a map from one distribution to another from paired training densities, and conditional settings in which the target is a full law p(yx)p(y\mid x) rather than a point prediction (Andreychenko et al., 2014, Laketa et al., 2022, Duda, 2022, Chen et al., 2018).

Setting Observed object Reconstructed object
Moment/cumulant inversion finitely many moments or gauge-invariant cumulants a density or mass function
Geometric inversion halfspace depth or geometric cdf support, atoms, or the full measure
Pushforward inverse problem ρy=G#ρx\rho_y=G_\#\rho_x an input measure ρx\rho_x
Indirect physical observation ensemble-averaged or transformed measurements a latent distribution over states or parameters

The same distinction appears in tree-indexed Markov random fields, where the object of interest is the residual distribution at the root: the law of the posterior vector ηT(L(n))\eta_T(L(n)) induced by a random boundary configuration. There, exact recursive computation is available in principle, but the support of the posterior law grows exponentially with depth, so reconstruction becomes a problem of controlled approximation rather than direct enumeration (0903.4812).

2. Moment, cumulant, and coefficient-based reconstruction

A large class of methods reconstructs distributions from finitely many statistics. In the Rice–Mele model, the starting point is the hierarchy of gauge-invariant cumulants CnC_n associated with the Zak phase. These are converted into gauge-invariant moments Z+\mathbb Z^+0, and the distribution is reconstructed by maximizing Shannon entropy under moment constraints. The resulting density has the standard form

Z+\mathbb Z^+1

with the coefficients obtained numerically by minimizing

Z+\mathbb Z^+2

The reconstructions use the first six gauge-invariant moments, start from the Gaussian defined by the first two cumulants, and use simulated annealing. When the Wannier functions are localized within one unit cell, the reconstructed probability corresponds to the squared modulus of the Wannier function (Yahyavi et al., 2017).

For one-dimensional discrete distributions on Z+\mathbb Z^+3, the same maximum-entropy principle leads to the exponential-family form

Z+\mathbb Z^+4

subject to finitely many non-central moments Z+\mathbb Z^+5. Because the support is unbounded, the main technical issue is not the formal entropy maximization itself but the adaptive support approximation used to localize computation where the main part of the probability mass is located. The dual objective

Z+\mathbb Z^+6

is minimized by a Levenberg–Marquardt iteration, with gradient and Hessian computed from truncated exponential sums (Andreychenko et al., 2014).

A different route replaces moment inversion by orthonormal expansions. For yield-curve parameters, the joint density is expanded on Z+\mathbb Z^+7 as

Z+\mathbb Z^+8

where the Z+\mathbb Z^+9 are orthonormal polynomials and the coefficients are estimated independently by sample means,

[0,1]d[0,1]^d0

After marginal normalization, the baseline is [0,1]d[0,1]^d1, so nonconstant coefficients act as cumulant-like perturbations from uniformity. The same logic underlies Hierarchical Correlation Reconstruction for conditional densities, where

[0,1]d[0,1]^d2

and each coefficient [0,1]d[0,1]^d3 is predicted separately by MSE regression, optionally with CCA and lasso (Duda et al., 2018, Duda, 2022).

These constructions differ in algebraic form, but they share a common structure: a low-order statistical summary is first converted into a function-space representation, and reconstruction is then obtained by a variational principle, orthogonal projection, or coefficient-wise regression.

3. Geometric and measure-theoretic inversion

A second major paradigm reconstructs distributions from geometric summaries. For Tukey’s halfspace depth,

[0,1]d[0,1]^d4

full reconstruction fails in general: distinct measures can have the same halfspace depth function everywhere in [0,1]d[0,1]^d5. The paper nevertheless proves three kinds of recoverable information. First, the support is constrained by

[0,1]d[0,1]^d6

Second, an atom [0,1]d[0,1]^d7 with [0,1]d[0,1]^d8 must be an extreme point of every [0,1]d[0,1]^d9 for p(yx)p(y\mid x)0. Third, atoms create directional jumps in depth along lines passing through them. The framework relies crucially on minimizing flag halfspaces, which always exist even when minimizing ordinary halfspaces do not (Laketa et al., 2022).

By contrast, the geometric or spatial distribution function

p(yx)p(y\mid x)1

admits an explicit inversion. The key operator is

p(yx)p(y\mid x)2

and the reconstruction theorem states that

p(yx)p(y\mid x)3

in p(yx)p(y\mid x)4 for arbitrary Borel probability measures. For sufficiently regular densities, the same identity holds pointwise. The inversion is local in odd dimension and nonlocal in even dimension, because p(yx)p(y\mid x)5 is an integer in the first case and a half-integer in the second (Konen, 2022).

A third formulation places the inverse problem directly on probability measure space. Given p(yx)p(y\mid x)6, the overdetermined problem is posed as

p(yx)p(y\mid x)7

while the underdetermined problem is

p(yx)p(y\mid x)8

The paper shows that the choice of p(yx)p(y\mid x)9 or ρx\rho_x0 changes the meaning of “reconstruction.” With a ρx\rho_x1-divergence, the output reconstruction is the conditional distribution of ρx\rho_x2 on the attainable range ρx\rho_x3; with Wasserstein distance, it is the marginal distribution obtained by metric projection onto ρx\rho_x4. In the underdetermined case, entropy minimization yields a piecewise constant, fiber-uniform reconstruction, whereas minimizing the second moment yields the least-norm solution (Li et al., 26 Apr 2025).

4. Indirect observables, controls, and reconstructed latent laws

Many reconstruction problems do not observe the target distribution directly at all. Instead, the distribution is latent in a dynamical or experimental forward model. In the manifold-based formulation of inverse problems over the probability simplex, the observables are ρx\rho_x5, the data are ρx\rho_x6, and reconstruction minimizes a mismatch functional ρx\rho_x7. Because the feasible set is the interior of the simplex rather than a Euclidean space, the paper derives the manifold-based gradient descent rule

ρx\rho_x8

equivalently

ρx\rho_x9

In the protein-ensemble reconstruction example, this manifold-aware update recovers both the pseudo-experimental SAXS data and the pseudo-true distribution when the step size is sufficiently small (Oroguchi et al., 2024).

In NMR spin ensembles, the unknown object is a discrete approximation of the joint distribution of p(yx)p(y\mid x)0. The measurement model is linear in the probabilities,

p(yx)p(y\mid x)1

and identifiability is controlled by a control-dependent Gram matrix p(yx)p(y\mid x)2. The greedy reconstruction algorithm chooses controls so that p(yx)p(y\mid x)3 is positive definite, hence the constrained least-squares reconstruction is uniquely solvable. The later SPIRED implementation extends this program to inhomogeneities in all spatial directions and provides the corresponding code for GRA and OGRA (Buchwald et al., 2021, Buchwald et al., 2023).

Probability distribution reconstruction by circuit cutting targets yet another latent law: the bitstring distribution of a variational quantum circuit after the circuit has been decomposed into fragments. For one cut, the paper reconstructs the output probability p(yx)p(y\mid x)4 from weighted sums of empirical fragment probabilities p(yx)p(y\mid x)5; for a cut CZ gate, six circuits are run and post-selection produces ten cases. The sampling guarantee is of Hoeffding type, and the method is compared under “train-then-cut” and “cut-then-train.” In the 4-qubit classifier example with 4 cut CNOT gates and 24 trainable parameters, a full parameter-shift gradient in the cut-then-fit regime requires

p(yx)p(y\mid x)6

circuit evaluations (Neumann et al., 3 Oct 2025).

Two further examples sharpen the same pattern. In pulse-stream reconstruction, the distribution of short sample trains along a curve in p(yx)p(y\mid x)7 determines the underlying pulse uniquely when the curve is regularly parametrized, so the latent object is a signal recovered from the probability law of local observations (Rupniewski, 2023). In stochastic gravitational-wave backgrounds, the target is the full 1-point PDF p(yx)p(y\mid x)8 of strain fluctuations. There the reconstruction is performed by inverting a compound-Poisson characteristic function, not by estimating a mean or a power spectrum alone (Ginat et al., 2019).

5. Distribution-to-distribution regression and local probabilistic representations

A different branch of the literature reconstructs one distribution from another by learning a map between distributions. In structural health monitoring, the observed input is a density p(yx)p(y\mid x)9, the missing target is ρy=G#ρx\rho_y=G_\#\rho_x0, and the training data are paired PDFs ρy=G#ρx\rho_y=G_\#\rho_x1. The method first applies the log-quantile-density transform

ρy=G#ρx\rho_y=G_\#\rho_x2

then represents the transformed response by FPCA scores, and finally fits a function-to-vector RKHS regression

ρy=G#ρx\rho_y=G_\#\rho_x3

The predicted score vector reconstructs ρy=G#ρx\rho_y=G_\#\rho_x4, which is mapped back to density space by the inverse LQD transform. In general cases this LQD-RKHS method outperforms the conventional distribution-to-distribution regression and the distribution-to-warping function regression, but in extrapolation cases it is inferior to the latter (Chen et al., 2018).

Local probabilistic reconstruction also appears in hyperspectral anomaly detection. There, each pixel is represented by a latent multivariate Gaussian distribution via a VAE, and a local expected distribution is estimated from a Chebyshev neighborhood. The anomaly score is the modified Wasserstein distance between the pixel’s own latent Gaussian and the neighborhood average expectation. This is not full density reconstruction in the nonparametric sense; it is a local probabilistic representation and neighborhood-based distribution comparison (Yu et al., 2021).

The same local viewpoint is extended to weighted histopolation on triangular meshes. Edge measurements are interpreted as weighted moments under probability densities ρy=G#ρx\rho_y=G_\#\rho_x5 on ρy=G#ρx\rho_y=G_\#\rho_x6, and orthogonal polynomials associated with ρy=G#ρx\rho_y=G_\#\rho_x7 generate unisolvent triples for reconstruction in ρy=G#ρx\rho_y=G_\#\rho_x8. The paper proves a general framework for arbitrary polynomial order ρy=G#ρx\rho_y=G_\#\rho_x9, with Jacobi-type and Gegenbauer densities as explicit examples, and uses an adaptive parameter selection algorithm to minimize the global reconstruction error (Milovanovic et al., 11 Nov 2025).

6. Identifiability, regularity, and recurring limitations

The literature repeatedly emphasizes that reconstruction is rarely exact without structural assumptions. Finitely many moments do not uniquely determine a distribution in general, which is why maximum entropy is used as a selection principle in both the Rice–Mele and discrete ρx\rho_x0 settings (Yahyavi et al., 2017, Andreychenko et al., 2014). Halfspace depth does not determine a multivariate measure completely; at best it yields support restrictions, persistent extreme points, and atom-detection criteria (Laketa et al., 2022). In measure-space inversion, even the meaning of the inverse depends on whether the problem is overdetermined or underdetermined and on whether one chooses a ρx\rho_x1-divergence, Wasserstein distance, entropy, or second moment (Li et al., 26 Apr 2025).

Regularity is equally decisive. The PDE inversion of the geometric cdf is exact in ρx\rho_x2, but a continuous density need not generate a geometric cdf with enough regularity to reconstruct the density pointwise. The paper’s odd/even dichotomy makes this sharper: the inverse operator is local in odd dimension and completely nonlocal in even dimension (Konen, 2022). In circuit cutting, reconstructing the full output distribution becomes infeasible when most probabilities are exponentially small, so the method is practical only when a small number of states carry significant mass (Neumann et al., 3 Oct 2025).

Several methods also trade exact probabilistic validity for tractability. Polynomial density expansions on ρx\rho_x3 are not guaranteed nonnegative, and the yield-curve paper explicitly discusses negative densities as artifacts of the method (Duda et al., 2018). Hierarchical Correlation Reconstruction likewise produces a raw density-like function that must be passed through a positive calibration map and renormalized (Duda, 2022). On trees, exact recursive propagation of the residual distribution is infeasible because the support grows exponentially with depth, so the survey formalism keeps only a bounded-support compressed representation that still upper-bounds expected convex functionals of the true posterior law (0903.4812).

Taken together, these limitations show that probability distribution reconstruction is governed by three persistent issues: identifiability of the target law, regularity of the forward summary used for inversion, and numerical control of approximation error when the exact inverse is high-dimensional, nonlocal, or combinatorial.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (17)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Probability Distribution Reconstruction.