---
title: Window-Deconvolved Quasi-Optimal Estimators
url: https://www.emergentmind.com/topics/window-deconvolved-quasi-optimal-estimators
type: topic
---

# Window-Deconvolved Quasi-Optimal Estimators

Searching arXiv for the specified papers and closely related work to ground the article in current literature.
arxiv_search
{"query":"id:1010.0289 OR id:2107.06287 OR id:2404.07249 OR id:2012.09389 OR id:1308.3551", "max_results": 10, "sort_by": "submittedDate", "sort_order": "descending"}
arxiv_search
{"query":"\"Deconvolution of window effect in galaxy power spectrum analysis\" OR \"Cosmology Without Windows: Quadratic Estimators for the Galaxy Power Spectrum\" OR \"Cosmology Without Window Functions: Cubic Estimators for the Galaxy Bispectrum\" OR \"PolyBin3D: A Suite of Optimal and Efficient Power Spectrum and Bispectrum Estimators for Large-Scale Structure\"", "max_results": 10, "sort_by": "relevance", "sort_order": "descending"}
Window-deconvolved quasi-optimal estimators are estimators for cosmological two-, three-, and four-point statistics that remove the effects of survey masks, selection functions, beams, or pixellization from the measurement itself, so that the reported quantity can be compared directly to unconvolved theory. In the literature surveyed here, this is realized in two closely related ways: by explicit deconvolution using the Fourier deconvolution theorem for galaxy power-spectrum multipoles, and by quadratic, cubic, or quartic maximum-likelihood constructions in which a Fisher normalization carries the window response and its inverse de-mixes the estimator output. “Quasi-optimal” denotes the use of computationally tractable approximations to inverse-covariance weighting, most commonly FKP-type weights, which remain unbiased and are close to minimum variance under the relevant Gaussian, weakly non-Gaussian, or distant-observer assumptions [1010.0289] [2012.09389] [2107.06287] [2404.07249] [2502.04434].

## 1. Conceptual scope and historical development

The modern literature uses closely related estimator architectures across several domains of large-scale-structure and CMB analysis. In redshift surveys, Sato, Hütsi, and Yamamoto developed a deconvolution method for galaxy multipole power spectra based on the deconvolution theorem and compatible with fast Fourier transforms [1010.0289]. Sato, Hütsi, Nakamura, and Yamamoto then analyzed the same problem in both convolved and deconvolved forms, emphasizing multipole mixing kernels and the practical treatment of survey windows in SDSS LRG DR7 [1308.3551]. Philcox reformulated the power-spectrum problem as a quadratic unwindowed estimator [2012.09389], and later generalized the strategy to the galaxy bispectrum with cubic estimators and Fisher deconvolution [2107.06287]. PolyBin3D places these power-spectrum and bispectrum constructions in a common maximum-likelihood framework with FFT and Monte Carlo implementations [2404.07249], while analogous ideas have been extended to mask-deconvolved quasi-optimal CMB trispectrum estimators [2502.04434].

| Domain | Estimator class | Window removal mechanism |
|---|---|---|
| Galaxy power spectrum | Fourier deconvolution; quadratic estimators | Configuration-space ratio or Fisher inversion [1010.0289] [2012.09389] |
| Galaxy bispectrum | Cubic estimators | Fisher matrix deconvolution [2107.06287] |
| Joint LSS power/bispectrum pipelines | Maximum-likelihood “unwindowed” estimators | Fisher normalization with linear filter \(S\) [2404.07249] |
| CMB trispectrum | Quartic template estimators | \(S=P^\dagger C^{-1}\) weighting and deconvolved Fisher normalization [2502.04434] |

A common misconception is that “window deconvolution” refers only to explicit inversion of a precomputed mode-coupling matrix. The literature is broader. In the Sato et al. power-spectrum construction, deconvolution is performed by dividing inverse Fourier transforms in configuration space and then transforming back [1010.0289]. In the Philcox and PolyBin3D formalisms, the deconvolution is encoded in the Fisher matrix: the estimator numerator is measured on masked data, and \(F^{-1}\) removes the window-induced mixing [2012.09389] [2404.07249].

## 2. Power-spectrum formulation in galaxy redshift surveys

The redshift-survey formulations begin from the FKP galaxy-minus-random field. In the notation of Sato, Hütsi, and Yamamoto,
\[
F(\mathbf{x})=\psi(\mathbf{x})[n_g(\mathbf{x})-\alpha n_{\rm rnd}(\mathbf{x})],
\]
with \(\alpha \ll 1\), \(\psi(\mathbf{x})\) a real-space weight, and \(n_{\rm rnd}\) a random catalog tracing the selection function. The Fourier-space field is
\[
F(\mathbf{k})=\int d^3x\,\psi(\mathbf{x})[n_g(\mathbf{x})-\alpha n_{\rm rnd}(\mathbf{x})]e^{i\mathbf{k}\cdot\mathbf{x}},
\]
and the convolved power-spectrum estimator is
\[
P^{\rm conv}(\mathbf{k})=\frac{|F(\mathbf{k})|^2}{\int d^3x\,\bar n^2(\mathbf{x})\psi^2(\mathbf{x})}-(1+\alpha)S_0,
\]
with
\[
S_0=\frac{\int d^3x\,\bar n(\mathbf{x})\psi^2(\mathbf{x})}{\int d^3x\,\bar n^2(\mathbf{x})\psi^2(\mathbf{x})}.
\]
The corresponding window estimator is
\[
W(\mathbf{k})=\frac{\left|\int d^3x\,\alpha n_{\rm rnd}(\mathbf{x})\psi(\mathbf{x})e^{i\mathbf{k}\cdot\mathbf{x}}\right|^2}{\int d^3x\,\bar n^2(\mathbf{x})\psi^2(\mathbf{x})}-\alpha S_0.
\]
In the ideal FKP limit, the measured spectrum obeys the convolution relation
\[
P^{\rm conv}(\mathbf{k})=\frac{1}{(2\pi)^3}\int d^3k'\,P(\mathbf{k}')W(\mathbf{k}-\mathbf{k}').
\]
This is the fundamental window-smearing relation that the deconvolved estimator seeks to undo [1010.0289].

For anisotropic redshift-space clustering, the power depends on \(\mu=\hat{\mathbf{k}}\cdot\hat{\mathbf{e}}\), and the deconvolved multipoles are computed after reconstruction of the full three-dimensional \(P(\mathbf{k})\). Under the distant-observer approximation used for FFT compatibility,
\[
P_\ell(k)=\frac{1}{V_k}\int_{V_k} d^3k\,P(\mathbf{k})L_\ell(\hat{\mathbf{e}}\cdot\hat{\mathbf{k}}),
\]
with the note that this normalization differs from the conventional one by a factor \(2\ell+1\) [1010.0289]. The companion analysis of the window effect also writes the convolved multipoles in kernel form,
\[
P_\ell^{\rm conv}(k)=\sum_{\ell'}\int_0^\infty dk'\,k'^2 M_{\ell\ell'}(k,k')P_{\ell'}(k'),
\]
where the coupling matrix is determined by the multipole moments of the survey window [1308.3551].

A central distinction in the literature is therefore methodological rather than conceptual. One may keep the measurement convolved and convolve theory with \(M_{\ell\ell'}\), or one may deconvolve the measurement and compare it directly to theory. The former is the standard FKP route; the latter defines the window-deconvolved estimator [1308.3551].

## 3. Deconvolution and quasi-optimality

In the Sato et al. power-spectrum method, the key step is the deconvolution theorem. Starting from the convolution relation, the true spectrum is reconstructed as
\[
P(\mathbf{k})=\int d^3s\,e^{i\mathbf{k}\cdot\mathbf{s}}
\frac{\int d^3k'\,e^{-i\mathbf{k}'\cdot\mathbf{s}}P^{\rm conv}(\mathbf{k}')}{\int d^3k''\,e^{-i\mathbf{k}''\cdot\mathbf{s}}W(\mathbf{k}'')}.
\]
Operationally, the inverse transforms of \(P^{\rm conv}\) and \(W\) are divided pointwise in configuration space and then Fourier transformed back. This yields a deconvolved \(P(\mathbf{k})\) before multipole projection [1010.0289].

The term “quasi-optimal” enters through weighting. In the original FKP derivation, the weight
\[
w(\mathbf{x})=\frac{1}{1+\bar n(\mathbf{x})P_0}
\]
minimizes the variance of spherically averaged band-powers for Gaussian fields and slowly varying \(\bar n(\mathbf{x})\). Sato et al. allow such weights within the deconvolution pipeline, though their SDSS LRG DR7 application sets \(\psi=1\) [1010.0289]. Sato et al. also describe the FKP weight as “quasi-optimal” under Gaussianity, slowly varying \(P(k)\), and distant-observer assumptions [1308.3551].

Philcox recast this logic in a quadratic-estimator language. For an arbitrary symmetric, positive-definite pixel weight \(H\),
\[
\hat q_\alpha^{\rm QE}=\frac{1}{2}d^T H C_{,\alpha} H d,\qquad
F_{\alpha\beta}^{\rm QE}=\frac{1}{2}{\rm Tr}[H C_{,\alpha} H C_{,\beta}],
\]
with unbiased estimator
\[
\hat p_\alpha^{\rm QE}=p_\alpha^{\rm fid}+\sum_\beta (F^{-1,{\rm QE}})_{\alpha\beta}(\hat q_\beta^{\rm QE}-\bar q_\beta^{\rm QE}).
\]
For \(H=C_{\rm fid}^{-1}\), this becomes the maximum-likelihood estimator, which saturates the Cramér–Rao bound in the Gaussian limit when \(C_{\rm fid}=C_D\). Replacing \(H\) by an FKP-like diagonal approximation yields the quasi-optimal FKP-based estimator [2012.09389].

The same structure persists at higher order. For the bispectrum, the general quasi-optimal estimator is
\[
\hat b_\alpha=\sum_\beta F^{-1}_{\alpha\beta}\hat q_\beta,\qquad
\hat q_\alpha=\frac{1}{6}B^{ijk}_{,\alpha}[Hd]_i\left([Hd]_j[Hd]_k-3H_{jk}\right),
\]
with
\[
F_{\alpha\beta}=\frac{1}{6}B^{ijk}_{,\alpha}B^{lmn}_{,\beta}H_{il}H_{jm}H_{kn}.
\]
With inverse-covariance weights this is minimum variance in the weakly non-Gaussian limit; with FKP weights it is “close-to-optimal and easy to compute” [2107.06287]. PolyBin3D generalizes this to \(n\)-point functions by expressing the estimator as a numerator, bias term, and Fisher normalization for a general linear filter \(S\), with the window entering through \(P\) and \(N\) inside the Fisher and bias terms [2404.07249].

## 4. Numerical realization and algorithmic structure

The original FFT-compatible power-spectrum deconvolution pipeline is explicit. One builds random catalogs with \(\alpha=0.01\), chooses \(\psi(\mathbf{x})\), places galaxies and randoms on an FFT grid, estimates \(P^{\rm conv}(\mathbf{k})\) and \(W(\mathbf{k})\), computes the inverse FFTs
\[
N(\mathbf{s})\equiv \int d^3k'\,e^{-i\mathbf{k}'\cdot\mathbf{s}}P^{\rm conv}(\mathbf{k}'),
\qquad
D(\mathbf{s})\equiv \int d^3k''\,e^{-i\mathbf{k}''\cdot\mathbf{s}}W(\mathbf{k}''),
\]
forms \(R(\mathbf{s})=N(\mathbf{s})/D(\mathbf{s})\), and FFTs back to obtain \(P(\mathbf{k})\). The complexity is dominated by 3D FFTs, with each forward or inverse FFT scaling as \(\mathcal{O}(N_{\rm grid}\log N_{\rm grid})\) [1010.0289].

The quadratic-estimator implementation replaces explicit deconvolution of \(P^{\rm conv}\) by repeated applications of covariance operators. In the maximum-likelihood power-spectrum case, \(C_{\rm fid}^{-1}\) is applied with preconditioned conjugate gradient descent, using FFT representations of \(C\) and \(C_{,\alpha}\). One application of \(C\) requires \( \mathcal{O}\!\left(\frac{1}{2}(\ell_{\max}+1)(\ell_{\max}+2)\right)\) FFTs, one application of \(C_{,\alpha}\) requires \(\mathcal{O}(2\ell+1)\) FFTs, and convergence in \(N_{\rm it}\approx 50\) iterations is typical. Monte Carlo estimation of Fisher matrices and biases adds only a \(\sqrt{1+1/N_{\rm mc}}\) error-bar penalty, quoted as \(\approx 0.5\%\) for \(N_{\rm mc}=100\) [2012.09389].

For the bispectrum, the practical implementation relies on separable templates and FFT-based filtered maps. In the FKP-weighted form, the data bispectrum across \(\approx 10^3\) mocks requires \(\approx 100\) CPU-hours, while the Fisher matrix and Monte Carlo corrections are data-independent and require \(\approx \mathcal{O}(150)\) CPU-hours for FKP weighting or \(\approx \mathcal{O}(850)\) CPU-hours for ML weighting, with percent-level accuracy for \(N_{\rm mc}\approx 50\) [2107.06287]. PolyBin3D reports Monte Carlo convergence for masked bispectra using only a small number of iterations, “typically \(<10\) for bispectra,” and supports GPU acceleration using JAX, with speed-ups of \(\sim 5\times\) versus 64-core CPU nodes for representative power-spectrum numerators [2404.07249].

These implementations are unified by the same computational principle: shift the expensive mask treatment into precomputable normalization objects, then use FFT-amenable filtered maps for the data-dependent part.

## 5. Covariance structure, empirical performance, and survey demonstrations

The principal empirical motivation for deconvolution is covariance simplification. In the SDSS LRG DR7 study of Sato, Hütsi, and Yamamoto, 1000 LRG-like mocks show that convolved spectra exhibit strong off-diagonal correlations, especially for narrower subsamples, whereas the correlation matrices of deconvolved spectra are “practically diagonal” for \(\ell=0\), \(\ell=2\), and the cross-covariances \(r_{02}\). Deconvolution also corrects amplitudes measured from different survey divisions, and the BAO signature in deconvolved \(P_0(k)\) and \(P_2(k)\) shows larger peak–trough amplitudes and closely matches the “ideal” catalogs, particularly for subsamples where window smoothing is otherwise significant [1010.0289].

The SDSS application used the LRG DR7 sample with \(z\in[0.16,0.47]\), sky coverage \(7150\,{\rm deg}^2\), and \(N=100{,}157\) LRGs, with distances computed assuming flat \(\Lambda{\rm CDM}\) with \(\Omega_m=0.28\), \(\psi=1\), and angular subdivision into 18 subsamples of \(\approx 397\,{\rm deg}^2\) or 32 subsamples of mean \(\approx 223\,{\rm deg}^2\) so that a single line of sight \(\hat{\mathbf e}\) could be assigned per subsample [1010.0289]. The companion window-analysis paper further shows that the convolved monopole and quadrupole amplitudes decrease with decreasing patch size, and that correction factors derived from the measured window can restore consistent amplitudes across different patch divisions [1308.3551].

In the BOSS DR12 power-spectrum study, Philcox found that unwindowed band-powers from ML and FKP estimators are statistically consistent and have similar error bars for a low-density sample with a compact window. The unwindowed covariance is nearly diagonal with slight nearest-bin anti-correlations, whereas windowed estimates have smaller per-bin variance at high \(k\) because the window induces bin correlations. Compressed coefficients and cosmological posteriors were likewise statistically consistent with conventional windowed analyses, with shifts \(<0.8\sigma\) [2012.09389].

PolyBin3D extends these comparisons. For power spectra, unwindowed estimators recover the ideal unmasked spectrum across multipoles, while windowed FKP estimators show \(\mathcal{O}(\sigma)\) distortions on large scales, particularly in the quadrupole. For bispectra, unwindowed estimators recover the injected bispectrum within errors, while windowed estimators exhibit significant large-scale biases and strong bin-to-bin correlations. The use of optimal CGD weights reduces large-scale power-spectrum error bars by \(5\)–\(10\%\) relative to FKP in realistic tests [2404.07249].

## 6. Generalization to higher-order statistics and other observables

The bispectrum generalization makes the logic of window-deconvolved quasi-optimal estimation especially transparent. The observed bispectrum of the masked field involves a three-leg window convolution,
\[
B(\mathbf{k}_1,\mathbf{k}_2,\mathbf{k}_3)\rightarrow B_W(\mathbf{k}_1,\mathbf{k}_2,\mathbf{k}_3),
\]
which is a six-dimensional integral at every triangle configuration and MCMC step. Philcox’s cubic estimator avoids this by constructing an estimator for the unwindowed bispectrum,
\[
\hat{\mathbf b}=F^{-1}\mathbf E,
\]
where the cubic data statistic \(\mathbf E\) is measured directly from the masked survey and the Fisher matrix deconvolves the window. In the limit of weak non-Gaussianity, the inverse-covariance-weighted version is minimum variance; the FKP-weighted variant is close-to-optimal and easy to compute [2107.06287].

PolyBin3D recasts both power-spectrum and bispectrum estimation in a single maximum-likelihood formalism. For general \(n\)-point functions,
\[
\hat x_\alpha=\sum_\beta F^{-1}_{\alpha\beta}\left[x^{\rm num}_\beta-x^{\rm bias}_\beta\right],
\]
with the window entering only through the pointing operator \(P\) and the noise \(N\) inside the Fisher and bias terms. Because \(F\) carries the mask dependence via \([SP]\), applying \(F^{-1}\) de-mixes the window leakage and returns unwindowed band-powers or bispectrum coefficients directly comparable to theory [2404.07249].

An analogous strategy has been developed for the CMB trispectrum. There, observed maps are filtered with
\[
S_{\rm opt}=P^\dagger C^{-1},
\]
where \(P=WYB\) contains the mask, spherical-harmonic synthesis, and beam, and a quartic Edgeworth-derived estimator measures template amplitudes \(\hat A_\alpha=\sum_\beta F^{-1}_{\alpha\beta}\hat N_\beta\). The disconnected terms subtract the cut-sky mean field and Gaussian biases, while the same \(SP\) factors in the Fisher normalization provide exact window and beam deconvolution. The resulting estimators are described as unbiased, minimum variance, mask-deconvolved, and able to account for correlations between templates, including lensing and point sources [2502.04434].

A plausible implication is that “window-deconvolved quasi-optimal estimator” is best understood as a general estimation paradigm rather than a single algorithm: window treatment is transferred from forward-model convolution of theory to estimator normalization and filtering, with optimality controlled by how closely the adopted weighting approximates \(C^{-1}\).

## 7. Limitations, assumptions, and methodological trade-offs

The principal numerical vulnerability of explicit deconvolution is division by a small window transform. In the Sato et al. implementation, \(D(\mathbf{s})\) must not vanish; the FFT box should be chosen to match the survey footprint so that \(W(\mathbf{k})\) does not approach zero at problematic modes. If \(D(\mathbf{s})\) is small, the deconvolution can amplify noise, and while the paper does not introduce explicit regularization, it notes that masking low-\(D(\mathbf{s})\) cells, band-limiting, or imposing a floor would help [1010.0289]. The 2011 follow-up similarly emphasizes that deconvolution can amplify noise where the window suppresses signal strongly and may require regularization or restricted \(k\)-ranges [1308.3551].

Approximate optimality also has a domain of validity. FKP weights are justified when the field is Gaussian, \(P(k)\) is slowly varying, and one is content with near-minimum variance rather than exact inverse-covariance weighting [1308.3551]. In the quadratic-estimator framework, unbiasedness does not require Gaussianity, but optimality of the ML form does [2012.09389]. For the bispectrum and trispectrum, the Edgeworth expansions and Cramér–Rao statements are explicitly tied to weak non-Gaussianity [2107.06287] [2502.04434].

Geometric approximations matter as well. The redshift-space multipole implementations typically assume a distant-observer or Yamamoto-like line-of-sight prescription. In the SDSS LRG deconvolution study, a small low-\(k\) deviation in the quadrupole is attributed to limitations of the distant-observer approximation in finite-area subsamples [1010.0289]. The 2011 treatment therefore advocates patching the survey so that the line of sight is approximately constant within each patch [1308.3551].

Finally, there is a trade-off between robustness and formal optimality. Philcox notes that maximum-likelihood weighting is more exact but computationally heavier, whereas FKP weighting is quasi-optimal and often sufficient for compact, low-density surveys such as BOSS DR12 [2012.09389]. The bispectrum and PolyBin3D studies echo this balance: ML weighting can be numerically delicate in masked regions or for fine binning, while FKP weighting sacrifices a small amount of optimality for robustness, speed, and simpler implementations [2107.06287] [2404.07249].

Within those assumptions, window-deconvolved quasi-optimal estimators provide a coherent route to unbiased, nearly decorrelated cosmological summary statistics whose normalization absorbs survey geometry rather than forcing it into every theory evaluation.

Source: https://www.emergentmind.com/topics/window-deconvolved-quasi-optimal-estimators