---
title: '21cmPSDenoiser: Denoising in 21-cm Cosmology'
url: https://www.emergentmind.com/topics/21cmpsdenoiser
type: topic
---

# 21cmPSDenoiser: Denoising in 21-cm Cosmology

Searching arXiv for the cited papers and closely related work on 21-cm denoising / foreground mitigation.
21cmPSDenoiser denotes, in the provided literature, two technically distinct denoising pipelines for 21-cm cosmology. In its explicit and primary recent usage, it is a score-based diffusion model that takes a single, forward-modelled realisation of the cylindrical \(2\)D 21-cm power spectrum and predicts the corresponding population mean during Bayesian inference, with the stated aim of mitigating sample variance in reionisation-era analyses [2507.12545]. In the same literature package, the label is also applied to the earlier RPCA + GILC pipeline for 21-cm signal recovery from foreground-contaminated intensity-mapping data, where denoising is formulated as a decomposition of the frequency–frequency covariance into a low-rank foreground term and a sparse HI term [1801.04082]. The shared designation therefore refers to a family resemblance at the level of inverse-problem structure rather than to a single invariant algorithm.

## 1. Definition, target quantity, and nomenclature

The 2025 formulation introduces \texttt{21cmPSDenoiser} as a score-based diffusion model for denoising a single, forward-modelled realisation of the \(21\)-cm \(2\)D power spectrum, with the output interpreted as the population-mean cylindrical power spectrum. The motivating problem is that state-of-the-art simulations of reionisation-era \(21\)-cm signal have limited volumes, generally orders of magnitude smaller than observations, so the Fourier modes in common between simulation and observation have limited overlap, especially in cylindrical \(2\)D \(k\)-space natural for \(21\)-cm interferometry. In that setting, sample variance is treated as the dominant stochastic contaminant to be removed from a single realisation [2507.12545].

A recurrent misconception is to assimilate the method to a conventional emulator. The paper explicitly distinguishes it from emulators by stating that the denoiser is not tied to a particular model or simulator, since its input is a model-agnostic realisation of the \(2\)D \(21\)-cm power spectrum. It further reports that the model generalises to power spectra produced with a different \(21\)-cm simulator than those on which it was trained [2507.12545].

The label is, however, not unique to this diffusion-based setting in the provided material. The same name is attached there to the RPCA + GILC workflow summarised from the 2018 foreground-subtraction study. That earlier method addresses a different inverse problem: recovering the cosmological \(21\)-cm signal from intensity-mapping data dominated by coherent foreground radiation from the Milky Way and extragalactic radio sources. This suggests a useful terminological distinction between **power-spectrum sample-variance denoising** and **covariance-structured signal recovery**, even though both are described as denoising.

## 2. Diffusion-theoretic formulation

The theoretical framework in the 2025 method is a score-based diffusion model in which sample variance in a single \(2\)D power-spectrum realisation is viewed as “noise” to be removed in order to recover the underlying population-mean power spectrum. The construction proceeds by specifying a forward stochastic diffusion process that gradually corrupts the clean mean power spectrum into Gaussian white noise, and then learning the time-dependent score function that enables reversal of that process [2507.12545].

Let \(x_0 \in \mathbb{R}^D\) denote a clean mean \(2\)D power spectrum, and let \(x(t)\) be its diffused counterpart at time \(t \in [0,T]\). The forward variance-preserving SDE is
\[
\mathrm d x = f(x,t)\,\mathrm dt + g(t)\,\mathrm d w,
\]
with
\[
f(x,t) = -\tfrac12 \beta(t)\,x,
\qquad
g(t) = \sqrt{\beta(t)},
\]
where \(\beta(t)\) is a prescribed noise schedule and \(w\) is a standard Wiener process. As \(t \to T\), \(x(T) \sim \mathcal N(0,\mathbf I)\).

The reverse-time dynamics are written as
\[
\mathrm d x =
\bigl[f(x,t)-g^2(t)\,\nabla_x \log p_t(x)\bigr]\mathrm dt
+ g(t)\,\mathrm d \bar w,
\]
where \(p_t(x)\) is the marginal density at time \(t\). The equivalent probability-flow ODE,
\[
\mathrm d x =
\bigl[f(x,t)-\tfrac12 g^2(t)\,\nabla_x \log p_t(x)\bigr]\mathrm dt,
\]
is deterministic and has the same marginal laws. The score is parametrised by a neural network \(s_\theta(x,t)\approx \nabla_x \log p_t(x)\), trained with a continuous-time denoising score-matching objective. One form given in the summary is
\[
\mathcal L(\theta)
=
\mathbb{E}_{t\sim\mathcal U[0,T],\,x_0\sim p_{\rm data},\,\varepsilon\sim\mathcal N(0,\mathbf I)}
\Bigl[
\lambda(t)\,
\bigl\lVert
s_\theta\bigl(x(t),t\bigr)
-\nabla_x\log p_{0t}\bigl(x(t)\mid x_0\bigr)
\bigr\rVert^2
\Bigr],
\]
with the equivalent form
\[
\mathcal L(\theta)
=
\mathbb{E}_{x_0,t,\varepsilon}
\Bigl[
\bigl\lVert
s_\theta(x(t),t)
+\tfrac{\varepsilon}{\sigma(t)}
\bigr\rVert^2
\Bigr].
\]

In this formulation, the denoiser is not estimating instrument noise directly; it is learning the score of the diffused distribution over mean power spectra so that a noisy single-realisation input can be transported toward the corresponding population mean. That distinction is central to interpreting the method’s scope.

## 3. Network architecture, preprocessing, and training corpus

The neural backbone is a U-Net convolutional autoencoder with time- and noise-level conditioning. Its input is a single noisy \(2\)D power-spectrum realisation \(x(t)\) together with the scalar diffusion time \(t\), and its output is the estimated score \(s_\theta(x,t)\) [2507.12545].

The reported hyperparameters are: diffusion time horizon \(T \approx 1\); effectively infinite numbers of training timesteps because the loss is continuous-time; Adam optimizer with learning rate \(10^{-4}\); batch size \(\simeq 256\) \(2\)D power-spectrum patches; and weight decay \(10^{-6}\). Training is stated to converge in \(\sim 200\) epochs on \(3.6\) M samples [2507.12545].

The training database is derived from \(21\)cmFAST v3 coeval boxes of side length \(300\) cMpc and cell size \(1.5\) cMpc. It spans \(\sim 900\) parameter combinations with \(\sim 100\) IC realisations each, for \(\simeq 90\,\mathrm{k}\) light-cone boxes. For each box, the pipeline slices into \(40\) redshift bins and computes cylindrical power spectra on \(32\) log-spaced \(k_\perp \times 16\) \(k_\parallel\) bins, giving \(\simeq 3.6\) M \(2\)D power spectra in total. Preprocessing consists of taking \(\log_{10}\Delta^2_{21}\) and applying min–max scaling to \([-1,1]\) [2507.12545].

The astrophysical training distribution is also constrained rather than purely ad hoc. The summary states that \((f_{\rm esc}, f_\star, M_{\rm turn})\) are drawn from a posterior conditioned on UV luminosity functions, CMB \(\tau_e\), and Ly\(\alpha\) dark fraction, while the X-ray parameters are flat in \(\log_{10}L_X/\mathrm{SFR}\in[38,42]\) and \(E_0\in[200,1500]\) eV. This means that the denoiser is trained on a broad but structured astrophysical ensemble rather than on a single fiducial history.

## 4. Embedding in Bayesian inference

The method is designed to be inserted into an explicit Gaussian-likelihood Bayesian inference pipeline, such as \(21\)CMMC, immediately after forward-modelling the \(2\)D power spectrum from a single simulation realisation. The resulting workflow can be summarised as follows [2507.12545]:

1. Propose \(\theta\) and generate one \(3\)D \(21\)-cm light cone, then compute the cylindrical \(2\)D power-spectrum realisation \(\Delta^2_{21,i}(k_\perp,k_\parallel,\theta)\).
2. Optionally crop to the experiment’s EoR window in \(\mu \equiv k_\parallel/k \ge \mu_{\min}\).
3. If the denoiser is used, compute
   \[
   \widehat\mu_{2D}(k_\perp,k_\parallel\mid\theta)
   =
   \mathrm{Denoiser}\bigl(\Delta^2_{21,i}\bigr);
   \]
   otherwise set \(\widehat\mu_{2D}=\Delta^2_{21,i}\).
4. Spherically average \(\widehat\mu_{2D}\to \widehat\mu_{1D}(k\mid\theta)\).
5. Apply the instrument window \(W\) to both model and data.
6. Evaluate the Gaussian likelihood
   \[
   \ln\mathcal L
   =
   -{\tfrac12}\sum_{k,z}
   \frac{\bigl[W\,\widehat\mu_{1D}(k,z\mid\theta)-W\,\Delta^2_{21,\rm mock}(k,z)\bigr]^2}
        {\sigma^2(k,z;\theta)},
   \]
   with total variance \(\sigma^2=\sigma^2_{\rm sens}+\sigma^2_{\rm fm}\).

Two choices are given for the forward-model variance. In the “Sample variance” baseline,
\[
\sigma_{\rm fm}=\widehat\mu_{1D}/\sqrt{N_{\rm modes}(k)},
\]
while with the denoiser,
\[
\sigma_{\rm fm}=\widehat\mu_{1D}\times \sigma_{\rm denoiser}(k,z),
\]
where \(\sigma_{\rm denoiser}\) is measured on the validation set. This is methodologically important: the denoiser changes both the mean model entering the likelihood and the treatment of forward-model uncertainty.

The computational-cost claim is similarly specific. Denoising one \(2\)D power spectrum requires \(\sim 6\) s on a V100 GPU, quoted as the median over \(200\) probability-flow ODE draws. Fixing & Pairing, by contrast, requires a second full simulation per likelihood step, i.e. \(\times 2\) cost, whereas \texttt{21cmPSDenoiser} adds \(<10\%\) overhead to a typical \(\sim 30\) s emulator-based likelihood [2507.12545].

## 5. Performance metrics and reported results

Performance is quantified with a fractional error metric
\[
\mathrm{FE}(k,z)\;[\%]
=
100\times
\frac{\bigl|\widehat\mu(k,z)-\mu_{\rm true}(k,z)\bigr|}
     {\max\bigl[0.01,\mu_{\rm true}(k,z)\bigr]},
\]
where \(\mu_{\rm true}\) is the mean power spectrum from \(\sim 200\) IC realisations at each \(\theta\) [2507.12545].

On the test set, after cropping to \(\mu \ge 0.97\), the reported summary statistics are:

| Method | Median FE | 68% CL |
|---|---:|---:|
| Baseline (single realisation) | \(\simeq 12.1\%\) | \(\simeq 30.9\%\) |
| Fix & Pair | \(\simeq 9.0\%\) | \(\simeq 25.0\%\) |
| 21cmPSDenoiser | \(\simeq 2.2\%\) | \(\simeq 7.8\%\) |

The same study states that individual samples of \(2\)D Fourier amplitudes of wave modes relevant to current \(21\)-cm observations can deviate from the mean by over \(50\%\) for \(300\) cMpc simulations, even when only considering stochasticity due to sampling of Gaussian initial conditions, and that \texttt{21cmPSDenoiser} reduces this deviation by an order of magnitude. It is further reported to outperform Fixing & Pairing by a factor of few at almost no additional computational cost [2507.12545].

For parameter inference, the HERA-mock experiment isolates the practical consequence of denoising. The traditional pipeline with no \(\mu_{\min}\) cut and a single realisation yields highly biased posteriors, described as \(\sim 10\,\sigma\) off and overconfident. Applying only the \(\mu_{\min}=0.97\) cut makes the posteriors unbiased but much wider because sample variance dominates. Combining the denoiser with \(\mu_{\min}=0.97\) yields unbiased posteriors with \(\sim 50\%\) narrower \(95\%\) intervals on each astrophysical parameter, and the abstract summarises this as an unbiased posterior that is \(50\%\) narrower [2507.12545].

## 6. Relation to the RPCA + GILC pipeline and broader methodological scope

The 2018 foreground-subtraction method provides a distinct denoising paradigm that is also associated with the name in the provided materials. There, the observation model for intensity mapping is
\[
x_i(p)=f_i(p)+s_i(p)+n_i(p),
\qquad
\mathbf{x}(p)=\mathbf{f}(p)+\mathbf{s}(p)+\mathbf{n}(p),
\]
with frequency–frequency covariance
\[
R=\frac{1}{N_p}\,\bigl\langle \mathbf{x}\,\mathbf{x}^T\bigr\rangle_p
\approx R_f+R_{\rm HI},
\]
where \(R_f\) is the low-rank foreground covariance and \(R_{\rm HI}\) is the sparse \(21\)-cm covariance. Renaming \(R\mapsto M\), the method seeks
\[
M=L+S,
\]
with \(L\approx R_f\) and \(S\approx R_{\rm HI}\) [1801.04082].

The ideal RPCA formulation is
\[
\min_{L,S}\;\bigl[\rank(L)+\lambda\|S\|_0\bigr]
\quad \text{s.t.}\quad M=L+S,
\]
and its convex relaxation, Principal Component Pursuit, is
\[
\min_{L,S}\;\|L\|_*+\lambda\|S\|_1
\quad \text{s.t.}\quad M=L+S.
\]
The augmented Lagrangian is solved with Inexact ALM, alternating singular-value thresholding for \(L\), soft-thresholding for \(S\), and dual updates for \(Y\), with stopping criterion
\[
\|M-L-S\|_F \le \delta \|M\|_F,
\qquad
\delta\sim 10^{-8}\text{--}10^{-14}.
\]
The universal choice
\[
\lambda=\frac{1}{\sqrt{\max(m,n)}}
\]
is stated to work well without tuning, with \(m=n=N_\nu\) in this application [1801.04082].

The significance of this earlier method lies in the structural assumptions it exploits: frequency coherence makes the foreground covariance low-rank, whereas the \(21\)-cm signal covariance is assumed sparse and dominant on or near the diagonal. The summary explicitly notes that sparsity of the frequency covariance of the \(21\)-cm signal is first explored there. In simulations over \(700\!-\!800\) MHz with \(256\) channels and \(13.7'\) pixels, the RPCA + GILC pipeline recovers \(R_{\rm HI}\) to relative error
\[
\Delta=\frac{\|S_{\rm recovered}-R_{\rm HI}\|_F}{\|R_{\rm HI}\|_F}
\sim 10^{-1}\text{--}10^{-2},
\]
stated as an order of magnitude better than classic PCA. Angular power spectra and cross-spectra coincide at the \(\sim 1\)–\(5\%\) level over \(\ell\sim 10^2\!-\!10^3\), the transfer function is within a few percent of unity, and the \(1\)D line-of-sight power spectrum is also well recovered, especially at high \(k_\parallel\) [1801.04082].

The distinction between the 2018 and 2025 pipelines is therefore substantive. The RPCA + GILC method denoises covariance structure in foreground-contaminated map-space data, whereas the diffusion model denoises sample variance in cylindrical power-spectrum space. Their assumptions and failure modes are correspondingly different. For the RPCA approach, violations of low-rank foreground covariance or sparse signal covariance—such as calibration artifacts or polarization leakage—may break the separation, motivating extensions like Stable PCP with \(M=L+S+Z\) and the constraint \(\|M-L-S\|_F\le\delta\), partial-SVD solvers for large \(N_\nu\), and GNILC for local foreground-rank estimation in space and scale [1801.04082]. For the diffusion model, the decisive assumptions are those implicit in the training distribution, the score-learning objective, and the use of validation-set network uncertainty within the likelihood [2507.12545].

Taken together, the two usages define a broader methodological pattern: both versions of 21cmPSDenoiser exploit strong structural priors to transform a single contaminated or stochastic realisation into an estimator of an underlying cosmological quantity, but they do so on different observables and against different sources of error.

Source: https://www.emergentmind.com/topics/21cmpsdenoiser