---
title: '21-cm Forest: Probing Early Universe Structures'
url: https://www.emergentmind.com/topics/21-cm-forest
type: topic
---

# 21-cm Forest: Probing Early Universe Structures

The 21-cm forest is the ensemble of narrow absorption features produced when neutral hydrogen along a line of sight imprints the redshifted hyperfine transition on the spectra of distant, compact, radio-bright background sources during Cosmic Dawn and the Epoch of Reionization. In contrast to the global 21-cm signal and 3D 21-cm tomography, which probe diffuse brightness against the CMB, the forest is a pencil-beam observable that resolves small-scale neutral structure in frequency space and is therefore acutely sensitive to the thermal state of the neutral intergalactic medium, the abundance of minihalos and filamentary absorbers, and any physics that modifies structure formation below galactic scales [2606.24665][2407.14298][2006.10070].

## 1. Physical basis and radiative transfer

The forest arises because the observed continuum from a high-redshift radio source is attenuated at frequencies corresponding to intervening neutral hydrogen. A standard thermally broadened line-of-sight optical depth used in recent work is
$$
\tau_\nu = \frac{3 h_P c^3 A_{10}}{32 \pi^{3/2} k_B \nu_{21}^2} \int \frac{n_{\mathrm{HI}}(x)}{b(x)\,T_S(x)} \exp\!\left[-\frac{(u(\nu)-v(x))^2}{b^2(x)}\right] dx,
$$
with $b=(2k_B T_K/m_H)^{1/2}$, where $n_{\mathrm{HI}}$ is the neutral hydrogen density, $T_S$ the spin temperature, $T_K$ the kinetic temperature, and $v(x)$ the line-of-sight peculiar velocity. In the optically thin limit, the differential brightness temperature against the background radiation is
$$
\delta T_b(\hat{s},\nu) \approx \frac{[T_S(\hat{s},z)-T_\gamma(\hat{s},z)]\,\tau(\hat{s},z)}{1+z},
$$
with $z=\nu_{21}/\nu-1$ [2511.13092]. A widely used intuitive approximation for the line optical depth is
$$
\tau_{21} \approx \frac{3 c^3 h_P A_{10} n_{\mathrm{HI}}}{32 \pi k_B \nu_{21}^2 T_S (1+z) (dv_{\parallel}/dr_{\parallel})},
$$
which makes explicit the dependence on density, spin temperature, redshift, and velocity gradients [2511.13092][1501.04425].

Absorption requires $T_S<T_\gamma$. The spin temperature is set by the balance between CMB coupling, collisions, and Wouthuysen–Field coupling,
$$
T_S^{-1}=\frac{T_\gamma^{-1}+x_c T_K^{-1}+x_\alpha T_c^{-1}}{1+x_c+x_\alpha},
$$
so the forest strengthens when Ly$\alpha$ coupling drives $T_S\rightarrow T_K$ while the gas remains cold, and it fades once X-ray heating raises $T_K$ and hence $T_S$ well above the background temperature [1310.7936][1101.5431]. This dependence makes the forest simultaneously a probe of thermodynamics and of the local gas distribution.

A persistent misconception is that the forest is synonymous with isolated minihalo lines. Recent simulation-based studies treat it more broadly: diffuse neutral IGM structures, filaments, infall regions around halos, and collapsed starless systems can all contribute. In detailed radiative-hydrodynamic modeling, the strongest features can arise in moderately overdense filaments and halo outskirts, while unresolved minihalos would add an even narrower, deeper component [1510.02296][2307.04130].

## 2. Absorbers, thermal history, and small-scale structure

The forest is unusually sensitive to small-scale cosmological structure because the relevant absorbers live on mass and length scales far below those probed by standard 21-cm tomography. In one minihalo-based treatment, the dominant absorber masses are $M\sim10^5$–$10^6\,h^{-1}M_\odot$ just above the Jeans scale, corresponding roughly to comoving radii $R\sim0.01$–$0.1\,h^{-1}\,\mathrm{Mpc}$ and hence to modes $k\sim10$–$100\,h\,\mathrm{Mpc}^{-1}$ [1403.1605]. This is the regime where warm dark matter free streaming, neutrino mass, primordial running, primordial black holes, and baryon–dark-matter relative velocity can all leave detectable imprints [1403.1605][2501.00769][2606.24665].

Warm dark matter is a canonical example. Its finite free-streaming length suppresses the linear matter power spectrum on small scales, lowers the abundance of low-mass halos, and erodes the fine-grained absorption network. A standard transfer function used for intuition is
$$
T(k)=[1+(\alpha k)^{2\mu}]^{-5/\mu}, \qquad \mu \approx 1.12,
$$
with $\alpha$ decreasing as $m_{\mathrm{WDM}}$ increases [2511.13092]. In forest observables, this appears as a depletion of narrow absorbers, earlier mergers of neighboring troughs, and suppression of high-$k_\parallel$ power in the 1D line-of-sight spectrum [2307.04130][2411.17094].

Heating acts differently. X-ray emissivity is commonly parameterized by $f_X$, which rescales the X-ray output per unit star-formation rate. Larger $f_X$ raises the IGM kinetic temperature, pushes $T_S$ upward, and compresses absorption contrast. In semi-numerical models this suppresses the forest more uniformly across scales than warm dark matter does, because heating modifies optical depth amplitudes even when the underlying small-scale density ranking is approximately preserved [2511.13092][2307.04130]. This distinction underlies much of the modern statistical analysis of the forest.

The same logic extends beyond warm dark matter. Massive neutrinos suppress the matter power spectrum below the free-streaming scale and can, in favorable scenarios, be constrained by the forest 1D power spectrum to around $0.1\,\mathrm{eV}$ when the IGM temperature is externally constrained [2501.00769]. Primordial black holes have a two-sided effect: PBH shot-noise isocurvature enhances low-mass halo formation, increasing line counts, whereas PBH accretion heating raises the IGM temperature and suppresses them. Forecasts indicate competitive upper limits as low as $f_{\mathrm{PBH}}\sim10^{-3}$ at $M_{\mathrm{PBH}}=100\,M_\odot$ under suitable assumptions [2104.10695]. Subhalos within host minihalos can also matter: one study found that although the boost is negligible for $10^5\,M_\odot$ hosts, substructure can enhance optical depth by an order of magnitude for $10^7\,M_\odot$ hosts and increase the integrated absorber abundance by up to order $10\%$ [2209.01305].

## 3. Statistical descriptions of the forest

Direct line-by-line spectroscopy is only one way to use the forest. A central development has been the move toward statistical observables that aggregate information from many weak features. Along a sightline, the standard summary is the one-dimensional power spectrum,
$$
\delta\widetilde{T'}(\hat{s},k_{\parallel})=\int \delta T'_b(\hat{s},r_z)e^{-ik_{\parallel}r_z}dr_z,\qquad
P(\hat{s},k_{\parallel})=\frac{|\delta\widetilde{T'}(\hat{s},k_{\parallel})|^2}{\Delta r_z},
$$
with $P_{\mathrm{1D}}(k_\parallel)$ obtained by averaging over segments or sources [2407.14298][2006.10070]. This statistic is especially effective because heating mainly changes the overall amplitude, whereas small-scale structure suppression changes the scale dependence. In forecasts for the forest contribution to the 21-cm power spectrum, the absorption signal was found to dominate a distinctive high-$k$ region, $k\gtrsim0.5\,\mathrm{Mpc}^{-1}$, when the IGM is cool [1310.7936].

The 1D power spectrum also has practical advantages. It reduces cosmic variance through averaging over independent segments and can detect the forest even when individual absorbers are too weak to identify. A dedicated z=6 detectability study showed that with 10 radio-loud sources the 1D forest power spectrum is detectable for $k\lesssim8.5\,\mathrm{MHz}^{-1}$ with 500 hr on the uGMRT and for $k\lesssim32.4\,\mathrm{MHz}^{-1}$ with 50 hr on SKA1-low if the IGM is $25\%$ neutral and the neutral regions have spin temperature $\lesssim30\,\mathrm{K}$ [2412.06879]. Earlier work on narrow-sightline statistics similarly concluded that a $\sim1000$ hr campaign targeting $\sim100$ narrow sightlines to $\sim1$–10 mJy sources could detect the LOS power spectrum and discriminate reionization scenarios [2006.10070].

A more recent development is the use of explicitly non-Gaussian and topological summaries. In “Topological Signatures of Heating and Dark Matter in the 21 cm Forest” the forest is converted into a standardized 1D field $z(\nu)$, a sublevel filtration
$$
X_t=\{\nu\mid z(\nu)\le t\},
$$
and a Betti-0 curve
$$
\beta_0(t)=\#\{j\mid b_j\le t<d_j\},
$$
where each trough $j$ has birth $b_j$, death $d_j$, and persistence lifetime $\ell_j=d_j-b_j$ [2511.13092]. The resulting descriptors—trough-line density $\lambda(t_\star)$, lifetime variance $\sigma_\ell^2$, and lifetime skewness $\gamma_\ell$—were shown to respond in nearly orthogonal directions across the $(f_X,m_{\mathrm{WDM}})$ parameter space. In that analysis, a common persistence cut $\tau_\star=0.411$ and a common threshold $t_\star$ fixed from the baseline noiseless CDM model stabilized the statistics, and the topological signatures remained detectable under SKA1-Low-like thermal noise. This suggests that the forest contains merger-hierarchy and connectivity information not accessible to amplitude-only summaries.

## 4. Background sources, instrumental requirements, and detectability

The limiting resource for forest work is not only telescope sensitivity but also the availability of sufficiently bright, compact background sources. Radio-loud quasars are the most studied candidates. A recent physics-driven source-population forecast found that if the radio-loud fraction remains roughly constant with redshift, a one-year SKA-LOW survey could detect approximately 20 radio-loud quasars at $z\sim9$ bright enough to resolve individual forest lines; if the radio-loud fraction declines strongly with redshift, the yield drops sharply and line spectroscopy becomes much harder [2407.18136]. At lower redshift, prospects are already improving: more than 30 radio-loud quasars are now known at $z\gtrsim5.5$, which materially changes the observational outlook for late-reionization forest searches [2412.06879].

Fine spectral resolution is indispensable. Several studies adopt $\sim1$ kHz channelization for direct or statistical analyses because the narrowest features are kHz-scale. In the topological study, degrading the resolution from 1 kHz to 5–10 kHz merged neighboring troughs and lowered and narrowed the Betti-0 peak [2511.13092]. In direct-detection forecasts with SKA1-Low, line widths of order 1–5 kHz are explicitly resolved, and targeted observations at $\Delta\nu=5$–20 kHz over 1000 hr were judged capable of detecting strong features for favorable source fluxes in the $z=7.5$–15 range [1501.04425].

Individual-line detection remains difficult. One early simulation study concluded that statistically significant 21-cm absorption against a Cygnus A–type source at $z\sim9$ could be detected by SKA in less than a year, whereas significant detection of the detailed forest features would require nearly a decade under per-channel methods [1101.5431]. In detailed radiative-hydrodynamic modeling, LOFAR was estimated to detect a few strong absorption features over a few tens of MHz for a 20 mJy source at $z=10$, while the SKA would recover a much larger fraction of the absorption information for the same source [1510.02296]. The practical implication is that direct line spectroscopy is possible but selective, whereas statistical observables are likely to dominate the first detections.

Calibration and continuum control remain central. The forest is less vulnerable than diffuse 21-cm tomography to wide-field foreground confusion because it uses compact sources as backlights, but it is not immune to bandpass structure, chromatic synthesized PSF leakage, radio-frequency interference, and continuum-model residuals [2006.10070][2412.06879]. Because many forest statistics depend primarily on rank ordering or on differential fluctuations after continuum subtraction, they can be more robust than absolute-amplitude observables, but only if spectral calibration is stable on kHz scales [2511.13092].

## 5. Parameter inference and constraints on astrophysics and fundamental physics

The forest is now used not only for detection but for quantitative inference. In a Fisher-forecast framework based on the 1D power spectrum, simulated SKA1-LOW observations at $z=9$ with 10 sources of $S_{150}=10\,\mathrm{mJy}$, 100 segments of 10 cMpc, and 100 hr per source yielded marginalized uncertainties of $\sigma(m_{\mathrm{WDM}})\approx1.3\,\mathrm{keV}$ and $\sigma(T_K)\approx3.7\,\mathrm{K}$ for a mildly heated fiducial model with $m_{\mathrm{WDM}}=6\,\mathrm{keV}$ and $T_K=60\,\mathrm{K}$; the corresponding SKA2-LOW forecast gave $\sigma(m_{\mathrm{WDM}})\approx0.3\,\mathrm{keV}$ and $\sigma(T_K)\approx0.6\,\mathrm{K}$ [2307.04130]. Even for a strongly heated case with $T_K\approx600\,\mathrm{K}$, the same study found $\sigma(m_{\mathrm{WDM}})\approx0.6\,\mathrm{keV}$ and $\sigma(T_K)\approx88\,\mathrm{K}$ for SKA2-LOW.

Likelihood-free inference has pushed this further. A normalizing-flow pipeline trained on simulated 21-cm forest power spectra recovered, for a low-heating SKA1-LOW 100 hr scenario with true $m_{\mathrm{WDM}}=6\,\mathrm{keV}$ and $T_K=60\,\mathrm{K}$,
$$
m_{\mathrm{WDM}}=6.28^{+1.61}_{-1.59}\,\mathrm{keV}, \qquad
T_K=60.17^{+6.16}_{-6.66}\,\mathrm{K},
$$
and for a high-heating SKA2-LOW 200 hr case,
$$
m_{\mathrm{WDM}}=6.47^{+2.04}_{-1.88}\,\mathrm{keV}, \qquad
T_K=627.95^{+94.18}_{-91.07}\,\mathrm{K},
$$
while explicitly modeling the non-Gaussian distribution of the 1D power spectrum [2407.14298]. A later z=6 study compared five inference pipelines and found that the most effective method bypassed the power spectrum entirely: a 1D U-Net produced a 256-dimensional latent representation of the noisy spectrum, and XGBoost regression on that latent space yielded meaningful constraints on $\langle x_{\mathrm{HI}}\rangle$ and $\log_{10}f_X$ even from a single 50 hr uGMRT sightline, corresponding to an approximately one-order-of-magnitude reduction in integration time relative to earlier power-spectrum-based techniques [2507.11611].

The forest also supports broader fundamental-physics constraints. Under weak astrophysical heating, one forecast based on the forest 1D power spectrum found sensitivity to dark-matter annihilation at $\langle\sigma v\rangle\sim10^{-31}\,\mathrm{cm^3\,s^{-1}}$ and to decay lifetimes $\tau\sim10^{30}\,\mathrm{s}$ for $10\,\mathrm{GeV}$ particles, together with sensitivity to primordial black holes at $M_{\mathrm{PBH}}\sim10^{15}\,\mathrm{g}$ and abundance $f_{\mathrm{PBH}}\simeq10^{-13}$ [2509.05705]. For neutrino mass, another forecast argued that, in an ideal scenario with external temperature information, the forest could constrain the summed mass to around $0.1\,\mathrm{eV}$ [2501.00769]. These numbers are highly model-dependent, but they illustrate the parameter space that becomes available once the forest is treated as a high-dimensional statistical field rather than as a set of isolated lines.

## 6. Limitations, misconceptions, and outlook

The dominant limitations are astrophysical degeneracy, source scarcity, and realism of small-scale modeling. Heating is the most persistent degeneracy: X-ray preheating can suppress the forest so efficiently that warm-dark-matter suppression, PBH heating, or neutrino-induced small-scale damping become difficult to distinguish unless the thermal history is constrained independently [2307.04130][2509.05705]. Source scarcity remains a practical bottleneck at the highest redshifts, especially if the radio-loud fraction evolves downward [2407.18136]. On the modeling side, unresolved minihalo gas physics, self-shielding, radiative transfer, shock heating, peculiar velocities, and reionization patchiness all affect line statistics and high-$k$ power [1510.02296][2411.17094].

Several common misconceptions can therefore be stated precisely. The first is that the forest is only a direct-detection problem. In fact, a large fraction of current progress relies on statistical observables—LOS power spectra, fluctuation variance, topology, and machine-learning summaries—which remain informative when no individual line is significant [1101.5431][2412.06879][2511.13092]. The second is that the forest is just a high-redshift analog of the Ly$\alpha$ forest. The analogy is useful, but the 21-cm forest is sensitive to spin temperature and to the radio background in a way that makes early heating a central part of the signal model [2606.24665]. The third is that foregrounds cease to matter because the background source is compact. Smooth foregrounds are less central than in diffuse tomography, but bandpass structure, continuum subtraction, chromatic PSF response, and RFI remain decisive systematics [2006.10070][2412.06879].

The near-term outlook is nonetheless substantially stronger than it was a decade ago. Source catalogs are expanding, late-end reionization models leave open a z≈6 forest window, SKA-Low-like sensitivities make statistical detection plausible, and analysis frameworks now include halo-model predictions, simulation-based inference, deep latent-space compression, and topological data analysis [2412.06879][2407.18136][2507.11611][2511.13092]. A plausible implication is that the first robust forest measurements will be hybrid: a small number of deep spectra from bright radio-loud quasars analyzed jointly with power-spectrum, non-Gaussian, and topology-aware summaries, then interpreted in combination with external constraints from global 21-cm experiments, 3D tomography, Ly$\alpha$ forest data, and high-redshift source surveys [2606.24665].

Source: https://www.emergentmind.com/topics/21-cm-forest