---
title: Haar Wavelet Downsampling Overview
url: https://www.emergentmind.com/topics/haar-wavelet-based-downsampling
type: topic
---

# Haar Wavelet Downsampling Overview

Searching arXiv for the cited papers to ground the article in the current literature.
Haar wavelet-based downsampling is the use of the orthonormal Haar discrete wavelet transform (DWT) to couple dyadic decimation with multiresolution analysis and exact reconstruction. In one dimension it separates a signal into approximation and detail coefficients; in two dimensions it produces the four critically sampled subbands $LL$, $LH$, $HL$, and $HH$, so that low-frequency structure and oriented high-frequency structure are represented explicitly rather than being collapsed by noninvertible pooling. In recent arXiv literature, this mechanism is used both as a front-end for serial X-ray crystallography image segmentation and lossy compression, where approximation coefficients are suppressed to isolate Bragg peaks, and as a preprocessing stage for CNN-based image super-resolution, where wavelet subbands define the representation learned by the network [2605.19199] [2204.07862].

## 1. Mathematical basis of Haar downsampling

The orthonormal Haar (db1) filter bank uses the 1D scaling and wavelet analysis filters
$$
h = [1/\sqrt{2},\, 1/\sqrt{2}], \qquad g = [1/\sqrt{2},\, -1/\sqrt{2}].
$$
This choice preserves energy and permits simple inversion. The same source characterizes Haar by one vanishing moment and the “shortest possible filter support,” and notes that the coefficients become integer-valued for hardware when the $\sqrt{2}$ normalization is factored appropriately [2605.19199].

For a discrete signal $x_j[n]$ at level $j$, the approximation and detail coefficients at level $j+1$ are
$$
a_{j+1}[n] = \sum_k h[k]\,x_j[2n-k], \qquad
d_{j+1}[n] = \sum_k g[k]\,x_j[2n-k].
$$
Equivalently, the signal is convolved with $h$ and $g$ and then downsampled by $2$. Because the Haar bank is orthonormal, synthesis uses the same filters up to reversal, and reconstruction is
$$
x_j[n] = \sum_k h[k]\,a_{j+1}\!\uparrow\!2[n-k] + \sum_k g[k]\,d_{j+1}\!\uparrow\!2[n-k],
$$
where $\uparrow 2$ denotes upsampling by inserting zeros between samples [2605.19199] [2204.07862].

In two dimensions, the transform is separable. For an image $X_j[m,n]$, the level-$j$ subbands are
$$
LL_j[p,q] = \sum_u\sum_v h[u]h[v]\,X_j[2p-u,2q-v],
$$
$$
LH_j[p,q] = \sum_u\sum_v h[u]g[v]\,X_j[2p-u,2q-v],
$$
$$
HL_j[p,q] = \sum_u\sum_v g[u]h[v]\,X_j[2p-u,2q-v],
$$
$$
HH_j[p,q] = \sum_u\sum_v g[u]g[v]\,X_j[2p-u,2q-v].
$$
Each subband is decimated by $2$ in both dimensions. Thus, for an $M\times N$ input, $LL_j$ has shape $(M/2^j)\times(N/2^j)$ after $j$ levels, and the same dimensional rule holds for the detail subbands at that scale [2605.19199].

## 2. Multilevel decomposition, pooling, and implementation semantics

A multilevel Haar pyramid is formed by repeatedly applying the 2D DWT to the $LL$ subband. In the crystallography pipeline, levels $j=1,\dots,J$ are produced in this way, with $J=4$ reported as optimal for the tested datasets [2605.19199]. In the blockwise formulation used for adaptive 2D Haar compression, one level maps each $2\times 2$ block to four coefficients and therefore yields four $n/2 \times m/2$ matrices $A$, $V$, $H$, and $D$ from an $n\times m$ matrix; repeated decomposition then implies an $n/2^L \times m/2^L$ approximation after $L$ levels [1410.0705].

The low-frequency subband has a particularly direct interpretation for Haar. At level $1$, $LL$ is a local average over a $2\times 2$ support; at deeper levels it corresponds to progressively larger effective windows. This makes Haar downsampling analogous to average pooling at the level of local averaging, but not at the level of representation. The DWT is frequency-selective, exactly invertible in the orthonormal setting, and spatially localized because of compact support; average or max pooling does not supply an oriented frequency decomposition and is not invertible [2605.19199].

This distinction is central to the practical meaning of “downsampling” in Haar systems. The transform is critically sampled: the aggregate number of coefficients across $LL$, $LH$, $HL$, and $HH$ equals the original number of pixels, even though each subband has half resolution in each axis. A plausible implication is that Haar downsampling is best understood not as raw data reduction by itself, but as a reallocation of samples into subspaces that can then be selectively discarded, thresholded, or transmitted.

The same literature also emphasizes that dyadic decimation introduces shift variance: small translations can change which samples populate $LL$ versus $LH/HL/HH$ [2605.19199]. In super-resolution implementations, boundary extension must also be chosen explicitly. Symmetric extension is described as commonly used with orthonormal wavelets such as Haar to preserve edge energy and avoid artifacts, whereas periodic extension can introduce seams and zero-padding can dim borders; however, that super-resolution paper does not specify which padding policy it used [2204.07862].

## 3. Background suppression and peak isolation in serial crystallography

In serial X-ray crystallography, Haar wavelet-based downsampling is used not merely to reduce spatial resolution but to separate smooth background scatter from localized diffraction peaks. The reported pipeline applies a 2D Haar DWT with $J$ levels, zeros the approximation coefficients, reconstructs from detail subbands only, thresholds the reconstructed image, runs peak finding, and transmits only the identified peaks for lossy compression [2605.19199].

For one reconstruction stage with approximation suppressed, the detail-only inverse synthesis is written as
$$
\hat{X}_j = (h_s \otimes g_s) * (LH_j\uparrow 2) + (g_s \otimes h_s) * (HL_j\uparrow 2) + (g_s \otimes g_s) * (HH_j\uparrow 2),
$$
with the standard recursive multiresolution synthesis applied across levels when $LL_J$ is zeroed [2605.19199]. The rationale given is domain-specific: smooth background from water or air scatter is concentrated in $LL$, whereas sharp Bragg peaks populate detail subbands across scales.

The empirical results are stated on 100 simulated nanoBragg frames with known ground truth derived from noiseless diffraction via connected-components and a 5-pixel matching radius. With Haar, $J=4$, and threshold $\tau \approx 170$ photons, the pipeline achieves $F1 \approx 0.96$, precision $P \approx 1.00$, and recall $R \approx 0.92$. Under the same evaluation framework, peakfinder8 with $SNRmin \approx 3.0$ yields $F1 \approx 0.37$, $P \approx 0.94$, and $R \approx 0.24$ [2605.19199].

The paper further reports downstream crystallographic analysis on real ePix10kA data. The metrics $CC^*$ and $R_{\mathrm{split}}$ converge at $J=4$ and track the unprocessed baseline through the practical resolution limit of approximately $1.7$--$1.8$ Å. The same source defines
$$
R_{\mathrm{split}} = \frac{1}{\sqrt{2}}
\cdot
\frac{\sum_h |I_h^{(1)} - I_h^{(2)}|}
{0.5\cdot \sum_h (I_h^{(1)} + I_h^{(2)})},
$$
$$
CC_{1/2} =
\frac{\sum_h (I_h^{(1)}-\bar{I}^{(1)})(I_h^{(2)}-\bar{I}^{(2)})}
{\sqrt{\sum_h (I_h^{(1)}-\bar{I}^{(1)})^2 \cdot \sum_h (I_h^{(2)}-\bar{I}^{(2)})^2}},
$$
and
$$
CC^* = \sqrt{\frac{2\cdot CC_{1/2}}{1+CC_{1/2}}}.
$$
Here $CC^*$ is interpreted as estimating the correlation for the fully merged dataset from the half-dataset correlation, while $R_{\mathrm{split}}$ measures consistency between half datasets after merging and scaling [2605.19199].

## 4. Wavelet choice, noise sensitivity, and hardware realization

A recurrent issue in Haar wavelet-based downsampling is whether Haar is merely the simplest option or whether it is specifically advantageous. In the serial crystallography study, a comparison of 12 wavelet families at $J=4$ and over 100 frames found Haar to be optimal for Bragg-peak detection, with $F1 \approx 0.96$ and a precision-recall curve reaching the upper-right corner of the precision-recall space. Other families clustered at $F1 \approx 0.6$--$0.73$, with coif1, bior2.2, and db2 identified as the next-best performers [2605.19199]. The explanation given is that minimal 2-tap support maximizes spatial localization and preserves peak amplitude, whereas longer filters spread peak energy across more coefficients and weaken the reconstructed peaks.

The same study also identifies a limitation under added Gaussian noise with $\sigma = 10$--$210$ ADU. The DWT-based method retains $F1 \approx 0.96$ at low noise levels, specifically for $\sigma < \sim 30$ ADU, but precision degrades beyond approximately $50$ ADU because noise populates the high-frequency detail subbands and no denoising is applied before inverse synthesis. By contrast, peakfinder8 maintains a stable $F1 \approx 0.37$ across the tested noise levels because it uses adaptive radial background and noise estimation [2605.19199]. The reported remedies are wavelet-domain detail thresholding before synthesis, consideration of fewer decomposition levels, and evaluation of alternative wavelets or combinations of Haar with detail thresholding. The same source notes, however, that the ePixUHR noise floor is well below 1 photon, so the observed degradation above $50$ ADU lies above expected operating conditions.

The hardware literature in the same paper treats Haar downsampling as especially attractive for streaming architectures. An FPGA implementation was demonstrated on an AMD/Xilinx Alveo U200 at 200 MHz. The DWT analysis filters are realized as fixed conv2D kernels with stride-2 and four output channels corresponding to $LL/LH/HL/HH$, while separable 1D filters are preferred for resource efficiency [2605.19199]. For a full design of 6 ASICs $\times$ 8 partitions, giving 48 parallel cores, a separable float16 implementation is projected at approximately 3,312 DSP slices and about 260 BRAM blocks, corresponding to roughly 48% and 15% of U200 resources respectively. The same work states that a complete 4-level pipeline per core is projected at about 69 DSPs if the initiation interval is relaxed for deeper levels, and that single-layer cores run in approximately 24--27 $\mu$s at 200 MHz [2605.19199].

Boundary handling is treated operationally rather than abstractly in that hardware setting. The paper does not prescribe a specific mathematical extension rule, but suppresses edge artifacts through overlapped tiling: 3 columns per side for a $7\times 7$ kernel and 2 columns per side for a $5\times 5$ kernel. It further states that the ePixUHR ASIC streams $192\times 168$ pixels in 8 parallel partitions of 24 columns, that DWT cores operate per partition with overlap, and that the fixed coefficients provide a path from ePixUHR firmware to on-detector ASIC implementation in the SparkPix detector family [2605.19199].

## 5. Haar downsampling in wavelet-space super-resolution

In CNN-based image super-resolution, Haar wavelet-based downsampling serves as a representation transform rather than as a segmentation operator. A single-level 2D Haar DWT produces $LL$, $LH$, $HL$, and $HH$ from the low-resolution input; the network then predicts residual corrections for the detail subbands, and the inverse DWT combines the preserved low-frequency component with the corrected detail bands to reconstruct the output [2204.07862].

The reported single-level network uses 10 convolutional layers, each with 64 filters of size $3\times 3$ followed by ReLU,
$$
y_\ell = \max(0, x_\ell * K_\ell), \qquad \ell = 1,\dots,10,
$$
and is trained with an $L_2$ loss in wavelet space,
$$
\mathcal{L}_{L2}(y,x) = \| W_v(y),\, W_v(x) + \Delta W_v(x)\|_2^2,
$$
using Adam with learning rate $0.01$, momentum decay $\beta = 0.9$, for 50 epochs [2204.07862].

A central finding of that study is empirical invariance across 37 single-level wavelets drawn from the Haar, Daubechies, Biorthogonal, Reverse Biorthogonal, Coiflets, and Symlets families. Trained on DIV2K and evaluated on Set14, the metrics vary only slightly across wavelets. Haar reports PSNR $26.5901$, SSIM $0.8560$, FSIM $0.9616$, GSM $0.9945$, MAD $0.4160$, SR-SIM $0.9807$, and VIF $0.6243$, while the full range spans roughly PSNR $26.43$--$26.62$, SSIM $0.8498$--$0.8573$, FSIM $0.9591$--$0.9620$, and VIF $0.6118$--$0.6270$ [2204.07862]. The paper interprets this as evidence that, in the single-level setting, the four-band decomposition and the CNN’s operation in wavelet space matter more than fine-grained differences among scalar wavelet filters.

The same paper contrasts single-level Haar with the GHM multi-level multiwavelet transform. GHM uses multiple scaling and wavelet functions, a two-stage matrix filter bank with provided matrices $H_1$, $H_2$, $G_1$, and $G_2$, and yields 16 subbands rather than 4 [2204.07862]. On Set14, the table reported in that work gives bicubic / Haar / GHM values of PSNR $25.5616 / 26.5901 / 26.4518$, SSIM $0.8130 / 0.8560 / 0.7917$, FSIM $0.9418 / 0.9616 / 0.9868$, GSM $0.9917 / 0.9945 / 0.9978$, MAD $0.3962 / 0.4160 / 0.4183$, SR-SIM $0.9701 / 0.9807 / 0.9946$, and VIF $0.5210 / 0.6243 / 0.6554$ [2204.07862]. Thus, GHM improves FSIM, GSM, MAD, SR-SIM, and VIF over both bicubic and Haar, but yields slightly lower PSNR and SSIM than Haar on Set14. Within the scope of that evidence, Haar remains the pragmatic default for single-level wavelet-space CNNs because of simplicity and speed, whereas GHM is used when richer multiscale representation is prioritized.

## 6. Adaptive 2D Haar banks and compression-oriented variants

A distinct line of work treats Haar wavelet-based downsampling through adaptive 2D wavelet banks defined directly on the unit square rather than through standard separable 1D filter taps. In that formulation, Haar multiresolution analysis is generated by the characteristic function of $[0,1]\times[0,1]$, with expansion matrix
$$
M=\begin{pmatrix}2&0\\0&2\end{pmatrix},
$$
and three piecewise-constant wavelet functions $\eta_i(x,y)$ supported on the four quadrants of the unit square, each quadrant assigned constants $a_{ij}$ subject to orthonormality constraints [1410.0705].

The operational transform in that paper is blockwise. For a $2\times 2$ matrix $M=(m_{ij})$, the low-pass and detail coefficients are defined as
$$
a = (m_{11}+m_{12}+m_{21}+m_{22})/4,
$$
$$
v = \sum_{i=1}^{2}\sum_{j=1}^{2}\psi_{ij}^{1}m_{ij}, \qquad
h = \sum_{i=1}^{2}\sum_{j=1}^{2}\psi_{ij}^{2}m_{ij}, \qquad
d = \sum_{i=1}^{2}\sum_{j=1}^{2}\psi_{ij}^{3}m_{ij}.
$$
Applying this transform over disjoint $2\times 2$ blocks produces four matrices $A$, $V$, $H$, and $D$, each of size $n/2 \times m/2$ for an $n\times m$ input [1410.0705]. In this setting, the approximation matrix $A$ is explicitly the block-average image, so it acts as the antialiased downsampled representation.

The same work constructs four orthonormal bases in $L^2(\mathbb{R}^2)$: a classical Haar basis and vertical, horizontal, and diagonal variants [1410.0705]. Its adaptive Haar transform computes all four transforms for each $2\times 2$ block and selects the basis that minimizes
$$
v_k^2 + h_k^2 + d_k^2.
$$
This is a local, spatially adaptive criterion intended to reduce detail energy prior to quantization and coding.

The reported compression pipeline converts an image to three RGB matrices, applies the two-dimensional wavelet transformation to each color matrix, quantizes the coefficients, performs Huffman coding, and writes the binary stream to a file [1410.0705]. Compression rates are reported in the range of 30%--50%. For quantization, the stated rates are 54% for 64 levels, 49% for 32 levels, 45% for 16 levels, and 44% for 8 levels. For decomposition depth, the rates are 49% at level 1, 37% at level 2, 35% at level 3, and 34% at level 4 [1410.0705]. The same paper states that the adaptive scheme is worse than simple wavelet-basis compression on average by 1--3% because the basis-id matrix must also be stored, and that no correlation between image type and compression was revealed.

A common misconception is that “Haar downsampling” is synonymous with crude averaging. These papers show a narrower and more technical picture. In the adaptive compression setting, the approximation coefficient is indeed a $2\times 2$ average; in the separable DWT setting, however, the average is only one branch of a perfect-reconstruction filter bank whose detail channels carry orientation and scale information [1410.0705] [2605.19199]. A second misconception is that the simplest wavelet is necessarily the weakest. In serial crystallography, the evidence points in the opposite direction: Haar’s minimal support is precisely what makes it optimal among the tested wavelet families for localized Bragg-peak detection [2605.19199].

Source: https://www.emergentmind.com/topics/haar-wavelet-based-downsampling