---
title: Patch-Wise Blurring Diffusion
url: https://www.emergentmind.com/topics/patch-wise-blurring-diffusion
type: topic
---

# Patch-Wise Blurring Diffusion

Patch-wise blurring diffusion is a framework for image restoration that applies diffusion processes independently to spatially overlapping patches of an image, enabling localized modeling of noise and degradation artifacts. The approach was formalized for thermal imaging applications in the TDiff method, which addresses resolution loss, fixed pattern noise, and localized artifacts commonly found in thermal images from low-cost cameras. By modeling the diffusion prior at the patch level and integrating mechanisms for inverse problem guidance, patch-wise blurring diffusion achieves unified restoration across denoising, deblurring, and super-resolution tasks while managing limited and non-diverse training data [2510.06460].

## 1. Patch-based Diffusion Process

Patch-wise blurring diffusion decomposes a high-resolution image $x_0 \in \mathbb{R}^{H \times W}$ into a set of overlapping patches using an extraction operator $P_k$. Each patch is of size $ps \times ps$ and is processed independently through a denoising diffusion probabilistic model (DDPM) framework. For each patch $k$, the forward diffusion process is defined as a Markov chain:
\[
q(P_k(x_t) \mid P_k(x_{t-1})) = \mathcal{N}\left(P_k(x_t); \sqrt{\alpha_t} P_k(x_{t-1}), \beta_t I_{ps^2}\right)
\]
where the schedule $\{\beta_t\}$ defines the noise variance and $\alpha_t = 1 - \beta_t$, $\bar{\alpha}_t = \prod_{i=1}^t \alpha_i$. The process runs from $t = 1$ to $T$, and the perturbed patch at time $t$ can be sampled directly from $x_0$:
\[
P_k(x_t) = \sqrt{\bar{\alpha}_t} P_k(x_0) + \sqrt{1 - \bar{\alpha}_t} \epsilon, \quad \epsilon \sim \mathcal{N}(0, I_{ps^2})
\]
The diffusion process is applied to each patch independently, leveraging the localized nature of distortions commonly observed in thermal imaging [2510.06460].

## 2. Reverse Diffusion and Denoising

The reverse diffusion process, or denoising step, aims to recover clean patches from noisy observations by modeling:
\[
p_\theta(P_k(x_{t-1}) \mid P_k(x_t)) = \mathcal{N}\left(P_k(x_{t-1}); \mu_\theta(P_k(x_t), t), \Sigma_t I_{ps^2}\right)
\]
where the mean $\mu_\theta$ depends on a learned noise predictor $\epsilon_\theta$ realized by a time-conditional U-Net:
\[
\mu_\theta(P_k(x_t), t) = \frac{1}{\sqrt{\alpha_t}} \left(
    P_k(x_t) - \frac{\beta_t}{\sqrt{1 - \bar{\alpha}_t}} \epsilon_\theta(P_k(x_t), t)
\right)
\]
The objective for training $\epsilon_\theta$ on each patch is:
\[
L(\theta) = \mathbb{E}_{t, x_0, \epsilon} \left\| \epsilon - \epsilon_\theta(
    \sqrt{\bar{\alpha}_t} x_0 + \sqrt{1 - \bar{\alpha}_t} \epsilon, t
) \right\|^2
\]
This approach enables learning a prior over small localized regions, allowing effective denoising and restoration of localized degradations [2510.06460].

## 3. Patch Extraction, Tiling, and Reconstruction

The image is partitioned into a tiled grid of overlapping patches:
- For an image of width $W$ and height $H$, with patch size $ps$ and stride $s$, the number of patches horizontally is $M = \lfloor (W-ps)/s \rfloor+1$, vertically $N = \lfloor (H-ps)/s \rfloor+1$, and total patches $K = MN$.
- Each patch $P_k$ is indexed by starting coordinates $(i_k, j_k)$, determined by:
  \[
  i_k = \left\lfloor \frac{k-1}{M} \right\rfloor s, \quad j_k = [(k-1) \mod M] s
  \]
- Patches are extracted such that $[P_k(x)]_{u,v} = x_{i_k + u,\, j_k + v}$, for $0 \leq u, v < ps$.

Overlapping denoised patches are reassembled into a full-resolution image via smooth windowing and normalization to avoid seams. The window function is the 2D raised-cosine (Hann) window:
\[
w(u,v) = h(u) h(v), \quad h(r) = \frac{1}{2} [1 - \cos(2\pi r/(ps-1))]
\]
Reconstruction is performed by weighted averaging:
\[
W(i,j) = \sum_{k: (i,j)\in \text{patch } k} w(i-i_k,\,j-j_k)
\]
\[
\hat{x}_t(i,j) = \frac{1}{W(i,j)} \sum_{k: (i,j)\in \text{patch } k} w(i-i_k,\,j-j_k) [\hat{P}_k(x_0|t)]_{i-i_k,\,j-j_k}
\]
[2510.06460]

## 4. Inverse Problem Guidance and Measurement Conditioning

For applications such as deblurring or general inverse imaging, patch-wise blurring diffusion integrates explicit measurement guidance during reverse diffusion. Given degradation $y = A x_0 + \eta$ (with $A$ representing a linear degradation, e.g., blur or downsampling), at each time $t$ an estimated clean patch is computed:
\[
\hat{P}_k(x_0|t) = \frac{P_k(x_t) - \sqrt{1 - \bar{\alpha}_t}\, \epsilon_\theta(P_k(x_t), t)}{\sqrt{\bar{\alpha}_t}}
\]
Measurement-consistent updates are enforced per patch via:
- Back-projection: $g_{\text{BP}} = A^{\top} (A A^{\top} + \lambda I)^{-1} (A\, \hat{P}_k - y_k)$
- Least squares correction: $g_{\text{LS}} = c\, A^{\top} (A\, \hat{P}_k - y_k)$

The guided reverse step updates each $P_k(x_{t-1})$ by incorporating a weighted combination of $g_{\text{BP}}$ and $g_{\text{LS}}$, controlled by a time-dependent parameter $\delta_t$:
\[
P_k(x_{t-1}) = \sqrt{\alpha_t} \hat{P}_k(x_0|t) + \sqrt{1-\alpha_t} \epsilon_\theta(P_k(x_t),t) - \mu_t[(1-\delta_t)g_{\text{BP}} + \delta_t g_{\text{LS}}]
\]
This mechanism allows the framework to act as a plug-and-play prior for inverse problems within the diffusion process [2510.06460].

## 5. Architectural and Training Details

The core denoiser network $\epsilon_\theta$ is a grayscale, time-conditional U-Net, parameterized as follows:
- Base channel count: $c_0=64$ for $ps=128$, $c_0=32$ for $ps=64$
- Channel multipliers: $[1, 1, 2, 2, 4, 4]$
- Sinusoidal timestep embeddings are added at every resolution

For deblurring, the blur kernel or its frequency response is included as a second input channel, or provided via cross-attention at the bottleneck. The forward and reverse diffusion employ a schedule $\beta_{\text{start}}=10^{-4}$, $\beta_{\text{end}}=2 \times 10^{-2}$, $T=1\,000$ steps. Training uses the Adam optimizer with learning rate $2\times 10^{-4}$ and batch size approximately $200$ patches [2510.06460].

## 6. Inference Pipeline

At inference, the full restoration process proceeds as:
1. Initialize $x_T$ as independent Gaussian noise.
2. For $t = T \ldots 1$:
   - Extract patches $P_k(x_t)$ for all $k$.
   - For each $k$, compute $\epsilon_\theta(P_k(x_t),t)$ and $\hat{P}_k(x_0|t)$.
   - Compute $g_{\text{BP}}$, $g_{\text{LS}}$ for each patch using local measurements.
   - Update $P_k(x_{t-1})$ via the guided reverse step.
   - Merge patches by windowed average to obtain $x_{t-1}$.
3. After $t=0$, the restored image $x_0$ is obtained.

This patch-based diffusion with smooth blending is directly implementable in PyTorch or TensorFlow and realizes the TDiff restoration pipeline [2510.06460]. 

## 7. Significance and Applications

Patch-wise blurring diffusion, as instantiated in TDiff, is the first framework to apply a learned diffusion prior at the patch level for thermal image restoration across multiple tasks and measurement settings. The approach leverages local structure for robust restoration on limited training data, provides a unified pipeline for denoising, deblurring, and super-resolution, and enables consistent restoration even under real measurement conditions. Its generality suggests applicability beyond thermal to other imaging modalities exhibiting local, patch-dependent degradation [2510.06460].

Source: https://www.emergentmind.com/topics/patch-wise-blurring-diffusion