---
title: 'HPG-Diff: Hierarchical Physics-Guided Diffusion'
url: https://www.emergentmind.com/topics/hierarchical-physics-guided-diffusion-hpg-diff
type: topic
---

# HPG-Diff: Hierarchical Physics-Guided Diffusion

Searching arXiv for the cited HPG-Diff-related papers to ground the article in current records.
Hierarchical Physics-Guided Diffusion (HPG-Diff) most directly refers to a diffusion-based framework for topology optimization that combines hierarchical conditioning by precomputed physics features with a differentiable connectivity penalty against floating material [2607.07233]. In a broader methodological sense, closely related arXiv work applies analogous ideas—hierarchical conditioning, structured intermediate variables, or physics-constrained diffusion updates—to video frame interpolation, human motion generation, and multi-slice scientific imaging, although those systems use different names and do not share a single canonical formulation [2504.00380] [2212.02500] [2512.06977]. Across these settings, the common premise is that diffusion is more reliable when its denoising trajectory is constrained by problem-structured physical information rather than left to operate solely in an unconstrained latent or pixel space.

## 1. Terminology and conceptual scope

In the strict sense used by the paper titled “HPG-Diff: Hierarchical physics-guided diffusion with differentiable connectivity constraints for topology optimization,” HPG-Diff is a generative method for compliance-driven topology optimization under varying boundary conditions, loads, and volume fractions [2607.07233]. Its defining components are a DDPM-style UNet, hierarchical injection of three physics features—displacement \(U\), principal stress line (PSL), and strain energy density (SED)—and a Floating Material Suppression (FMS) loss.

A broader reading of the term is possible but must be qualified. The frame interpolation paper “Hierarchical Flow Diffusion for Efficient Frame Interpolation” replaces latent-image denoising with hierarchical diffusion over bilateral optical flow; this is a hierarchical and physically meaningful intermediate representation, but its “physics” is motion structure rather than an external simulator or mechanics prior [2504.00380]. “PhysDiff: Physics-Guided Human Motion Diffusion Model” inserts a physics-based projection module into the diffusion sampling loop, yet the paper explicitly notes that it is not hierarchical in the classical coarse-to-fine or multi-resolution sense [2212.02500]. The multi-slice reconstruction framework presented as DART and DRIFT couples diffusion priors with modality-specific forward models and partitions the slice volume across GPUs; its hierarchy is architectural and computational, rather than a named HPG-Diff formalism [2512.06977].

A common misconception is that “physics-guided diffusion” always denotes inference-time simulation inside each reverse step. That is not generally true. In topology optimization, the physics-guided components primarily shape training and conditioning, while inference remains simple stochastic diffusion sampling after one preprocessing FEA pass [2607.07233]. By contrast, PhysDiff modifies the sampling loop directly through iterative projection in a physics simulator [2212.02500].

## 2. Core architecture in topology optimization

The topology-optimization HPG-Diff model is built on a standard DDPM-style UNet with reverse transition
\[
p_\theta(x_{t-1} \mid x_t, c) = \mathcal{N}\left(x_{t-1} \mid \mu_\theta(x_t, t, c), \Sigma_\theta(t)\right),
\]
where \(x_t\) is the noisy density field, \(x_{t-1}\) is the denoised sample, and \(c\) denotes the physics-guidance condition [2607.07233]. The model denoises a topology image or density map from noise to a binary-like material layout.

The central architectural idea is that the UNet already processes information hierarchically: shallow layers preserve local detail, mid layers capture intermediate structure, and deep layers encode global layout. HPG-Diff aligns this hierarchy with three precomputed FEA-derived features rather than concatenating all physics inputs into a single conditioning tensor. The three features are:

- **Displacement \(U\)**: a low-level local deformation field.
- **Principal Stress Line (PSL)**: a mid-level representation of load paths.
- **Strain Energy Density (SED)**: a high-level indicator of stiffness-critical regions.

These are obtained from an FEA solve on a uniform initial density field under prescribed loads and supports. The paper gives the equilibrium relation
\[
KU = F,
\]
the compliance objective
\[
c(x)=F^T U = U^T K U,
\]
and the total potential energy
\[
\Pi = \frac{1}{2}U^T K U - U^T F.
\]
SED is derived from strain energy, and PSL is computed from the stress tensor principal directions,
\[
\boldsymbol{\sigma} =
\begin{bmatrix}
\sigma_x & \tau_{xy}\\
\tau_{xy} & \sigma_y
\end{bmatrix},
\qquad
\boldsymbol{\sigma}\mathbf{v}_i = \sigma_i \mathbf{v}_i,\quad i=1,2.
\]

The layer assignment is explicit: \(U\) is injected into shallow layers, PSL into mid layers, and SED into deep layers via layer-specific cross-attention modules. The intended effect is semantic alignment between the abstraction level of the physics signal and the abstraction level of the denoising network. The paper interprets this mechanically: \(U\) provides fine-grained boundary and deformation guidance, PSL encodes skeleton-like load-transfer paths, and SED biases global material allocation toward stiffness-critical regions [2607.07233].

This arrangement is presented as the main reason the model generalizes more effectively to unseen boundary conditions than approaches conditioned only on a single global field such as SED. A plausible implication is that the hierarchy is not merely a representational convenience; it functions as an inductive bias that ties denoising decisions to mechanics-relevant scales.

## 3. Differentiable connectivity and floating material suppression

A major failure mode in generative topology optimization is floating material: disconnected dense regions that consume volume but do not contribute to load transfer. HPG-Diff addresses this with the Floating Material Suppression loss, a differentiable connectivity penalty inspired by virtual heat propagation from the load location [2607.07233].

The construction begins from a binary seed mask \(\mathbf{S}_{\mathrm{load}}\) at the load position, the predicted density field \(\hat{\mathbf{x}}_{\mathrm{pred}}\), and an iterative connectivity map \(\mathbf{D}^{(k)}\). The local propagation operator is defined over a \(3\times 3\) neighborhood:
\[
\big(\mathrm{Max}_{3\times3}(\mathbf{D})\big)_{ij} = \max_{(u,v)\in\mathcal{N}_1(i,j)} D_{uv}.
\]
Initialization is
\[
\mathbf{D}^{(0)} = \mathbf{S}_{\mathrm{load}}.
\]
Propagation then proceeds by
\[
D_{ij}^{(k+1)} = \max\Big( D_{ij}^{(k)}, (\mathrm{Max}_{3\times3}(\mathbf{D}^{(k)}))_{ij} \cdot (\hat{\mathbf{x}}_{\text{pred}})_{ij} \Big).
\]
After \(k\) steps, the final connectivity map is \(\mathbf{D}_{\mathrm{final}} = \mathbf{D}^{(k)}\).

Conceptually, the mechanism acts as a reachability proxy. Propagation starts at the load seed and expands only through predicted solid regions. Dense areas connected to the seed receive high propagated values; dense disconnected islands do not. The paper makes the binary-limit interpretation explicit through a masked reachability dilation,
\[
\mathcal{R}^{(0)} = S_{\mathrm{load}}, \quad
\mathcal{R}^{(k+1)} = \mathcal{R}^{(k)} \cup \{i\in\Omega_s \mid i \text{ is adjacent to some } j\in\mathcal{R}^{(k)}\}.
\]

The resulting loss is
\[
\mathcal{L}_{\mathrm{fm}} = \mathbb{E}_{\mathbf{x}_0,\epsilon,t} \left[ w(t)\; \frac{1}{|\Omega|} \sum_{(i,j)\in\Omega} (\hat{\mathbf{x}}_{\mathrm{pred}})_{ij} \big(1-(\mathbf{D}_{\mathrm{final}})_{ij}\big) \right],
\]
with time weighting
\[
w(t)=\exp(-c\,t/T).
\]
The full training objective is
\[
\mathcal{L} = \mathcal{L}_{\mathrm{noise}} + \lambda_{\mathrm{fm}}\mathcal{L}_{\mathrm{fm}}.
\]

The time weighting is operationally important. Early in denoising, the predicted density field is too noisy for connectivity to be meaningful, so the exponential term suppresses FMS. Near the end of the reverse chain, when topology becomes structured, the penalty strengthens. The paper states that this avoids gradient conflict with the standard diffusion noise loss. It also states that FMS is a soft differentiable penalty rather than a hard graph-theoretic guarantee of connectivity [2607.07233].

## 4. Training protocol, benchmarks, and quantitative performance

HPG-Diff is trained and evaluated on the topology optimization dataset from Mazé et al. / TopoDiff. The benchmark design space consists of \(64\times64\) square grids with volume fraction \(v_f\) ranging from 0.3 to 0.5 in steps of 0.02, random loads on unconstrained boundary nodes, load directions over \([0,\pi]\) in \(\pi/6\) increments, 42 boundary conditions used in training, and 5 unseen boundary conditions reserved for out-of-distribution testing [2607.07233]. The dataset contains 30,000 training samples, 1,800 in-distribution test samples, and 1,000 out-of-distribution test samples.

Evaluation uses binarized outputs with threshold \(\tau = 0.5\). The principal metrics are Compliance Error (CE), Volume Fraction Error (VFE), and Floating Material ratio (FM). On the in-distribution Test 1 split, HPG-Diff reports Avg CE \(0.87\%\), Med CE \(0.16\%\), Avg VFE \(1.38\%\), and FM \(2.90\%\). On the out-of-distribution Test 2 split, it reports Avg CE \(5.29\%\), Med CE \(0.61\%\), Avg VFE \(1.45\%\), and FM \(2.44\%\) [2607.07233].

Compared with the baselines listed in the paper, the margins are substantial. For Test 1, TopologyGAN reports Avg CE \(48.51\%\) and FM \(46.78\%\), TopoDiff-Guided reports Avg CE \(4.39\%\) and FM \(5.54\%\), and DOM w/ TA reports Avg CE \(4.44\%\) and FM \(6.72\%\). For Test 2, TopologyGAN reports Avg CE \(143.08\%\) and FM \(67.90\%\), TopoDiff-Guided reports Avg CE \(18.40\%\) and FM \(6.21\%\), and DOM w/ TA reports Avg CE \(32.19\%\) and FM \(14.20\%\).

The OOD result is especially emphasized. In the ablation study, removing hierarchical physics guidance causes Avg CE on Test 2 to jump from \(5.29\%\) to \(36.62\%\), indicating that the hierarchical conditioning rather than the diffusion backbone alone is the principal source of generalization under unseen boundary conditions [2607.07233]. The paper also reports a long-tailed error distribution on Test 2: Std CE \(31.76\%\), Max CE \(791.07\%\), and CE \(> 30\%\) for only \(2.70\%\) of samples. This suggests that most generated samples remain accurate, while a small number of failures dominate the average.

The paper further compares HPG-Diff with iterative post-processing. DOM w/ TA + SIMP\(_5\) reports Mdn CE \(1.89\%\) and FM \(10.19\%\), DOM w/ TA + SIMP\(_{10}\) reports Mdn CE \(1.15\%\) and FM \(2.61\%\), and HPG-Diff reports Mdn CE \(0.61\%\) and FM \(2.44\%\). The stated significance is that HPG-Diff reaches competitive or better quality without requiring iterative post-processing during inference [2607.07233].

## 5. Adaptation, scope conditions, and limitations

Beyond the square benchmark domain, the paper studies adaptation to a \(60\times20\) 3:1 rectangular domain using LoRA fine-tuning on 1,000 samples [2607.07233]. The reported case studies include a shelf bracket in the square domain, a cantilever beam in the 3:1 domain, and a classic bridge in the 3:1 domain. For the shelf bracket, the ground-truth compliance is 5.0506 and generated variants are 4.9962, 5.0398, and 5.0772. For the cantilever beam, the ground-truth compliance is 182.8988 and generated results are 179.1958 and 185.5106. For the classic bridge, the ground-truth compliance is 13.0730 and generated results are 12.9581 and 12.9663.

On an independent 200-sample rectangular-domain evaluation, the LoRA-adapted model reports Avg CE \(2.86\%\), Med CE \(0.58\%\), Avg VFE \(1.54\%\), and FM \(3.17\%\). The paper characterizes this as preliminary evidence that lightweight adaptation can transfer a pretrained square-domain prior to non-square domains without retraining from scratch [2607.07233].

The limitations are explicit. The framework is still limited to the benchmark setting and requires broader validation on more diverse geometries and problems. It assumes a single-load benchmark when instantiating FMS from the load seed; multi-load problems would require a multi-seed extension. It does not yet include nonlinear materials, fatigue, dynamic loads, overhang limits, or minimum length scales. The FMS loss reduces floating material substantially but does not mathematically guarantee perfect connectivity. LoRA reduces adaptation cost but still requires target-domain solver-generated data [2607.07233].

A frequent misreading is that HPG-Diff replaces classical topology optimization. The more precise interpretation supported by the paper is that it moves generative topology optimization closer to classical physics-based optimization in quality while retaining generative efficiency and diversity.

## 6. Related hierarchical and physics-guided diffusion formulations

The frame interpolation method “Hierarchical Flow Diffusion for Efficient Frame Interpolation” provides a closely related but domain-specific analogue [2504.00380]. Its central move is to avoid denoising a large latent image space and instead denoise bilateral optical flow fields \(f_0\) and \(f_1\) in a coarse-to-fine hierarchy, conditioned on multiscale image features. The denoiser predicts
\[
\tilde{f}_0^i, \tilde{f}_1^i = u(F_0^i, F_1^i, f_0^t, f_1^t, t; \theta),
\]
and the predicted flows drive a flow-guided image synthesizer,
\[
I_t = g(I_0, I_1, f_0, f_1; \Phi).
\]
The final rendered frame is
\[
\tilde{I}_t = M \odot w(I_0, \tilde{f}_0) + (1 - M) \odot w(I_1, \tilde{f}_1) + \Delta I.
\]
The paper reports state-of-the-art results on SNU-FILM, Xiph, DAVIS, and Vimeo-90k, as well as inference times on an RTX-4090 with \(1024 \times 1024\) input of 8.3 s for LDMVFI, 2.1 s for CBBD, 0.19 s for SGM-VFI, and 0.20 s for the proposed method, supporting the claim that it is \(10+\times\) faster than other diffusion-based methods [2504.00380]. In this case, the hierarchy is explicitly coarse-to-fine, and the structured intermediate variable is bilateral optical flow.

PhysDiff occupies a different point in the design space [2212.02500]. It is best understood as a plug-and-play physics-guided diffusion sampler for human motion generation. A pretrained motion denoiser \(D\) is wrapped by a physics-based projection operator \(\mathcal{P}_\pi\) implemented through a motion imitation policy in IsaacGym. The denoised motion is projected to a physically plausible motion and fed back into the next diffusion step. The method targets floating, foot sliding or skating, and ground penetration, summarized by
\[
\text{Phys-Err} = \text{Penetrate} + \text{Float} + \text{Skate}.
\]
On HumanML3D, MDM reports FID 0.544 and Phys-Err 31.572, while PhysDiff w/ MDM reports FID 0.433 and Phys-Err 4.111; on UESTC, MDM reports Phys-Err 28.371, while PhysDiff w/ MDM reports 1.463 [2212.02500]. The paper explicitly states that this is a scheduled physics-guided diffusion method rather than a true hierarchical multi-stage method; more projection steps generally help physical plausibility, but applying them later in the diffusion process works best.

The multi-slice reconstruction framework described as DART and DRIFT supplies another related formulation [2512.06977]. It uses a video diffusion prior over a multi-slice object \(X \in \mathbb{C}^{S \times N \times N}\), partitions slices into groups across GPUs, and alternates or combines diffusion denoising with modality-specific physics updates for MRI and 4D-STEM. For MRI, the physics step is a proximal gradient update on the undersampled Fourier fidelity term; for 4D-STEM, it is a proximal gradient update on intensity mismatch under the multislice forward model. DART interleaves diffusion and physics at each step, whereas DRIFT samples multiple diffusion candidates, selects the best initialization via SSIM, and then performs about 100 physics gradient steps. On MRI in-distribution evaluation, DART reports SSIM \(0.968 \pm 0.011\), exceeding Projection-Based \(0.844 \pm 0.027\), CS MRI \(0.804 \pm 0.025\), and TV \(0.803 \pm 0.024\); on OOD MRI Roots, DART reports SSIM \(0.813 \pm 0.130\) [2512.06977]. Here the hierarchy is partitioned and distributed rather than layer-specific or simulator-scheduled.

Taken together, these papers suggest that “hierarchical physics-guided diffusion” is less a single algorithm than a recurring design principle. The hierarchy may be over UNet depth, flow scales, diffusion timesteps, or slice partitions; the physics guidance may be precomputed FEA fields, simulator projection, or forward-model data consistency. What remains stable across the literature surveyed here is the rejection of unconstrained denoising in favor of denoising trajectories that are aligned with domain structure and physically meaningful constraints [2607.07233] [2504.00380] [2212.02500] [2512.06977].

Source: https://www.emergentmind.com/topics/hierarchical-physics-guided-diffusion-hpg-diff