Papers
Topics
Authors
Recent
Search
2000 character limit reached

HPG-Diff: Hierarchical Physics-Guided Diffusion

Updated 12 July 2026
  • Hierarchical Physics-Guided Diffusion is a generative method that integrates precomputed physics features with diffusion models to optimize material topology under varying boundary conditions.
  • It leverages a DDPM-style UNet with hierarchical injection of displacement, principal stress line, and strain energy density to align denoising with mechanical scales.
  • The approach employs a differentiable floating material suppression loss to reduce disconnected material regions, thereby enhancing compliance and generalization compared to baselines.

Searching arXiv for the cited HPG-Diff-related papers to ground the article in current records. Hierarchical Physics-Guided Diffusion (HPG-Diff) most directly refers to a diffusion-based framework for topology optimization that combines hierarchical conditioning by precomputed physics features with a differentiable connectivity penalty against floating material (Yang et al., 8 Jul 2026). In a broader methodological sense, closely related arXiv work applies analogous ideas—hierarchical conditioning, structured intermediate variables, or physics-constrained diffusion updates—to video frame interpolation, human motion generation, and multi-slice scientific imaging, although those systems use different names and do not share a single canonical formulation (Hai et al., 1 Apr 2025, Yuan et al., 2022, Valdy et al., 7 Dec 2025). Across these settings, the common premise is that diffusion is more reliable when its denoising trajectory is constrained by problem-structured physical information rather than left to operate solely in an unconstrained latent or pixel space.

1. Terminology and conceptual scope

In the strict sense used by the paper titled “HPG-Diff: Hierarchical physics-guided diffusion with differentiable connectivity constraints for topology optimization,” HPG-Diff is a generative method for compliance-driven topology optimization under varying boundary conditions, loads, and volume fractions (Yang et al., 8 Jul 2026). Its defining components are a DDPM-style UNet, hierarchical injection of three physics features—displacement UU, principal stress line (PSL), and strain energy density (SED)—and a Floating Material Suppression (FMS) loss.

A broader reading of the term is possible but must be qualified. The frame interpolation paper “Hierarchical Flow Diffusion for Efficient Frame Interpolation” replaces latent-image denoising with hierarchical diffusion over bilateral optical flow; this is a hierarchical and physically meaningful intermediate representation, but its “physics” is motion structure rather than an external simulator or mechanics prior (Hai et al., 1 Apr 2025). “PhysDiff: Physics-Guided Human Motion Diffusion Model” inserts a physics-based projection module into the diffusion sampling loop, yet the paper explicitly notes that it is not hierarchical in the classical coarse-to-fine or multi-resolution sense (Yuan et al., 2022). The multi-slice reconstruction framework presented as DART and DRIFT couples diffusion priors with modality-specific forward models and partitions the slice volume across GPUs; its hierarchy is architectural and computational, rather than a named HPG-Diff formalism (Valdy et al., 7 Dec 2025).

A common misconception is that “physics-guided diffusion” always denotes inference-time simulation inside each reverse step. That is not generally true. In topology optimization, the physics-guided components primarily shape training and conditioning, while inference remains simple stochastic diffusion sampling after one preprocessing FEA pass (Yang et al., 8 Jul 2026). By contrast, PhysDiff modifies the sampling loop directly through iterative projection in a physics simulator (Yuan et al., 2022).

2. Core architecture in topology optimization

The topology-optimization HPG-Diff model is built on a standard DDPM-style UNet with reverse transition

pθ(xt−1∣xt,c)=N(xt−1∣μθ(xt,t,c),Σθ(t)),p_\theta(x_{t-1} \mid x_t, c) = \mathcal{N}\left(x_{t-1} \mid \mu_\theta(x_t, t, c), \Sigma_\theta(t)\right),

where xtx_t is the noisy density field, xt−1x_{t-1} is the denoised sample, and cc denotes the physics-guidance condition (Yang et al., 8 Jul 2026). The model denoises a topology image or density map from noise to a binary-like material layout.

The central architectural idea is that the UNet already processes information hierarchically: shallow layers preserve local detail, mid layers capture intermediate structure, and deep layers encode global layout. HPG-Diff aligns this hierarchy with three precomputed FEA-derived features rather than concatenating all physics inputs into a single conditioning tensor. The three features are:

  • Displacement UU: a low-level local deformation field.
  • Principal Stress Line (PSL): a mid-level representation of load paths.
  • Strain Energy Density (SED): a high-level indicator of stiffness-critical regions.

These are obtained from an FEA solve on a uniform initial density field under prescribed loads and supports. The paper gives the equilibrium relation

KU=F,KU = F,

the compliance objective

c(x)=FTU=UTKU,c(x)=F^T U = U^T K U,

and the total potential energy

Π=12UTKU−UTF.\Pi = \frac{1}{2}U^T K U - U^T F.

SED is derived from strain energy, and PSL is computed from the stress tensor principal directions,

σ=[σxτxy τxyσy],σvi=σivi,i=1,2.\boldsymbol{\sigma} = \begin{bmatrix} \sigma_x & \tau_{xy}\ \tau_{xy} & \sigma_y \end{bmatrix}, \qquad \boldsymbol{\sigma}\mathbf{v}_i = \sigma_i \mathbf{v}_i,\quad i=1,2.

The layer assignment is explicit: pθ(xt−1∣xt,c)=N(xt−1∣μθ(xt,t,c),Σθ(t)),p_\theta(x_{t-1} \mid x_t, c) = \mathcal{N}\left(x_{t-1} \mid \mu_\theta(x_t, t, c), \Sigma_\theta(t)\right),0 is injected into shallow layers, PSL into mid layers, and SED into deep layers via layer-specific cross-attention modules. The intended effect is semantic alignment between the abstraction level of the physics signal and the abstraction level of the denoising network. The paper interprets this mechanically: pθ(xt−1∣xt,c)=N(xt−1∣μθ(xt,t,c),Σθ(t)),p_\theta(x_{t-1} \mid x_t, c) = \mathcal{N}\left(x_{t-1} \mid \mu_\theta(x_t, t, c), \Sigma_\theta(t)\right),1 provides fine-grained boundary and deformation guidance, PSL encodes skeleton-like load-transfer paths, and SED biases global material allocation toward stiffness-critical regions (Yang et al., 8 Jul 2026).

This arrangement is presented as the main reason the model generalizes more effectively to unseen boundary conditions than approaches conditioned only on a single global field such as SED. A plausible implication is that the hierarchy is not merely a representational convenience; it functions as an inductive bias that ties denoising decisions to mechanics-relevant scales.

3. Differentiable connectivity and floating material suppression

A major failure mode in generative topology optimization is floating material: disconnected dense regions that consume volume but do not contribute to load transfer. HPG-Diff addresses this with the Floating Material Suppression loss, a differentiable connectivity penalty inspired by virtual heat propagation from the load location (Yang et al., 8 Jul 2026).

The construction begins from a binary seed mask pθ(xt−1∣xt,c)=N(xt−1∣μθ(xt,t,c),Σθ(t)),p_\theta(x_{t-1} \mid x_t, c) = \mathcal{N}\left(x_{t-1} \mid \mu_\theta(x_t, t, c), \Sigma_\theta(t)\right),2 at the load position, the predicted density field pθ(xt−1∣xt,c)=N(xt−1∣μθ(xt,t,c),Σθ(t)),p_\theta(x_{t-1} \mid x_t, c) = \mathcal{N}\left(x_{t-1} \mid \mu_\theta(x_t, t, c), \Sigma_\theta(t)\right),3, and an iterative connectivity map pθ(xt−1∣xt,c)=N(xt−1∣μθ(xt,t,c),Σθ(t)),p_\theta(x_{t-1} \mid x_t, c) = \mathcal{N}\left(x_{t-1} \mid \mu_\theta(x_t, t, c), \Sigma_\theta(t)\right),4. The local propagation operator is defined over a pθ(xt−1∣xt,c)=N(xt−1∣μθ(xt,t,c),Σθ(t)),p_\theta(x_{t-1} \mid x_t, c) = \mathcal{N}\left(x_{t-1} \mid \mu_\theta(x_t, t, c), \Sigma_\theta(t)\right),5 neighborhood: pθ(xt−1∣xt,c)=N(xt−1∣μθ(xt,t,c),Σθ(t)),p_\theta(x_{t-1} \mid x_t, c) = \mathcal{N}\left(x_{t-1} \mid \mu_\theta(x_t, t, c), \Sigma_\theta(t)\right),6 Initialization is

pθ(xt−1∣xt,c)=N(xt−1∣μθ(xt,t,c),Σθ(t)),p_\theta(x_{t-1} \mid x_t, c) = \mathcal{N}\left(x_{t-1} \mid \mu_\theta(x_t, t, c), \Sigma_\theta(t)\right),7

Propagation then proceeds by

pθ(xt−1∣xt,c)=N(xt−1∣μθ(xt,t,c),Σθ(t)),p_\theta(x_{t-1} \mid x_t, c) = \mathcal{N}\left(x_{t-1} \mid \mu_\theta(x_t, t, c), \Sigma_\theta(t)\right),8

After pθ(xt−1∣xt,c)=N(xt−1∣μθ(xt,t,c),Σθ(t)),p_\theta(x_{t-1} \mid x_t, c) = \mathcal{N}\left(x_{t-1} \mid \mu_\theta(x_t, t, c), \Sigma_\theta(t)\right),9 steps, the final connectivity map is xtx_t0.

Conceptually, the mechanism acts as a reachability proxy. Propagation starts at the load seed and expands only through predicted solid regions. Dense areas connected to the seed receive high propagated values; dense disconnected islands do not. The paper makes the binary-limit interpretation explicit through a masked reachability dilation,

xtx_t1

The resulting loss is

xtx_t2

with time weighting

xtx_t3

The full training objective is

xtx_t4

The time weighting is operationally important. Early in denoising, the predicted density field is too noisy for connectivity to be meaningful, so the exponential term suppresses FMS. Near the end of the reverse chain, when topology becomes structured, the penalty strengthens. The paper states that this avoids gradient conflict with the standard diffusion noise loss. It also states that FMS is a soft differentiable penalty rather than a hard graph-theoretic guarantee of connectivity (Yang et al., 8 Jul 2026).

4. Training protocol, benchmarks, and quantitative performance

HPG-Diff is trained and evaluated on the topology optimization dataset from Mazé et al. / TopoDiff. The benchmark design space consists of xtx_t5 square grids with volume fraction xtx_t6 ranging from 0.3 to 0.5 in steps of 0.02, random loads on unconstrained boundary nodes, load directions over xtx_t7 in xtx_t8 increments, 42 boundary conditions used in training, and 5 unseen boundary conditions reserved for out-of-distribution testing (Yang et al., 8 Jul 2026). The dataset contains 30,000 training samples, 1,800 in-distribution test samples, and 1,000 out-of-distribution test samples.

Evaluation uses binarized outputs with threshold xtx_t9. The principal metrics are Compliance Error (CE), Volume Fraction Error (VFE), and Floating Material ratio (FM). On the in-distribution Test 1 split, HPG-Diff reports Avg CE xt−1x_{t-1}0, Med CE xt−1x_{t-1}1, Avg VFE xt−1x_{t-1}2, and FM xt−1x_{t-1}3. On the out-of-distribution Test 2 split, it reports Avg CE xt−1x_{t-1}4, Med CE xt−1x_{t-1}5, Avg VFE xt−1x_{t-1}6, and FM xt−1x_{t-1}7 (Yang et al., 8 Jul 2026).

Compared with the baselines listed in the paper, the margins are substantial. For Test 1, TopologyGAN reports Avg CE xt−1x_{t-1}8 and FM xt−1x_{t-1}9, TopoDiff-Guided reports Avg CE cc0 and FM cc1, and DOM w/ TA reports Avg CE cc2 and FM cc3. For Test 2, TopologyGAN reports Avg CE cc4 and FM cc5, TopoDiff-Guided reports Avg CE cc6 and FM cc7, and DOM w/ TA reports Avg CE cc8 and FM cc9.

The OOD result is especially emphasized. In the ablation study, removing hierarchical physics guidance causes Avg CE on Test 2 to jump from UU0 to UU1, indicating that the hierarchical conditioning rather than the diffusion backbone alone is the principal source of generalization under unseen boundary conditions (Yang et al., 8 Jul 2026). The paper also reports a long-tailed error distribution on Test 2: Std CE UU2, Max CE UU3, and CE UU4 for only UU5 of samples. This suggests that most generated samples remain accurate, while a small number of failures dominate the average.

The paper further compares HPG-Diff with iterative post-processing. DOM w/ TA + SIMPUU6 reports Mdn CE UU7 and FM UU8, DOM w/ TA + SIMPUU9 reports Mdn CE KU=F,KU = F,0 and FM KU=F,KU = F,1, and HPG-Diff reports Mdn CE KU=F,KU = F,2 and FM KU=F,KU = F,3. The stated significance is that HPG-Diff reaches competitive or better quality without requiring iterative post-processing during inference (Yang et al., 8 Jul 2026).

5. Adaptation, scope conditions, and limitations

Beyond the square benchmark domain, the paper studies adaptation to a KU=F,KU = F,4 3:1 rectangular domain using LoRA fine-tuning on 1,000 samples (Yang et al., 8 Jul 2026). The reported case studies include a shelf bracket in the square domain, a cantilever beam in the 3:1 domain, and a classic bridge in the 3:1 domain. For the shelf bracket, the ground-truth compliance is 5.0506 and generated variants are 4.9962, 5.0398, and 5.0772. For the cantilever beam, the ground-truth compliance is 182.8988 and generated results are 179.1958 and 185.5106. For the classic bridge, the ground-truth compliance is 13.0730 and generated results are 12.9581 and 12.9663.

On an independent 200-sample rectangular-domain evaluation, the LoRA-adapted model reports Avg CE KU=F,KU = F,5, Med CE KU=F,KU = F,6, Avg VFE KU=F,KU = F,7, and FM KU=F,KU = F,8. The paper characterizes this as preliminary evidence that lightweight adaptation can transfer a pretrained square-domain prior to non-square domains without retraining from scratch (Yang et al., 8 Jul 2026).

The limitations are explicit. The framework is still limited to the benchmark setting and requires broader validation on more diverse geometries and problems. It assumes a single-load benchmark when instantiating FMS from the load seed; multi-load problems would require a multi-seed extension. It does not yet include nonlinear materials, fatigue, dynamic loads, overhang limits, or minimum length scales. The FMS loss reduces floating material substantially but does not mathematically guarantee perfect connectivity. LoRA reduces adaptation cost but still requires target-domain solver-generated data (Yang et al., 8 Jul 2026).

A frequent misreading is that HPG-Diff replaces classical topology optimization. The more precise interpretation supported by the paper is that it moves generative topology optimization closer to classical physics-based optimization in quality while retaining generative efficiency and diversity.

The frame interpolation method “Hierarchical Flow Diffusion for Efficient Frame Interpolation” provides a closely related but domain-specific analogue (Hai et al., 1 Apr 2025). Its central move is to avoid denoising a large latent image space and instead denoise bilateral optical flow fields KU=F,KU = F,9 and c(x)=FTU=UTKU,c(x)=F^T U = U^T K U,0 in a coarse-to-fine hierarchy, conditioned on multiscale image features. The denoiser predicts

c(x)=FTU=UTKU,c(x)=F^T U = U^T K U,1

and the predicted flows drive a flow-guided image synthesizer,

c(x)=FTU=UTKU,c(x)=F^T U = U^T K U,2

The final rendered frame is

c(x)=FTU=UTKU,c(x)=F^T U = U^T K U,3

The paper reports state-of-the-art results on SNU-FILM, Xiph, DAVIS, and Vimeo-90k, as well as inference times on an RTX-4090 with c(x)=FTU=UTKU,c(x)=F^T U = U^T K U,4 input of 8.3 s for LDMVFI, 2.1 s for CBBD, 0.19 s for SGM-VFI, and 0.20 s for the proposed method, supporting the claim that it is c(x)=FTU=UTKU,c(x)=F^T U = U^T K U,5 faster than other diffusion-based methods (Hai et al., 1 Apr 2025). In this case, the hierarchy is explicitly coarse-to-fine, and the structured intermediate variable is bilateral optical flow.

PhysDiff occupies a different point in the design space (Yuan et al., 2022). It is best understood as a plug-and-play physics-guided diffusion sampler for human motion generation. A pretrained motion denoiser c(x)=FTU=UTKU,c(x)=F^T U = U^T K U,6 is wrapped by a physics-based projection operator c(x)=FTU=UTKU,c(x)=F^T U = U^T K U,7 implemented through a motion imitation policy in IsaacGym. The denoised motion is projected to a physically plausible motion and fed back into the next diffusion step. The method targets floating, foot sliding or skating, and ground penetration, summarized by

c(x)=FTU=UTKU,c(x)=F^T U = U^T K U,8

On HumanML3D, MDM reports FID 0.544 and Phys-Err 31.572, while PhysDiff w/ MDM reports FID 0.433 and Phys-Err 4.111; on UESTC, MDM reports Phys-Err 28.371, while PhysDiff w/ MDM reports 1.463 (Yuan et al., 2022). The paper explicitly states that this is a scheduled physics-guided diffusion method rather than a true hierarchical multi-stage method; more projection steps generally help physical plausibility, but applying them later in the diffusion process works best.

The multi-slice reconstruction framework described as DART and DRIFT supplies another related formulation (Valdy et al., 7 Dec 2025). It uses a video diffusion prior over a multi-slice object c(x)=FTU=UTKU,c(x)=F^T U = U^T K U,9, partitions slices into groups across GPUs, and alternates or combines diffusion denoising with modality-specific physics updates for MRI and 4D-STEM. For MRI, the physics step is a proximal gradient update on the undersampled Fourier fidelity term; for 4D-STEM, it is a proximal gradient update on intensity mismatch under the multislice forward model. DART interleaves diffusion and physics at each step, whereas DRIFT samples multiple diffusion candidates, selects the best initialization via SSIM, and then performs about 100 physics gradient steps. On MRI in-distribution evaluation, DART reports SSIM Π=12UTKU−UTF.\Pi = \frac{1}{2}U^T K U - U^T F.0, exceeding Projection-Based Π=12UTKU−UTF.\Pi = \frac{1}{2}U^T K U - U^T F.1, CS MRI Π=12UTKU−UTF.\Pi = \frac{1}{2}U^T K U - U^T F.2, and TV Π=12UTKU−UTF.\Pi = \frac{1}{2}U^T K U - U^T F.3; on OOD MRI Roots, DART reports SSIM Π=12UTKU−UTF.\Pi = \frac{1}{2}U^T K U - U^T F.4 (Valdy et al., 7 Dec 2025). Here the hierarchy is partitioned and distributed rather than layer-specific or simulator-scheduled.

Taken together, these papers suggest that “hierarchical physics-guided diffusion” is less a single algorithm than a recurring design principle. The hierarchy may be over UNet depth, flow scales, diffusion timesteps, or slice partitions; the physics guidance may be precomputed FEA fields, simulator projection, or forward-model data consistency. What remains stable across the literature surveyed here is the rejection of unconstrained denoising in favor of denoising trajectories that are aligned with domain structure and physically meaningful constraints (Yang et al., 8 Jul 2026, Hai et al., 1 Apr 2025, Yuan et al., 2022, Valdy et al., 7 Dec 2025).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Hierarchical Physics-Guided Diffusion (HPG-Diff).