Papers
Topics
Authors
Recent
Search
2000 character limit reached

HDFlow: Hierarchical Diffusion-Flow Patterns

Updated 6 July 2026
  • HDFlow is a design pattern that organizes generative processes into hierarchical levels using coarse-to-fine diffusion and structured latent states.
  • It leverages multiscale feature pyramids, staged refinement, and explicit flow or diffusion-only modules to reduce search space and boost performance.
  • Hybrid implementations show that combining diffusion with flow components can achieve state-of-the-art speed and accuracy in applications like video interpolation and graph generation.

Searching arXiv for the cited papers and related HDFlow terminology to ground the article. {"query":"ti:\"Hierarchical Flow Diffusion for Efficient Frame Interpolation\" OR ti:\"Hierarchical Diffusion Autoencoders and Disentangled Image Manipulation\" OR ti:\"HDFlow: Enhancing LLM Complex Problem-Solving with Hybrid Thinking and Dynamic Workflows\" OR ti:\"Simple Hierarchical Planning with Diffusion\" OR ti:\"Diffusion-Based Hierarchical Image Steganography\"","max_results":10,"sort_by":"relevance"} {"query":"(Lu et al., 2023, Hai et al., 1 Apr 2025, Chen et al., 2024, Xu et al., 2024, Boget et al., 31 Mar 2026, Zeng et al., 26 Sep 2025, Zang et al., 29 Jun 2025, Zhou et al., 8 Oct 2025, Cheng et al., 12 Nov 2025, Liu et al., 2024)","max_results":15,"sort_by":"submittedDate"} Hierarchical Diffusion-Flow (HDFlow) denotes a non-uniform research theme rather than a single canonical architecture. In its most direct generative-modeling sense, it refers to hierarchical, coarse-to-fine diffusion over a structured flow representation, most clearly exemplified by “Hierarchical Flow Diffusion for Efficient Frame Interpolation” (Hai et al., 1 Apr 2025). Adjacent work uses similar multiscale principles without an explicit flow component, or combines diffusion with flow in modular ways, while the paper titled “HDFlow” (Yao et al., 2024) concerns hybrid LLM reasoning rather than diffusion modeling. This suggests that HDFlow is best understood as a design pattern centered on multilevel structure, staged refinement, and structured intermediate variables.

1. Terminology and scope

Current arXiv usage separates at least four meanings that are easy to conflate. First, “Hierarchical Flow Diffusion” is an explicit video-frame-interpolation method built around bilateral optical flow and coarse-to-fine diffusion (Hai et al., 1 Apr 2025). Second, several hierarchical diffusion papers are conceptually close to HDFlow but explicitly state that they contain no normalizing flow, flow matching, or invertible transport machinery, as in Hierarchical Diffusion Autoencoders and Hierarchical Diffuser (Lu et al., 2023). Third, some hybrid systems genuinely combine diffusion and flow components, but define hierarchy in ways other than scale, such as importance tiers in steganography (Xu et al., 2024). Fourth, the title “HDFlow” itself is already occupied by a framework for hybrid fast/slow LLM reasoning, not by a hierarchical diffusion-flow generator (Yao et al., 2024).

Work Domain Relation to “HDFlow”
“Hierarchical Flow Diffusion for Efficient Frame Interpolation” (Hai et al., 1 Apr 2025) Video frame interpolation Most direct HDFlow-style usage
“Hierarchical Diffusion Autoencoders and Disentangled Image Manipulation” (Lu et al., 2023) Image autoencoding and editing Hierarchical diffusion, explicitly not a flow model
“Diffusion-Based Hierarchical Image Steganography” (Xu et al., 2024) Steganography Modular diffusion-plus-flow pipeline
“Hierarchical Discrete Flow Matching for Graph Generation” (Boget et al., 31 Mar 2026) Graph generation Close discrete HDFlow variant
“HDFlow: Enhancing LLM Complex Problem-Solving with Hybrid Thinking and Dynamic Workflows” (Yao et al., 2024) LLM reasoning Official HDFlow title, unrelated to diffusion modeling

A useful working definition therefore treats HDFlow as a family of architectures in which hierarchy constrains a diffusion-like or flow-like generation process through coarse-to-fine scales, structured latent states, or staged transport. The precise meaning of “hierarchy” varies substantially across papers.

2. Canonical formulation in frame interpolation

The clearest generative instantiation is the frame-interpolation system that models bilateral optical flow explicitly by hierarchical diffusion and then synthesizes the middle frame with a flow-guided decoder (Hai et al., 1 Apr 2025). Its central claim is that denoising in optical-flow space is more effective than denoising directly in image-latent space because the flow space is a much smaller search space in the denoising procedure.

The formulation uses bilateral flow anchored at the unknown middle frame. If I0I_0 and I1I_1 are the two input frames and ItI_t is the interpolated frame, the image synthesizer is

It=g(I0,I1,f0,f1;Φ),I_t = g(I_0, I_1, f_0, f_1; \Phi),

where f0f_0 and f1f_1 are the optical flows from ItI_t to I0I_0 and I1I_1. During training, pseudo bilateral flow is obtained from a pretrained RAFT model. The synthesizer is a multiscale ResNet-based encoder-decoder that predicts a 1-channel blending mask MM and a 3-channel residual map I1I_10, with final reconstruction

I1I_11

Hierarchy is realized as multi-scale coarse-to-fine diffusion over a feature pyramid. The encoder extracts feature pairs I1I_12, level I1I_13 has spatial resolution I1I_14 of the original image, and diffusion is performed across three pyramid levels, from I1I_15 to I1I_16 to I1I_17. At each level, the denoising U-Net predicts clean bilateral flow from noisy flow and conditioning features:

I1I_18

The diffusion timeline is partitioned across pyramid levels, and when the process moves to a finer stage the coarser prediction is upsampled by I1I_19 and re-noised before further denoising. The flow model is supervised at all scales with

ItI_t0

while the synthesizer uses

ItI_t1

with ItI_t2 and ItI_t3.

The architecture shares the denoising U-Net across feature levels, with separate feature and flow projectors per scale. After separate training of the synthesizer and the diffusion model, both are jointly fine-tuned end to end. Empirically, the method is reported as state of the art in accuracy and 10+ times faster than other diffusion-based methods; on the same RTX-4090 and ItI_t4 input pair, runtime is ItI_t5 s for LDMVFI, ItI_t6 s for CBBD, ItI_t7 s for SGM-VFI, and ItI_t8 s for the proposed method. For 256×256-resolution diffusion inference, the reported total runtime is ItI_t9 ms, broken down into encoder It=g(I0,I1,f0,f1;Φ),I_t = g(I_0, I_1, f_0, f_1; \Phi),0 ms, scale-It=g(I0,I1,f0,f1;Φ),I_t = g(I_0, I_1, f_0, f_1; \Phi),1 It=g(I0,I1,f0,f1;Φ),I_t = g(I_0, I_1, f_0, f_1; \Phi),2 ms, scale-It=g(I0,I1,f0,f1;Φ),I_t = g(I_0, I_1, f_0, f_1; \Phi),3 It=g(I0,I1,f0,f1;Φ),I_t = g(I_0, I_1, f_0, f_1; \Phi),4 ms, scale-It=g(I0,I1,f0,f1;Φ),I_t = g(I_0, I_1, f_0, f_1; \Phi),5 It=g(I0,I1,f0,f1;Φ),I_t = g(I_0, I_1, f_0, f_1; \Phi),6 ms, and decoder It=g(I0,I1,f0,f1;Φ),I_t = g(I_0, I_1, f_0, f_1; \Phi),7 ms (Hai et al., 1 Apr 2025).

3. Hierarchical latent organization in image, language, and tree generation

A second major line of work uses hierarchy to organize semantics across latent levels without adding a flow component. Hierarchical Diffusion Autoencoders replace the single bottleneck of earlier diffusion autoencoders with a coarse-to-fine feature hierarchy, typically derived from U-Net feature levels, and argue that higher hierarchy levels capture high-level semantics while lower levels capture fine appearance and detail (Lu et al., 2023). Linear probing on CelebA-HQ shows this scale-wise specialization: the high-level code is stronger on attributes such as Smile, Eyeglass, and Young, while the low-level code is stronger on Blackhair, PaleSkin, Brownhair, and 5_o_Clock_Shadow. A capacity-matched comparison is especially important: after 1000 training steps, validation reconstruction MSE is It=g(I0,I1,f0,f1;Φ),I_t = g(I_0, I_1, f_0, f_1; \Phi),8 for DAE(2560) and It=g(I0,I1,f0,f1;Φ),I_t = g(I_0, I_1, f_0, f_1; \Phi),9 for HDAE, while parameter counts are f0f_00M and f0f_01M respectively. The paper is explicit that this is not a flow model; the contribution is hierarchical latent organization for diffusion decoding, plus a truncated-feature-based approach for disentangled editing.

In language modeling, Hierarchical Diffusion LLMs define hierarchy over semantic abstraction levels rather than over spatial scale (Zhou et al., 8 Oct 2025). The main instantiation uses the chain

f0f_02

where f0f_03 is a word token, f0f_04 is a cluster-level token, and f0f_05 is mask. The forward marginal is

f0f_06

so corruption moves each token independently to a higher-level ancestor with coarser semantics. The reverse model always predicts word-level probabilities, but the ELBO decomposes into a cluster-to-word cross-entropy term and a mask-to-cluster cross-entropy term. The paper shows that MDLM is recovered as a special case when the hierarchy collapses to one cluster, so the method generalizes masked discrete diffusion by replacing flat noise with ancestor-based semantic noise.

HDTree introduces another non-flow interpretation of hierarchy, this time as a rooted binary tree over quantized latent codes (Zang et al., 29 Jun 2025). Inputs are encoded as f0f_07, then quantized into a root-to-leaf path

f0f_08

where each f0f_09 chooses the nearest valid child code at depth f1f_10. The diffusion process itself remains a conditional DDPM in data space, with f1f_11 acting as the conditioning variable. This produces a unified hierarchical codebook plus quantized diffusion process, and it avoids branch-specific modules by using one encoder, one shared tree-structured codebook, and one diffusion decoder.

4. Hybrid diffusion–flow variants and graph formulations

Some papers combine diffusion and flow more literally, but define hierarchy in domain-specific ways. In diffusion-based hierarchical image steganography, hierarchy is defined by payload importance and robustness rather than by multiresolution feature depth (Xu et al., 2024). Tier-1 content is concealed through a diffusion inversion backbone, then refined by a two-stage Enhance-Flow; Tier-2 images or text are embedded into the generated container through an invertible Embed-Flow. The result is a modular pipeline in which diffusion provides robust Tier-1 concealment and flow modules provide reversible enhancement and high-capacity secondary hiding. The paper explicitly notes that this is not a single unified probabilistic diffusion-flow model, but it is a clear hybrid architecture.

On the graph side, hierarchical discrete flow matching is close to an HDFlow formulation in a stricter algorithmic sense (Boget et al., 31 Mar 2026). Graph generation is factorized across levels as

f1f_12

where f1f_13 is a deterministic expansion of the coarser graph into a spanning supergraph at the finer level. Hierarchy is induced by node clustering and quotient-graph construction, while discrete flow matching refines each expanded sparse graph into its target graph. The main computational claim is that hierarchy reduces complexity from f1f_14 pairwise evaluation to f1f_15, and the reported number of function evaluations drops to f1f_16, f1f_17, or f1f_18, compared with f1f_19 or more in diffusion baselines. On ZINC250k, for example, dense DFM with 128 NFEs takes ItI_t0 s, while HDFM with 128 NFEs takes ItI_t1 s and HDFM with 32 NFEs takes ItI_t2 s.

A different graph formulation appears in continuous geometry-aware graph diffusion via hyperbolic neural PDE, where hierarchy is encoded geometrically by negative curvature rather than by explicit coarsening (Liu et al., 2024). The method reformulates propagation as a continuous-time diffusion-flow on the Poincaré ball, with node-wise attention acting as diffusivity inside a hyperbolic graph diffusion equation. Its local, global, and local-global diffusivity schemes operationalize diffusion over low- and high-order proximity. This is not a multiresolution HDFlow model, but it is an explicit continuous-time diffusion-flow framework for hierarchical graph structure.

Flow-only work can also inform the HDFlow concept. Fractal Flow introduces recursive normalizing flows with LDA-structured latent priors and a coarse-to-fine recursive architecture, but it is not a diffusion model (Zhang et al., 27 Aug 2025). Its relevance is architectural: it demonstrates that hierarchical latent semantics, recursive refinement, and interpretable module design can be built into exact generative models without branch-specific transforms.

5. Hierarchical planning and control

In offline reinforcement learning and planning, hierarchy is usually temporal rather than spatial. Hierarchical Diffuser defines a two-level planner in which a high-level Sparse Diffuser generates jumpy subgoals every ItI_t3 steps and a low-level diffusion model fills in dense segments between them (Chen et al., 2024). The hierarchy is simple and explicit: subgoals are every ItI_t4-th state, and lower-level generation is conditioned by clamping segment endpoints to adjacent subgoals. The method contains no flow component, but it shows that temporal abstraction can improve both receptive field and efficiency. On AntMaze-Large, Diffuser obtains ItI_t5 while HD reaches ItI_t6; on Maze2D-Medium, training time per 100 updates is ItI_t7 s for HD and ItI_t8 s for Diffuser.

CHD sharpens the planning argument by diagnosing the failure of weakly coupled hierarchy (Hao et al., 12 May 2025). The paper attributes long-horizon degradation to loose coupling between high-level sub-goal selection and low-level trajectory generation, then proposes Coupled Hierarchical Diffusion, in which HL sub-goals and LL trajectories are modeled within a unified hierarchical objective and a shared classifier passes LL feedback upstream so that sub-goals self-correct while sampling proceeds. The practical reverse process factorizes HL and LL denoising kernels, but restores bidirectional interaction through classifier guidance and an asynchronous parallel schedule. This is still diffusion-only, yet it is especially relevant to HDFlow because it makes coupling—not just multiscale decomposition—the central design requirement for coherent long-horizon plans.

SIHD extends the same area by replacing a fixed two-layer temporal hierarchy with an adaptively constructed one derived from state-graph structural information (Zeng et al., 26 Sep 2025). It builds a k-nearest-neighbor state graph, minimizes structural entropy to obtain an encoding tree, and uses the resulting communities to define multiple temporal scales. Lower diffusion layers are conditioned on structural information gain rather than only on reward prediction, and the base layer includes a structural entropy regularizer intended to encourage exploration of underrepresented states while avoiding extrapolation errors. Reported gains are largest on long-horizon navigation tasks: average improvements are ItI_t9 on single-task Maze2D, I0I_00 on multi-task Maze2D, and I0I_01 on AntMaze, with an I0I_02 reduction in training time and a I0I_03 reduction in planning time relative to Diffuser on Maze2D.

6. Empirical themes, misconceptions, and open directions

A recurring empirical theme is that hierarchical structure often matters more than merely enlarging a flat latent or increasing denoising budget. HDAE’s comparison against DAE(2560) shows better reconstruction at nearly identical parameter count, which the authors interpret as evidence that hierarchical conditioning is superior to simply enlarging a flat bottleneck (Lu et al., 2023). In graph generation, HDFM shows that a hierarchy can simultaneously reduce the search space and improve fidelity by constraining generation to sparse expanded supergraphs rather than dense arbitrary graphs (Boget et al., 31 Mar 2026). In planning, CHD argues that hierarchy without cross-level correction is insufficient, and SIHD argues that fixed hierarchy without adaptive temporal scales is insufficient (Hao et al., 12 May 2025).

Several misconceptions are therefore worth excluding. One is that every HDFlow-like method contains an explicit flow model. This is false for HDAE, Hierarchical Diffuser, SIHD, HDLM, and HDTree, all of which are diffusion-first or diffusion-only formulations (Yao et al., 2024). A second is that “hierarchy” always means the same thing. Across the cited literature it can mean multiscale feature pyramids, temporal abstraction, semantic-scale abstraction, quotient-graph coarsening, importance-dependent robustness tiers, or hyperbolic geometry (Xu et al., 2024). A third is that the title “HDFlow” uniquely identifies a hierarchical diffusion-flow generator; in current arXiv usage it also names a framework for hybrid fast/slow LLM reasoning (Yao et al., 2024).

The open research direction suggested by this body of work is not the replacement of diffusion by flow, but the integration of their strongest motifs. A plausible implication is that a more complete HDFlow formulation would combine adaptive hierarchy construction, explicit coupled coarse-to-fine inference, and either discrete or continuous transport across levels. The surveyed papers already provide most of these ingredients separately: structured hierarchical latent spaces, coarse-to-fine diffusion schedules, hybrid diffusion-plus-flow modules, discrete flow matching over hierarchical supports, and coupled multilevel planning objectives. What remains uneven across the literature is a single unified formulation that combines hierarchical latent design, cross-level correction, and explicit flow machinery within one end-to-end model.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Hierarchical Diffusion-Flow (HDFlow).