HDFlow: Hierarchical Diffusion-Flow Patterns
- HDFlow is a design pattern that organizes generative processes into hierarchical levels using coarse-to-fine diffusion and structured latent states.
- It leverages multiscale feature pyramids, staged refinement, and explicit flow or diffusion-only modules to reduce search space and boost performance.
- Hybrid implementations show that combining diffusion with flow components can achieve state-of-the-art speed and accuracy in applications like video interpolation and graph generation.
Searching arXiv for the cited papers and related HDFlow terminology to ground the article. {"query":"ti:\"Hierarchical Flow Diffusion for Efficient Frame Interpolation\" OR ti:\"Hierarchical Diffusion Autoencoders and Disentangled Image Manipulation\" OR ti:\"HDFlow: Enhancing LLM Complex Problem-Solving with Hybrid Thinking and Dynamic Workflows\" OR ti:\"Simple Hierarchical Planning with Diffusion\" OR ti:\"Diffusion-Based Hierarchical Image Steganography\"","max_results":10,"sort_by":"relevance"} {"query":"(Lu et al., 2023, Hai et al., 1 Apr 2025, Chen et al., 2024, Xu et al., 2024, Boget et al., 31 Mar 2026, Zeng et al., 26 Sep 2025, Zang et al., 29 Jun 2025, Zhou et al., 8 Oct 2025, Cheng et al., 12 Nov 2025, Liu et al., 2024)","max_results":15,"sort_by":"submittedDate"} Hierarchical Diffusion-Flow (HDFlow) denotes a non-uniform research theme rather than a single canonical architecture. In its most direct generative-modeling sense, it refers to hierarchical, coarse-to-fine diffusion over a structured flow representation, most clearly exemplified by “Hierarchical Flow Diffusion for Efficient Frame Interpolation” (Hai et al., 1 Apr 2025). Adjacent work uses similar multiscale principles without an explicit flow component, or combines diffusion with flow in modular ways, while the paper titled “HDFlow” (Yao et al., 2024) concerns hybrid LLM reasoning rather than diffusion modeling. This suggests that HDFlow is best understood as a design pattern centered on multilevel structure, staged refinement, and structured intermediate variables.
1. Terminology and scope
Current arXiv usage separates at least four meanings that are easy to conflate. First, “Hierarchical Flow Diffusion” is an explicit video-frame-interpolation method built around bilateral optical flow and coarse-to-fine diffusion (Hai et al., 1 Apr 2025). Second, several hierarchical diffusion papers are conceptually close to HDFlow but explicitly state that they contain no normalizing flow, flow matching, or invertible transport machinery, as in Hierarchical Diffusion Autoencoders and Hierarchical Diffuser (Lu et al., 2023). Third, some hybrid systems genuinely combine diffusion and flow components, but define hierarchy in ways other than scale, such as importance tiers in steganography (Xu et al., 2024). Fourth, the title “HDFlow” itself is already occupied by a framework for hybrid fast/slow LLM reasoning, not by a hierarchical diffusion-flow generator (Yao et al., 2024).
| Work | Domain | Relation to “HDFlow” |
|---|---|---|
| “Hierarchical Flow Diffusion for Efficient Frame Interpolation” (Hai et al., 1 Apr 2025) | Video frame interpolation | Most direct HDFlow-style usage |
| “Hierarchical Diffusion Autoencoders and Disentangled Image Manipulation” (Lu et al., 2023) | Image autoencoding and editing | Hierarchical diffusion, explicitly not a flow model |
| “Diffusion-Based Hierarchical Image Steganography” (Xu et al., 2024) | Steganography | Modular diffusion-plus-flow pipeline |
| “Hierarchical Discrete Flow Matching for Graph Generation” (Boget et al., 31 Mar 2026) | Graph generation | Close discrete HDFlow variant |
| “HDFlow: Enhancing LLM Complex Problem-Solving with Hybrid Thinking and Dynamic Workflows” (Yao et al., 2024) | LLM reasoning | Official HDFlow title, unrelated to diffusion modeling |
A useful working definition therefore treats HDFlow as a family of architectures in which hierarchy constrains a diffusion-like or flow-like generation process through coarse-to-fine scales, structured latent states, or staged transport. The precise meaning of “hierarchy” varies substantially across papers.
2. Canonical formulation in frame interpolation
The clearest generative instantiation is the frame-interpolation system that models bilateral optical flow explicitly by hierarchical diffusion and then synthesizes the middle frame with a flow-guided decoder (Hai et al., 1 Apr 2025). Its central claim is that denoising in optical-flow space is more effective than denoising directly in image-latent space because the flow space is a much smaller search space in the denoising procedure.
The formulation uses bilateral flow anchored at the unknown middle frame. If and are the two input frames and is the interpolated frame, the image synthesizer is
where and are the optical flows from to and . During training, pseudo bilateral flow is obtained from a pretrained RAFT model. The synthesizer is a multiscale ResNet-based encoder-decoder that predicts a 1-channel blending mask and a 3-channel residual map 0, with final reconstruction
1
Hierarchy is realized as multi-scale coarse-to-fine diffusion over a feature pyramid. The encoder extracts feature pairs 2, level 3 has spatial resolution 4 of the original image, and diffusion is performed across three pyramid levels, from 5 to 6 to 7. At each level, the denoising U-Net predicts clean bilateral flow from noisy flow and conditioning features:
8
The diffusion timeline is partitioned across pyramid levels, and when the process moves to a finer stage the coarser prediction is upsampled by 9 and re-noised before further denoising. The flow model is supervised at all scales with
0
while the synthesizer uses
1
with 2 and 3.
The architecture shares the denoising U-Net across feature levels, with separate feature and flow projectors per scale. After separate training of the synthesizer and the diffusion model, both are jointly fine-tuned end to end. Empirically, the method is reported as state of the art in accuracy and 10+ times faster than other diffusion-based methods; on the same RTX-4090 and 4 input pair, runtime is 5 s for LDMVFI, 6 s for CBBD, 7 s for SGM-VFI, and 8 s for the proposed method. For 256×256-resolution diffusion inference, the reported total runtime is 9 ms, broken down into encoder 0 ms, scale-1 2 ms, scale-3 4 ms, scale-5 6 ms, and decoder 7 ms (Hai et al., 1 Apr 2025).
3. Hierarchical latent organization in image, language, and tree generation
A second major line of work uses hierarchy to organize semantics across latent levels without adding a flow component. Hierarchical Diffusion Autoencoders replace the single bottleneck of earlier diffusion autoencoders with a coarse-to-fine feature hierarchy, typically derived from U-Net feature levels, and argue that higher hierarchy levels capture high-level semantics while lower levels capture fine appearance and detail (Lu et al., 2023). Linear probing on CelebA-HQ shows this scale-wise specialization: the high-level code is stronger on attributes such as Smile, Eyeglass, and Young, while the low-level code is stronger on Blackhair, PaleSkin, Brownhair, and 5_o_Clock_Shadow. A capacity-matched comparison is especially important: after 1000 training steps, validation reconstruction MSE is 8 for DAE(2560) and 9 for HDAE, while parameter counts are 0M and 1M respectively. The paper is explicit that this is not a flow model; the contribution is hierarchical latent organization for diffusion decoding, plus a truncated-feature-based approach for disentangled editing.
In language modeling, Hierarchical Diffusion LLMs define hierarchy over semantic abstraction levels rather than over spatial scale (Zhou et al., 8 Oct 2025). The main instantiation uses the chain
2
where 3 is a word token, 4 is a cluster-level token, and 5 is mask. The forward marginal is
6
so corruption moves each token independently to a higher-level ancestor with coarser semantics. The reverse model always predicts word-level probabilities, but the ELBO decomposes into a cluster-to-word cross-entropy term and a mask-to-cluster cross-entropy term. The paper shows that MDLM is recovered as a special case when the hierarchy collapses to one cluster, so the method generalizes masked discrete diffusion by replacing flat noise with ancestor-based semantic noise.
HDTree introduces another non-flow interpretation of hierarchy, this time as a rooted binary tree over quantized latent codes (Zang et al., 29 Jun 2025). Inputs are encoded as 7, then quantized into a root-to-leaf path
8
where each 9 chooses the nearest valid child code at depth 0. The diffusion process itself remains a conditional DDPM in data space, with 1 acting as the conditioning variable. This produces a unified hierarchical codebook plus quantized diffusion process, and it avoids branch-specific modules by using one encoder, one shared tree-structured codebook, and one diffusion decoder.
4. Hybrid diffusion–flow variants and graph formulations
Some papers combine diffusion and flow more literally, but define hierarchy in domain-specific ways. In diffusion-based hierarchical image steganography, hierarchy is defined by payload importance and robustness rather than by multiresolution feature depth (Xu et al., 2024). Tier-1 content is concealed through a diffusion inversion backbone, then refined by a two-stage Enhance-Flow; Tier-2 images or text are embedded into the generated container through an invertible Embed-Flow. The result is a modular pipeline in which diffusion provides robust Tier-1 concealment and flow modules provide reversible enhancement and high-capacity secondary hiding. The paper explicitly notes that this is not a single unified probabilistic diffusion-flow model, but it is a clear hybrid architecture.
On the graph side, hierarchical discrete flow matching is close to an HDFlow formulation in a stricter algorithmic sense (Boget et al., 31 Mar 2026). Graph generation is factorized across levels as
2
where 3 is a deterministic expansion of the coarser graph into a spanning supergraph at the finer level. Hierarchy is induced by node clustering and quotient-graph construction, while discrete flow matching refines each expanded sparse graph into its target graph. The main computational claim is that hierarchy reduces complexity from 4 pairwise evaluation to 5, and the reported number of function evaluations drops to 6, 7, or 8, compared with 9 or more in diffusion baselines. On ZINC250k, for example, dense DFM with 128 NFEs takes 0 s, while HDFM with 128 NFEs takes 1 s and HDFM with 32 NFEs takes 2 s.
A different graph formulation appears in continuous geometry-aware graph diffusion via hyperbolic neural PDE, where hierarchy is encoded geometrically by negative curvature rather than by explicit coarsening (Liu et al., 2024). The method reformulates propagation as a continuous-time diffusion-flow on the Poincaré ball, with node-wise attention acting as diffusivity inside a hyperbolic graph diffusion equation. Its local, global, and local-global diffusivity schemes operationalize diffusion over low- and high-order proximity. This is not a multiresolution HDFlow model, but it is an explicit continuous-time diffusion-flow framework for hierarchical graph structure.
Flow-only work can also inform the HDFlow concept. Fractal Flow introduces recursive normalizing flows with LDA-structured latent priors and a coarse-to-fine recursive architecture, but it is not a diffusion model (Zhang et al., 27 Aug 2025). Its relevance is architectural: it demonstrates that hierarchical latent semantics, recursive refinement, and interpretable module design can be built into exact generative models without branch-specific transforms.
5. Hierarchical planning and control
In offline reinforcement learning and planning, hierarchy is usually temporal rather than spatial. Hierarchical Diffuser defines a two-level planner in which a high-level Sparse Diffuser generates jumpy subgoals every 3 steps and a low-level diffusion model fills in dense segments between them (Chen et al., 2024). The hierarchy is simple and explicit: subgoals are every 4-th state, and lower-level generation is conditioned by clamping segment endpoints to adjacent subgoals. The method contains no flow component, but it shows that temporal abstraction can improve both receptive field and efficiency. On AntMaze-Large, Diffuser obtains 5 while HD reaches 6; on Maze2D-Medium, training time per 100 updates is 7 s for HD and 8 s for Diffuser.
CHD sharpens the planning argument by diagnosing the failure of weakly coupled hierarchy (Hao et al., 12 May 2025). The paper attributes long-horizon degradation to loose coupling between high-level sub-goal selection and low-level trajectory generation, then proposes Coupled Hierarchical Diffusion, in which HL sub-goals and LL trajectories are modeled within a unified hierarchical objective and a shared classifier passes LL feedback upstream so that sub-goals self-correct while sampling proceeds. The practical reverse process factorizes HL and LL denoising kernels, but restores bidirectional interaction through classifier guidance and an asynchronous parallel schedule. This is still diffusion-only, yet it is especially relevant to HDFlow because it makes coupling—not just multiscale decomposition—the central design requirement for coherent long-horizon plans.
SIHD extends the same area by replacing a fixed two-layer temporal hierarchy with an adaptively constructed one derived from state-graph structural information (Zeng et al., 26 Sep 2025). It builds a k-nearest-neighbor state graph, minimizes structural entropy to obtain an encoding tree, and uses the resulting communities to define multiple temporal scales. Lower diffusion layers are conditioned on structural information gain rather than only on reward prediction, and the base layer includes a structural entropy regularizer intended to encourage exploration of underrepresented states while avoiding extrapolation errors. Reported gains are largest on long-horizon navigation tasks: average improvements are 9 on single-task Maze2D, 0 on multi-task Maze2D, and 1 on AntMaze, with an 2 reduction in training time and a 3 reduction in planning time relative to Diffuser on Maze2D.
6. Empirical themes, misconceptions, and open directions
A recurring empirical theme is that hierarchical structure often matters more than merely enlarging a flat latent or increasing denoising budget. HDAE’s comparison against DAE(2560) shows better reconstruction at nearly identical parameter count, which the authors interpret as evidence that hierarchical conditioning is superior to simply enlarging a flat bottleneck (Lu et al., 2023). In graph generation, HDFM shows that a hierarchy can simultaneously reduce the search space and improve fidelity by constraining generation to sparse expanded supergraphs rather than dense arbitrary graphs (Boget et al., 31 Mar 2026). In planning, CHD argues that hierarchy without cross-level correction is insufficient, and SIHD argues that fixed hierarchy without adaptive temporal scales is insufficient (Hao et al., 12 May 2025).
Several misconceptions are therefore worth excluding. One is that every HDFlow-like method contains an explicit flow model. This is false for HDAE, Hierarchical Diffuser, SIHD, HDLM, and HDTree, all of which are diffusion-first or diffusion-only formulations (Yao et al., 2024). A second is that “hierarchy” always means the same thing. Across the cited literature it can mean multiscale feature pyramids, temporal abstraction, semantic-scale abstraction, quotient-graph coarsening, importance-dependent robustness tiers, or hyperbolic geometry (Xu et al., 2024). A third is that the title “HDFlow” uniquely identifies a hierarchical diffusion-flow generator; in current arXiv usage it also names a framework for hybrid fast/slow LLM reasoning (Yao et al., 2024).
The open research direction suggested by this body of work is not the replacement of diffusion by flow, but the integration of their strongest motifs. A plausible implication is that a more complete HDFlow formulation would combine adaptive hierarchy construction, explicit coupled coarse-to-fine inference, and either discrete or continuous transport across levels. The surveyed papers already provide most of these ingredients separately: structured hierarchical latent spaces, coarse-to-fine diffusion schedules, hybrid diffusion-plus-flow modules, discrete flow matching over hierarchical supports, and coupled multilevel planning objectives. What remains uneven across the literature is a single unified formulation that combines hierarchical latent design, cross-level correction, and explicit flow machinery within one end-to-end model.