Papers
Topics
Authors
Recent
Search
2000 character limit reached

Hybrid-Balance GFlowNet (HBG) Overview

Updated 14 July 2026
  • The paper introduces HBG as a framework that hybridizes trajectory balance and detailed balance to blend global credit assignment with local feedback.
  • HBG employs measure-theoretic corrections, including Radon–Nikodym derivatives, to manage hybrid discrete–continuous state spaces.
  • The framework is applied to vehicle routing, demonstrating improved performance in CVRP and TSP by reducing variance and bias in training signals.

Searching arXiv for the cited HBG-related papers to ground the article in current records. arxiv_search.query{"search_query":"id:(Lahlou et al., 2023) OR id:(Madan et al., 2022) OR id:(Zhang et al., 6 Oct 2025) OR id:(Chen et al., 2 Feb 2026)","max_results":10} arxiv_search.search{"query":"id:(Lahlou et al., 2023) OR id:(Madan et al., 2022) OR id:(Zhang et al., 6 Oct 2025) OR id:(Chen et al., 2 Feb 2026)","max_results":10} arxiv.search{"query":"(Lahlou et al., 2023, Madan et al., 2022, Zhang et al., 6 Oct 2025, Chen et al., 2 Feb 2026)","max_results":10} Hybrid-Balance GFlowNet (HBG) denotes a class of Generative Flow Network formulations in which distinct balance mechanisms are combined, reweighted, or made explicit in order to preserve reward-proportional sampling while improving credit assignment, extending the state space, or controlling exploration and exploitation. In explicit nomenclature, HBG is a framework for vehicle routing that integrates Trajectory Balance (TB) and Detailed Balance (DB) in a single training and inference scheme (Zhang et al., 6 Oct 2025). In generalized continuous-state GFlowNet theory, the same expression naturally refers to TB or DB instantiated on hybrid discrete–continuous spaces with Radon–Nikodym derivatives and change-of-variables corrections (Lahlou et al., 2023). Related literature also maps the idea of “hybrid balance” to λ\lambda-weighted mixtures of subtrajectory constraints and to α\alpha-weighted forward/backward mixing, although those papers do not explicitly introduce the HBG name (Madan et al., 2022, Chen et al., 2 Feb 2026). Taken together, these works suggest that HBG is best understood not as a single universally fixed objective, but as a family of balance constructions that hybridize local and global consistency, discrete and continuous transitions, or forward and backward dynamics.

1. Scope, nomenclature, and canonical GFlowNet setting

A GFlowNet learns a forward policy PFP_F over a constructive process so that terminal objects xx are sampled proportionally to a nonnegative reward R(x)R(x). In the DAG formulation, a trajectory τ=(s0,,sT=x)\tau=(s_0,\ldots,s_T=x) satisfies

PF(τ)=t=0T1PF(st+1st),P_F(\tau)=\prod_{t=0}^{T-1} P_F(s_{t+1}\mid s_t),

and the target condition is

R(x)=Zτ:sT=xPF(τ),R(x)=Z\sum_{\tau:s_T=x} P_F(\tau),

with Z=xXR(x)Z=\sum_{x\in X}R(x) at optimum (Madan et al., 2022). In the combinatorial optimization formulation used for vehicle routing problems, a state ss encodes a partial construction, an action appends the next node subject to feasibility, the backward policy α\alpha0 reverses a forward step, the flow α\alpha1 is a scalar learned on each state, and the partition function α\alpha2 is the source flow at α\alpha3 (Zhang et al., 6 Oct 2025).

Within this general setting, the expression “Hybrid-Balance GFlowNet” has multiple technical uses. The vehicle-routing formulation introduces HBG as a training and inference framework that “uniquely integrates TB and DB in a principled and adaptive manner” and augments AGFN and GFACS (Zhang et al., 6 Oct 2025). The continuous-state theory describes HBG as the trajectory-balance objective specialized to hybrid spaces, where continuous corrections must be made explicit when a policy is parameterized through actions and deterministic transformations (Lahlou et al., 2023). The SubTB(α\alpha4) work states that it does not explicitly introduce or name HBG, but characterizes SubTB(α\alpha5) as embodying a hybrid-balance principle because it blends local and trajectory-wide training signals through a α\alpha6-weighted mixture of subtrajectories (Madan et al., 2022). The α\alpha7-GFN work likewise states that it does not use the term HBG explicitly, but presents a tunable hybridization of forward and backward components through

α\alpha8

which is interpreted there as a hybrid-balance construction (Chen et al., 2 Feb 2026).

2. Balance laws from which HBG is constructed

The central ingredients of HBG are the standard GFlowNet balance objectives. For a complete trajectory α\alpha9, TB imposes

PFP_F0

or equivalently

PFP_F1

DB instead enforces local edge constraints,

PFP_F2

with terminal condition PFP_F3. FM imposes local conservation of edge flows,

PFP_F4

and uses the induced forward policy

PFP_F5

These objectives were originally contrasted as trajectory-wide, transition-wise, and state-local constraints (Madan et al., 2022).

In measurable-state GFlowNet theory, the same ideas are expressed with measures and kernels. A flow is a pair PFP_F6 with PFP_F7 and PFP_F8, while the backward kernel satisfies PFP_F9. The resulting Radon–Nikodym derivatives

xx0

unify discrete probabilities and continuous densities (Lahlou et al., 2023). In this generalized formulation, HBG is obtained by instantiating TB or DB in hybrid spaces and, when actions rather than state kernels are parameterized directly, by inserting the correct Radon–Nikodym or Jacobian terms.

A recurrent theme across these balance laws is the local–global tradeoff. DB and FM yield low-variance but biased training signals because credit is assigned locally, whereas TB yields low-bias global credit assignment but higher variance because it propagates reward over full trajectories (Madan et al., 2022). HBG variants differ in where they intervene in this tradeoff: some make continuous corrections explicit, some interpolate across subtrajectory lengths, and some directly sum local and global losses.

3. HBG in hybrid discrete–continuous state spaces

The generalized continuous-state theory models a hybrid state space through a measurable pointed graph

xx1

with distinguished source xx2 and sink xx3. A hybrid state space can be instantiated as

xx4

where xx5 is countable or finite and xx6 is a manifold or Euclidean subset; xx7 is the Borel xx8-algebra induced by the disjoint union topology; and the reference measure is naturally chosen as counting measure on xx9 Lebesgue measure on R(x)R(x)0, plus Dirac masses at R(x)R(x)1 and R(x)R(x)2 (Lahlou et al., 2023).

The measure-theoretic FM condition is

R(x)R(x)3

for bounded measurable R(x)R(x)4 with R(x)R(x)5. In density form this becomes

R(x)R(x)6

R(x)R(x)7-almost surely on R(x)R(x)8. Terminal balance is expressed by reward matching,

R(x)R(x)9

or, in densities,

τ=(s0,,sT=x)\tau=(s_0,\ldots,s_T=x)0

τ=(s0,,sT=x)\tau=(s_0,\ldots,s_T=x)1-almost surely on τ=(s0,,sT=x)\tau=(s_0,\ldots,s_T=x)2 (Lahlou et al., 2023). The paper’s correctness theorem states that if FM and reward matching hold with respect to a positive finite τ=(s0,,sT=x)\tau=(s_0,\ldots,s_T=x)3, then the terminating measure τ=(s0,,sT=x)\tau=(s_0,\ldots,s_T=x)4 is a probability measure and

τ=(s0,,sT=x)\tau=(s_0,\ldots,s_T=x)5

for all measurable τ=(s0,,sT=x)\tau=(s_0,\ldots,s_T=x)6.

The distinctive HBG issue arises when the forward policy is parameterized in an action space

τ=(s0,,sT=x)\tau=(s_0,\ldots,s_T=x)7

and then pushed forward to next states through a deterministic or stochastic mechanism τ=(s0,,sT=x)\tau=(s_0,\ldots,s_T=x)8. In that case the induced state-transition density must include the continuous correction

τ=(s0,,sT=x)\tau=(s_0,\ldots,s_T=x)9

where

PF(τ)=t=0T1PF(st+1st),P_F(\tau)=\prod_{t=0}^{T-1} P_F(s_{t+1}\mid s_t),0

For diffeomorphic continuous maps with Lebesgue reference measure, this simplifies to

PF(τ)=t=0T1PF(st+1st),P_F(\tau)=\prod_{t=0}^{T-1} P_F(s_{t+1}\mid s_t),1

whereas for discrete branches PF(τ)=t=0T1PF(st+1st),P_F(\tau)=\prod_{t=0}^{T-1} P_F(s_{t+1}\mid s_t),2 on the corresponding atom of PF(τ)=t=0T1PF(st+1st),P_F(\tau)=\prod_{t=0}^{T-1} P_F(s_{t+1}\mid s_t),3 (Lahlou et al., 2023).

The hybrid trajectory-balance equality for

PF(τ)=t=0T1PF(st+1st),P_F(\tau)=\prod_{t=0}^{T-1} P_F(s_{t+1}\mid s_t),4

is

PF(τ)=t=0T1PF(st+1st),P_F(\tau)=\prod_{t=0}^{T-1} P_F(s_{t+1}\mid s_t),5

with loss

PF(τ)=t=0T1PF(st+1st),P_F(\tau)=\prod_{t=0}^{T-1} P_F(s_{t+1}\mid s_t),6

When one works directly in state space with PF(τ)=t=0T1PF(st+1st),P_F(\tau)=\prod_{t=0}^{T-1} P_F(s_{t+1}\mid s_t),7, the correction PF(τ)=t=0T1PF(st+1st),P_F(\tau)=\prod_{t=0}^{T-1} P_F(s_{t+1}\mid s_t),8 is already included in PF(τ)=t=0T1PF(st+1st),P_F(\tau)=\prod_{t=0}^{T-1} P_F(s_{t+1}\mid s_t),9, and HBG reduces to the standard TB loss. The practical significance is narrow but important: HBG is useful precisely when the policy is modeled in action space and one must translate it to the state-space kernel correctly.

4. HBG as interpolation between local and trajectory-wide credit assignment

Subtrajectory Balance introduces a different kind of hybridization. For any contiguous subtrajectory

R(x)=Zτ:sT=xPF(τ),R(x)=Z\sum_{\tau:s_T=x} P_F(\tau),0

the paper shows that DB is equivalent to the subtrajectory constraint

R(x)=Zτ:sT=xPF(τ),R(x)=Z\sum_{\tau:s_T=x} P_F(\tau),1

with R(x)=Zτ:sT=xPF(τ),R(x)=Z\sum_{\tau:s_T=x} P_F(\tau),2 if R(x)=Zτ:sT=xPF(τ),R(x)=Z\sum_{\tau:s_T=x} P_F(\tau),3 is terminal. The associated loss is

R(x)=Zτ:sT=xPF(τ),R(x)=Z\sum_{\tau:s_T=x} P_F(\tau),4

SubTB(R(x)=Zτ:sT=xPF(τ),R(x)=Z\sum_{\tau:s_T=x} P_F(\tau),5) then aggregates subtrajectory losses as

R(x)=Zτ:sT=xPF(τ),R(x)=Z\sum_{\tau:s_T=x} P_F(\tau),6

where R(x)=Zτ:sT=xPF(τ),R(x)=Z\sum_{\tau:s_T=x} P_F(\tau),7 weighs by subtrajectory length (Madan et al., 2022).

The limiting cases recover the two standard extremes. As R(x)=Zτ:sT=xPF(τ),R(x)=Z\sum_{\tau:s_T=x} P_F(\tau),8, only 1-step subtrajectories contribute, matching the average DB loss over edges. As R(x)=Zτ:sT=xPF(τ),R(x)=Z\sum_{\tau:s_T=x} P_F(\tau),9, the longest subtrajectory dominates, recovering TB. At Z=xXR(x)Z=\sum_{x\in X}R(x)0, all subtrajectories receive uniform weight. The gradient can be computed with one forward and one backward pass through the networks for Z=xXR(x)Z=\sum_{x\in X}R(x)1, Z=xXR(x)Z=\sum_{x\in X}R(x)2, and Z=xXR(x)Z=\sum_{x\in X}R(x)3; Z=xXR(x)Z=\sum_{x\in X}R(x)4 linear operations combine logits across subtrajectories, while deep network evaluation remains Z=xXR(x)Z=\sum_{x\in X}R(x)5.

The paper does not explicitly introduce or name HBG, but it states that SubTB(Z=xXR(x)Z=\sum_{x\in X}R(x)6) “embodies the hybrid-balance principle by blending local (edge/state) and trajectory-wide training signals through a Z=xXR(x)Z=\sum_{x\in X}R(x)7-weighted mixture of subtrajectories” (Madan et al., 2022). That interpretation is supported by the empirical bias–variance analysis: DB has the highest self-consistency and lowest variance, TB the lowest self-consistency and highest variance, and SubTB(Z=xXR(x)Z=\sum_{x\in X}R(x)8) lies in between. On hypergrid tasks, SubTB(Z=xXR(x)Z=\sum_{x\in X}R(x)9) converges faster and with less variability across seeds than TB, and in the very sparse setting with background reward ss0, TB fails to discover all modes beyond ss1, whereas SubTB(ss2) still finds all and matches the target distribution well. On AMP, SubTB(ss3) attains reward ss4 and diversity ss5, compared with TB at ss6 and ss7; on GFP, SubTB(ss8) reaches reward ss9 with diversity α\alpha00, while TB reaches α\alpha01 with diversity α\alpha02 (Madan et al., 2022). These results place HBG, in this interpretive sense, within the broader program of balancing gradient bias and variance rather than merely combining two named losses.

5. Explicit HBG for vehicle routing problems

The 2025 vehicle-routing paper introduces Hybrid-Balance GFlowNet as a solver framework for CVRP and TSP. Its stated premise is that TB is well aligned with the global objective of minimizing total tour length but yields diffuse credit assignment in long-horizon VRPs, whereas DB produces strong local feedback but lacks a global perspective. HBG therefore combines the two in a single objective,

α\alpha03

and also studies a weighted variant

α\alpha04

Empirically, a fixed α\alpha05 achieves a favorable balance across CVRP and TSP in AGFN and GFACS (Zhang et al., 6 Oct 2025).

The TB component uses the solver-specific shaped reward α\alpha06:

α\alpha07

The DB component is defined stepwise:

α\alpha08

with trajectory loss α\alpha09. The local energy term is

α\alpha10

where α\alpha11 is the local transition cost, namely the distance of the last edge in α\alpha12.

The framework also specifies VRP-specific backward probabilities. At trajectory level, if α\alpha13 contains a multi-route decomposition with multi-node sub-route count α\alpha14 and single-node sub-route count α\alpha15, then

α\alpha16

At the step level,

α\alpha17

The learned flow head is

α\alpha18

A further distinctive feature is depot-centric inference for CVRP:

α\alpha19

Feasibility is enforced by masking unvisited customers whose demand exceeds the remaining capacity, and depot return is forced when no feasible unvisited customer remains. This asymmetry is motivated by the observation that only the depot has multiple valid predecessors in the backward dynamics.

The reported results are consistently favorable. On synthetic CVRP, AGFN improves from gap α\alpha20 to α\alpha21 at α\alpha22, from α\alpha23 to α\alpha24 at α\alpha25, and from α\alpha26 to α\alpha27 at α\alpha28 when HBG is added. GFACS improves from α\alpha29 to α\alpha30, from α\alpha31 to α\alpha32, and from α\alpha33 to α\alpha34 at the same sizes; with local search, the corresponding improvements are smaller but still present, for example α\alpha35 to α\alpha36 at α\alpha37. On TSP, AGFN improves from α\alpha38 to α\alpha39 at α\alpha40, and GFACS improves from α\alpha41 to α\alpha42. Runtime overhead is reported as negligible, with AGFN adding α\alpha43–α\alpha44s and GFACS remaining unchanged within measurement precision. Ablations show that DB-only is worse than TB and HBG, and that fixed α\alpha45 yields the most stable improvements across CVRP and TSP (Zhang et al., 6 Oct 2025).

6. Markov-chain reinterpretation and α\alpha46-hybridization

A different formalization of HBG emerges from the Markov-chain perspective on GFlowNets. The α\alpha47-GFN paper shows that standard GFlowNet objectives correspond to reversibility of the equally mixed kernel

α\alpha48

and generalizes this to

α\alpha49

The one-step reversibility relation is

α\alpha50

and the flows act as an unnormalized probability measure through

α\alpha51

For a partial trajectory segment α\alpha52, the α\alpha53-SubTB target is

α\alpha54

with analogous α\alpha55-DB, α\alpha56-TB, and forward-looking variants (Chen et al., 2 Feb 2026).

The paper states that these α\alpha57-objectives are equivalent to reversibility of the Markov chain with kernel α\alpha58, and that their convergence to unique flows is similar to vanilla objectives for all α\alpha59. The conditions used include finite state space, a pointed DAG with source and sink, irreducibility and positive recurrence of the induced Markov chain, and positive rewards on terminal states so that α\alpha60 is well defined.

The practical role of α\alpha61 is to control exploration and exploitation. For α\alpha62-SubTB, the gradient-level characterization is

α\alpha63

For α\alpha64, the added term is positive and larger when α\alpha65 is small, pushing low-probability paths down faster and sharpening mass around high-reward trajectories; for α\alpha66, it is negative and promotes exploration. The paper recommends a two-stage schedule in which α\alpha67 is first held away from α\alpha68 and then exponentially annealed back toward α\alpha69:

α\alpha70

The paper explicitly notes that it does not use the term HBG, but in the supplied interpretation HBG corresponds to this α\alpha71-weighted hybridization of forward and backward components (Chen et al., 2 Feb 2026). The empirical effect is substantial: across Set, Bit Sequence, and Molecule Generation, α\alpha72-GFN objectives are reported to outperform previous GFlowNet objectives, with up to a α\alpha73 increase in the number of discovered modes. The conceptual significance is that hybrid balance can be framed not only as combining local and global constraints, but also as changing the reversible mixture that underlies those constraints.

7. Practical guidance, limitations, and recurrent misconceptions

A recurrent misconception is that HBG denotes a single standardized loss. The available literature does not support that claim. One paper explicitly names HBG for VRP and defines it as α\alpha74 plus a depot-centric inference rule (Zhang et al., 6 Oct 2025). Another uses the term for hybrid discrete–continuous TB with explicit Radon–Nikodym corrections (Lahlou et al., 2023). The SubTB(α\alpha75) and α\alpha76-GFN papers both state that they do not explicitly use the HBG name, although they can be interpreted as hybrid-balance constructions (Madan et al., 2022, Chen et al., 2 Feb 2026). This suggests that HBG is an umbrella description for several non-identical balance hybridizations.

A second misconception is that hybridization merely means summing losses. In continuous or mixed spaces, the central issue may instead be measure-theoretic correctness. The continuous-state theory explicitly warns not to replace integrals over states by actions without Radon–Nikodym corrections, and emphasizes that support mismatch or ill-conditioned Jacobians can lead to biased or unstable training (Lahlou et al., 2023). In this setting, HBG should be used when the forward policy is parameterized in an action space with deterministic hybrid maps to states; if the model works directly in state space through α\alpha77 densities with respect to α\alpha78, then standard TB, DB, or FM already include the proper correction and are simpler to implement.

A third misconception is that stronger local balance is always sufficient. The VRP results contradict that view: DB-only is worse than TB and worse than the combined HBG objective, with gaps up to α\alpha79 for AGFN and α\alpha80 for GFACS at α\alpha81 in the cited ablations (Zhang et al., 6 Oct 2025). Conversely, purely global TB can exhibit diffuse credit assignment on long-horizon problems, which is why SubTB(α\alpha82) and the VRP HBG framework both introduce intermediate or additive local structure (Madan et al., 2022, Zhang et al., 6 Oct 2025).

The main practical limitations are likewise heterogeneous. In the generalized theory, correctness relies on finitely absorbing structure, accessibility, absolute continuity of kernels, and existence of the backward reference kernel α\alpha83; trajectory length and Jacobian conditioning directly affect memory and numerical stability (Lahlou et al., 2023). In SubTB(α\alpha84), the α\alpha85 enumeration of subtrajectories adds overhead in linear operations, and choosing α\alpha86 too close to the TB or DB extremes can reintroduce high variance or high bias; the paper reports that fixed α\alpha87 near α\alpha88–α\alpha89 is a strong default and that truncation to short subtrajectories can still work well (Madan et al., 2022). In the VRP formulation, α\alpha90 is the most stable setting in the reported experiments, but benefits still depend on the underlying solver, and depot-centric inference is structurally most advantageous for depot problems such as CVRP rather than TSP (Zhang et al., 6 Oct 2025). In α\alpha91-GFNs, fixed extreme α\alpha92 can reduce reward fitting, so annealing back to α\alpha93 is recommended unless a persistent exploration or exploitation bias is explicitly desired (Chen et al., 2 Feb 2026).

Within these constraints, the unifying principle remains stable across the literature: HBG refers to a GFlowNet construction in which balance is hybridized so that reward-proportional terminal sampling is preserved or approximated while the model gains a more useful training signal, a more general state space, or a more controllable exploration–exploitation profile.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Hybrid-Balance GFlowNet (HBG).