---
title: 'Torsional-GFN: Conditional Conformation Sampling'
url: https://www.emergentmind.com/topics/torsional-gfn
type: topic
---

# Torsional-GFN: Conditional Conformation Sampling

Searching arXiv for the primary paper and closely related work on torsion-focused molecular conformation generation.
Torsional-GFN is a conditional conformation generator for small molecules that uses a GFlowNet over torsion angles to sample conformations approximately proportionally to the Boltzmann distribution, conditioned on a molecular graph and its local structure, namely bond lengths and bond angles [2507.11759]. Its central design choice is to factor molecular conformation into local structure and torsional degrees of freedom, then train a reward-driven generative model whose target density is proportional to \(\exp(-E/k_B T)\). Within this formulation, Torsional-GFN is positioned as a torsion-centric alternative to molecular dynamics and to likelihood-based generative approaches, with an emphasis on reward-only training, off-policy trajectories, and conditional zero-shot transfer to unseen bond lengths and angles for the molecules studied [2507.11759].

## 1. Scope and scientific setting

The target problem is to sample independent conformations \(c \in C_G\) for a molecular graph \(G=(V,E)\) from
\[
p(c \mid G) = \frac{1}{Z(G)} \exp\left(\frac{-E(c)}{k_B T}\right),
\]
where \(E(c)\) is the internal energy and \(Z(G)\) the unknown normalizer [2507.11759]. The paper motivates this distribution by its role in thermodynamic estimates, including free energies and binding affinity calculations in drug discovery.

The method is framed against two broad baselines. Molecular dynamics is described as accurate because it simulates physical time evolution, but expensive because uncorrelated samples require many small time steps. Other generative approaches, including Boltzmann generators and diffusion models, are described as potentially faster but often reliant on likelihood, Kullback-Leibler objectives, or importance reweighting, with associated mode-seeking, mean-seeking, or high-variance behavior [2507.11759]. Torsional-GFN is introduced as a GFlowNet-based alternative that can train directly from reward signals and can use off-policy trajectory data without importance sampling.

A further premise is that the rigid-rotor approximation is insufficient if the full Boltzmann ensemble is sought, because bond lengths and bond angles fluctuate at finite temperature. Accordingly, the method does not fix local structure completely. Instead, it conditions on local structure and restricts the generative model itself to torsions [2507.11759]. This places Torsional-GFN in the class of intrinsic-coordinate conformation models, but with the generative burden concentrated on the torus of rotatable dihedrals.

## 2. Factorization into local structure and torsional variables

The conformation is written in intrinsic coordinates as local structure \(L\), consisting of bond lengths and bond angles, together with torsions \(\Phi=(\phi^1,\dots,\phi^m)\), the \(m\) rotatable dihedral angles [2507.11759]. The paper expresses the factorization as
\[
p(c \mid G) \approx p(L\mid G)\, p(\Phi \mid G, L).
\]
In this decomposition, Torsional-GFN learns only the conditional torsional component,
\[
p_\top^\theta(\Phi \mid G,L),
\]
while \(L\) is supplied externally.

The intrinsic-to-extrinsic conversion is written as
\[
p(c \mid G) = \frac{1}{\sqrt{\det(g)}}\, p(\Phi\mid G,L)\, p(L\mid G),
\]
where \(\det(g)\) is the metric or Jacobian factor for the coordinate transformation [2507.11759]. In the reported experiments, the model sets \(\det(g)=1\) for simplicity, while noting that the factor matters in a fully correct treatment. This detail is significant because it locates the method between a practical empirical implementation and a more exact statistical-mechanical formulation.

The reward is defined by the Boltzmann factor
\[
R(\Phi\mid G,L)=\exp\left(-\frac{E(c(\Phi,L))}{k_B T}\right).
\]
The conditional generative objective is therefore to sample torsion angles with probability proportional to this reward, given the graph and local structure [2507.11759]. In this sense, Torsional-GFN is not merely a conformer proposal model; it is explicitly aimed at reward-proportional sampling on torsional space.

## 3. GFlowNet construction on the torsion torus

The state space is the \(m\)-dimensional torus
\[
\mathcal{X}=[0,2\pi]^m,
\]
with one angle per rotatable bond [2507.11759]. A trajectory is a sequence of torsion states
\[
\tau = (\Phi_0 \rightarrow \Phi_1 \rightarrow \cdots \rightarrow \Phi_n=\Phi),
\]
starting from a source state \(\Phi_0\) and progressing through sequential updates. The model uses a forward policy \(P_F(\Phi_t \mid \Phi_{t-1})\) and a backward policy \(P_B(\Phi_{t-1} \mid \Phi_t)\), both parameterized as mixtures of von Mises distributions, which the paper treats as a natural parameterization for circular variables.

The training objective is the VarGrad GFlowNet objective. For a batch of trajectories \(B_\tau\), the loss is
\[
\mathcal{L}_{\mathrm{VG}}(B_\tau;\theta \mid G,L) = \mathbb{E}_{\tau\in B_\tau} \left[ \log Z_\theta(G,L) - \log \mathcal{C}_\theta(\tau\mid G,L) \right]^2,
\]
with
\[
\mathcal{C}_\theta(\tau\mid G,L) = \frac{P_B^\theta(\tau\mid \Phi_n,G,L)\,R(\Phi_n\mid G,L)} {P_F^\theta(\tau\mid G,L)},
\]
and
\[
\log Z_\theta(G,L) = \mathbb{E}_{\tau'\in B_\tau} \log \mathcal{C}_\theta(\tau'\mid G,L).
\]
The full dataset loss is
\[
\mathcal{L}_D = \mathbb{E}_{G_i,L_i\sim D}\, \mathcal{L}_{\mathrm{VG}}(B_\tau;\theta\mid G_i,L_i)
\]
[2507.11759].

The stated interpretation is that the objective drives the trajectory-wise quantity \(\mathcal{C}_\theta\) toward a constant \(Z_\theta\), which in turn implies that terminal torsion states are sampled in proportion to the reward [2507.11759]. The paper emphasizes an operational advantage over importance-sampling-based energy training: the behavior policy may be any full-support off-policy distribution, and no importance weighting is required. This distinguishes the method from energy-based training schemes whose gradients depend directly on importance ratios.

## 4. Policy parameterization and VectorGNN

To support a single model across multiple molecules, the paper introduces VectorGNN, a graph neural network designed to be reflection-equivariant and to infer torsion-relevant geometry from the molecular graph and 3D coordinates [2507.11759]. This is the principal architectural component used to amortize the forward and backward torsional policies across different conditional inputs.

The forward and backward policies are written as mixtures of von Mises distributions:
\[
P^\theta_F(\phi_{t+1}\mid \phi_t) = \sum_{k=1}^{K} w^\theta_{k,F}\, \mathrm{VM}\!\left(\phi_{t+1}\mid \mu^\theta_{k,F}(\phi_t),\kappa^\theta_{k,F}(\phi_t)\right),
\]
\[
P^\theta_B(\phi_t\mid \phi_{t+1}) = \sum_{k=1}^{K} w^\theta_{k,B}\, \mathrm{VM}\!\left(\phi_t\mid \mu^\theta_{k,B}(\phi_{t+1}),\kappa^\theta_{k,B}(\phi_{t+1})\right).
\]
VectorGNN first performs invariant message passing to obtain atom embeddings, then predicts pairwise force-like quantities, converts them into torques around rotatable bonds, and outputs pseudo-scalars used to parameterize the von Mises mixture:
\[
w_k(\phi) = (o_\phi)^2_{k,1}, \quad \mu_k(\phi) = (o_\phi)_{k,2}, \quad \kappa_k(\phi) = (o_\phi)^2_{k,3}.
\]
The paper presents this architecture as more suitable than a multilayer perceptron for structure sharing across molecules [2507.11759].

At inference time, the model is conditioned on a molecular graph \(G\) and an externally obtained local structure \(L\). It initializes from a source torsion state \(\Phi_0\) sampled uniformly on \([0,2\pi]^m\), repeatedly samples updates from the learned forward policy, and after \(n\) steps constructs the full 3D conformation \(c(\Phi,L)\) from the fixed local geometry and the sampled torsions [2507.11759]. The paper notes that the arbitrary reference atoms used to define each torsion only translate the reward landscape on the torus by a multiple of \(2\pi\), which it treats as a robustness property of the formulation.

## 5. Training protocol and empirical findings

The reported experiments use six molecules from FreeSolv for training, each with two rotatable torsions, and two additional molecules for testing [2507.11759]. Ground-truth conformations are obtained from molecular dynamics simulations using OpenMM and OpenFF 2.1.1, run for 2 ns at 1 fs and subsampled to 1 ps decorrelated frames, yielding 2001 conformations per molecule. For training, bond lengths and bond angles are fixed to those from one arbitrary molecular dynamics conformation per molecule, and the energy function is MMFF94s. Before GFlowNet training, VectorGNN is pretrained on supervised prediction of energy, \(\sin\phi\), and \(\cos\phi\) using 10,000 uniformly sampled torsion configurations per molecule [2507.11759].

Evaluation uses three metrics. The first is the Jensen-Shannon divergence between the discretized learned terminal distribution and the discretized target distribution,
\[
\mathrm{JSD}^P = \mathrm{JSD}(P_\top^\theta(\Phi_i\mid G,L)\,\|\,P(\Phi_i\mid G,L)).
\]
The second is the log-probability versus log-reward correlation,
\[
\rho_{\log p_\top^\theta,\log R}
=
\frac{\mathrm{cov}(\log p_\top^\theta(\Phi_i\mid G,L),\log R(\Phi_i\mid G,L))}
{\sigma_{\log p_\top^\theta}\sigma_{\log R}}.
\]
The third is the energy histogram Jensen-Shannon divergence ratio,
\[
\mathrm{JSD}^E_{\mathrm{GFN}/\mathrm{JSD}^E_{\mathrm{rand}}}
=
\frac{\mathrm{JSD}(P_{\mathrm{MD}(E)}\|P_{\mathrm{GFN}(E)})}
{\mathrm{JSD}(P_{\mathrm{MD}(E)}\|P_{\mathrm{rand}(E)})},
\]
for which values below 1 indicate that Torsional-GFN is closer to molecular dynamics than random torsions [2507.11759].

On the training molecules, the paper reports strong agreement with the target landscape. Representative results include **CCC**, with \(\mathrm{JSD}^P = 0.0100\), \(\rho = 0.9740\), and energy ratio \(0.0433\); **C[C@@H]1CCCC[C@@H]1C**, with \(\mathrm{JSD}^P = 0.0270\), \(\rho = 0.9561\), and energy ratio \(0.1391\); and **COC=O**, with \(\mathrm{JSD}^P = 0.0358\), \(\rho = 0.7068\), and energy ratio \(0.0201\) [2507.11759]. The paper further states that, for several training molecules, the energy histograms of Torsional-GFN samples nearly overlap those of molecular dynamics, and the sampled probability landscapes visually track the molecular-dynamics reward landscape.

The results on held-out molecules are described more cautiously. One test molecule shows poor correlation and incomplete mode coverage, whereas the other shows partial generalization but not a perfect match [2507.11759]. The empirical picture is therefore asymmetric: strong training-set performance, clear evidence of conditional transfer to unseen local structures for seen molecules, and only limited evidence for generalization to unseen molecules.

## 6. Generalization claims, limitations, and prospective extensions

A principal claim of the paper is zero-shot generalization to unseen local structures. The model is trained with fixed bond lengths and angles per molecule, then evaluated on new bond lengths and angles sampled from molecular dynamics for the same molecules [2507.11759]. The reported qualitative two-dimensional plots show that the model can follow shifts in the energy landscape induced by these new local structures. This suggests that the learned torsion policy is not tied only to a single frozen geometry, but responds to conditioned local geometric changes.

The same section of the paper is explicit that unseen-molecule generalization is not yet established. The evidence is limited to two held-out molecules, one of which behaves poorly and one of which is only somewhat closer to molecular dynamics than random torsions [2507.11759]. The authors accordingly describe this as potential rather than as a resolved capability.

The reported limitations are substantial and define the current scope of Torsional-GFN. Training is described as computationally expensive, with even the small experimental setting requiring a GPU with at least 40 GB memory. The experiments are restricted to molecules with only two rotatable torsions. Local structure is not generated by the GFlowNet; it is sampled from molecular dynamics and provided as conditioning input. In addition, the implementation sets \(\det(g)=1\) rather than fully accounting for the intrinsic-to-extrinsic volume correction [2507.11759].

The paper identifies several future directions: training on larger datasets and larger molecular systems, extending the GFlowNet to generate local structures \(L\) as well as torsions, exploring alternative policy parameterizations, and possibly using simulation-free objectives [2507.11759]. A plausible implication is that the present method should be understood as a proof of concept for conditional Boltzmann-targeted torsional sampling rather than as a complete solution to unconditional conformer generation.

## 7. Relation to torsion-centric generative modeling

Torsional-GFN belongs to a broader line of work that concentrates conformation generation on flexible torsional coordinates rather than on full Cartesian space. The most immediate comparison in the supplied literature is “Torsional Diffusion for Molecular Conformer Generation,” which likewise focuses on torsion angles but uses a diffusion process on the hypertorus \(\mathbb{T}^m\) with an extrinsic-to-intrinsic score model rather than a GFlowNet objective [2206.01729].

The distinction is methodological and objective-level. Torsional diffusion learns a continuous-time score model and samples conformers by reverse diffusion; it is explicitly described as not building trajectories by sequentially adding actions to maximize a flow objective [2206.01729]. Torsional-GFN, by contrast, is a conditional GFlowNet whose forward and backward policies define torsional trajectories, and whose terminal distribution is trained to be proportional to a reward defined by the Boltzmann factor [2507.11759]. Both approaches concentrate the generative process on flexible dihedral variables, but only Torsional-GFN is cast as reward-proportional sampling via GFlowNet balance conditions.

This comparison also clarifies a common source of ambiguity in the term “torsional” within molecular machine learning. In this context it refers neither to mechanical torsion in materials nor to generic torsional dynamics in continuum mechanics. It denotes a torsion-space formulation of molecular conformation generation, where the primary latent or state variables are rotatable dihedral angles [2206.01729]. Within that category, Torsional-GFN is specifically a conditional continuous GFlowNet on the torus of torsion angles, trained from reward alone and conditioned on local molecular geometry [2507.11759].

Source: https://www.emergentmind.com/topics/torsional-gfn