Q-Guided Value-Gradient Matching (Q-VGM)
- Q-VGM is an off-policy reinforcement learning method that transforms critic gradients into denoising-time velocity corrections to fine-tune VLA policies.
- It overcomes limitations of policy-gradient and direct Q-maximization by aligning the policy’s residual velocity field with value gradients without backpropagating through the denoising chain.
- Empirical results on LIBERO, RoboTwin 2.0, and real-robot tasks show significant improvements, validating Q-VGM's effectiveness over existing approaches.
Searching arXiv for the named method and closely related work to ground the article in current papers. {"query":"Q-Guided Value-Gradient Matching for Flow-Matching VLA Policies arXiv (Wang et al., 6 Jun 2026)", "max_results": 5} Q-Guided Value-Gradient Matching (Q-VGM) is an off-policy reinforcement learning method for fine-tuning flow-matching vision-language-action (VLA) policies against a learned Q-function. Its defining move is to transform end-point action gradients from a critic into a denoising-time value-gradient field, and then to align the policy’s residual velocity field to that target without backpropagating through the denoising chain and without requiring an action likelihood. In the formulation introduced for a few-shot-supervised VLA, Q-VGM operates on a fixed replay buffer, uses self-generated rollout data rather than additional expert supervision, and reports average success-rate improvements from to on LIBERO, from to on RoboTwin 2.0, and from to on two real-robot tabletop tasks (Wang et al., 6 Jun 2026).
1. Methodological context and problem setting
Q-VGM addresses a specific difficulty in applying RL to flow-matching VLA policies: the policy is an iterative denoising model over action chunks, while the critic evaluates clean, executable chunks at . In the underlying setup, the base policy is a flow-matching VLA model, specifically from Physical Intelligence. At each decision step , given observation 0, language instruction 1, and proprioception 2, it outputs an action chunk
3
with a frozen vision-language backbone and a fine-tuned flow-based action expert (Wang et al., 6 Jun 2026).
The paper isolates two standard RL routes as problematic for such policies. Policy-gradient-style RL requires a tractable action likelihood 4 or a usable surrogate, but flow policies generate actions by iterative denoising and do not expose a simple likelihood. Naïve value-based methods are also unsatisfactory: direct Q-maximization through the denoising chain yields extremely long gradient paths and numerical instability at VLA scale; test-time Q selection or guidance does not amortize improvement into policy parameters; and Q-guided action distillation supervises only the terminal clean action while ignoring the intermediate velocity field that the flow actually learns (Wang et al., 6 Jun 2026).
This positioning makes Q-VGM distinct from nearby approaches. “GoldenStart: Q-Guided Priors and Entropy Control for Distilling Flow Policies” shapes the initial noise distribution and the distillation objective with Q-guided mechanisms, but it does not implement explicit value-gradient matching in ODE space (Zhang et al., 15 Mar 2026). “Guided Action Flow” uses a learned critic only at inference time to guide a frozen SmolVLA reverse-time sampler through action gradients, again without updating the base policy (Yang et al., 2 Jul 2026). By contrast, Q-VGM is a training-time method that converts critic information into supervision for the denoising-time velocity field itself (Wang et al., 6 Jun 2026).
A recurrent misconception is that Q-VGM is merely a variant of Q-improved action distillation. The method is explicitly framed otherwise: actions improved by local 5-ascent are not the sole training targets; instead, the corresponding displacement is converted into a time-dependent velocity correction, and the policy’s residual velocity is trained to match that correction across denoising steps (Wang et al., 6 Jun 2026).
2. Flow formulation and value-gradient view
The action expert is a conditional flow model with time convention 6 for clean actions and 7 for Gaussian noise. For a clean chunk 8 and noise 9, the interpolation path is
0
with 1 and 2. The expert predicts a velocity field
3
where 4 is the conditioning context composed of VLM prefix features and proprioception. Sampling uses an Euler discretization over a schedule 5: 6 yielding 7 (Wang et al., 6 Jun 2026).
Q-VGM adopts the value-gradient interpretation supplied by “Value Gradient Guidance for Flow Matching Alignment,” which formulates flow fine-tuning as a KL-regularized stochastic optimal control problem. In that view, the updated flow is decomposed into a pretrained base field and a residual control field: 8 With terminal reward 9, the KL-regularized improved distribution is written as
0
and the corresponding control objective is
1
VGG-Flow shows that the optimal residual velocity takes the value-gradient form
2
with terminal boundary condition
3
Q-VGM instantiates this by setting
4
so that the terminal value gradient is induced by 5 (Liu et al., 4 Dec 2025).
The result is a precise flow-control reinterpretation of actor improvement. Rather than differentiating the critic through the full denoising chain, Q-VGM approximates the denoising-time value-gradient field constructively from terminal Q-gradients and local Euler geometry, then trains the flow’s residual velocity to match it (Wang et al., 6 Jun 2026).
3. Construction of the denoising-time value-gradient field
The central construction in Q-VGM is the effective denoising-time correction field 6. For each denoising step 7, the method first defines a look-forward clean-action estimate under the frozen base flow: 8 This is a one-step Euler projection of the noisy state 9 to an estimated clean action at 0, computed under the base policy and treated with stop-gradient (Wang et al., 6 Jun 2026).
Starting from that estimate, Q-VGM performs iterative Q-gradient ascent in action space: 1
2
Among the sequence 3, it then applies keep-best selection: 4 Because 5 is the unmodified base action, this mechanism effectively turns off guidance when local ascent is not beneficial (Wang et al., 6 Jun 2026).
The selected action displacement is translated into a velocity correction by dividing by the remaining denoising time: 6 The policy is then trained with a gated local regression objective,
7
where the denoising-time gate is 8. The gate emphasizes times closer to clean actions, where the one-step look-forward is more accurate (Wang et al., 6 Jun 2026).
Several implementation choices are integral to the stability claim. Gradients do not propagate through the denoising chain: 9, 0, the critic, and 1 are all treated with stop-gradient. The denoising state is advanced as
2
so the policy update remains strictly local in time. This is the operative sense in which Q-VGM avoids both action likelihoods and gradient-through-time, while still using first-order critic information 3 (Wang et al., 6 Jun 2026).
4. Critic design and action-sensitive representation
The critic in Q-VGM is a chunk-level Q-function trained entirely offline on a fixed replay buffer. Each transition has the form
4
with chunk reward
5
Each critic head 6 is trained with a Cal-QL objective
7
where
8
and
9
After training, the ensemble mean
0
is frozen and used for action gradients 1 (Wang et al., 6 Jun 2026).
State representation is compressed through compact RLT features rather than raw VLA prefix tokens. The VLM prefix is mapped to a RL Token vector 2 via a small transformer autoencoder; proprioception 3 is projected by a linear map 4; and the critic state is
5
The VLM and the RLT encoder are frozen. This design keeps critic inputs compact while retaining VLA-derived state information (Wang et al., 6 Jun 2026).
A second design feature is per-layer action injection. Flattened action chunks are concatenated not only at the input but at every hidden layer: 6
7
8
The paper’s stated purpose is to preserve local sensitivity of Q to small perturbations in 9, which is essential because the actor update uses 0 rather than scalar values alone. Ablations report that removing per-layer action injection degrades LIBERO performance from 1 to 2, while using a single critic head yields 3 (Wang et al., 6 Jun 2026).
This architecture places Q-VGM in a broader lineage of value-gradient methods. Classical value-gradient learning treats gradients as the core control signal rather than scalar values alone (0803.3539). More recent “Continuous Q-Score Matching” similarly links a diffusion-policy score to the action gradient of a learned continuous Q-function by dynamic programming, using
4
as the target score field in continuous time (Hua et al., 20 Oct 2025). Q-VGM differs in formulation and domain, but the shared emphasis on first-order critic structure is explicit.
5. Empirical performance and comparative behavior
Q-VGM is evaluated on three settings: LIBERO, RoboTwin 2.0, and two real-robot tabletop tasks. In all cases, it starts from a few-shot supervised 5 policy, uses self-generated rollout data to train the critic, and then performs purely offline Q-VGM fine-tuning of the action expert (Wang et al., 6 Jun 2026).
| Setting | Baseline | Q-VGM |
|---|---|---|
| LIBERO average success | 75.0% | 92.5% |
| RoboTwin 2.0 average success | 76.4% | 87.2% |
| Real-robot average success | 40.0% | 67.5% |
On LIBERO, the suite-level results are Spatial 6, Object 7, Goal 8, and Long 9, with the largest gain on the longest-horizon suite. On the same backbone and same critic, the reported baselines are Test-time Q Selection at 0, Test-time Q Guidance at 1, Q-Improved Action Distillation at 2, and Diffusion-QL at 3, which is worse than the SFT baseline (Wang et al., 6 Jun 2026).
On RoboTwin 2.0, Q-VGM improves the average success rate from 4 to 5, compared with 6 for Test-time Q Selection, 7 for Test-time Q Guidance, and 8 for Q-Distillation. Large task-level gains are reported for place_shoe 9, handover_mic 0, and lift_pot 1 (Wang et al., 6 Jun 2026).
On real robots, Pick Peach improves from 2 3 to 4 5, and Stack Bowls improves from 6 7 to 8 9. The average is therefore 00 (Wang et al., 6 Jun 2026).
The ablation pattern clarifies what the method is using the critic for. On LIBERO, replacing RLT with ResNet drops performance to 01; removing per-layer action injection gives 02; a single critic head yields 03; removing keep-best produces 04; using uniform guidance 05 gives 06; and removing the frozen base anchor for look-forward yields 07 (Wang et al., 6 Jun 2026). These measurements support a narrow conclusion: the empirical advantage is associated not only with having a critic, but with turning its action gradients into a structured denoising-time target field and keeping that target anchored to the behavior policy’s support.
6. Relation to adjacent Q-guided flow methods and open questions
Q-VGM sits at the intersection of several strands of work on value-guided generative control. VGG-Flow provides the most direct theoretical precursor: it derives the optimal residual between a finetuned and pretrained velocity field as a scaled negative value gradient,
08
and trains flow models by matching the residual velocity field to a learned value-gradient oracle (Liu et al., 4 Dec 2025). Q-VGM inherits this velocity-matching perspective, but replaces a terminal reward model with a learned action-value critic over VLA action chunks (Wang et al., 6 Jun 2026).
“Guided Action Flow” shows a different use of the same underlying signal. There, a frozen SmolVLA policy is guided only at inference time by a learned action-chunk critic through
09
which is injected into the reverse-time sampler via a guided velocity update. The method improves validation success on LIBERO but identifies critic generalization and uncertainty-aware guidance as central bottlenecks (Yang et al., 2 Jul 2026). Q-VGM can be read as amortizing this kind of test-time guidance into policy parameters, rather than paying its cost during inference. This suggests a conceptual division: Guided Action Flow is Q-guided sampling for frozen policies, whereas Q-VGM is Q-guided velocity-field learning for trainable policies.
GoldenStart provides another neighboring template. It introduces a Q-guided prior over the initial noise for distilled flow policies and combines it with entropy-regularized actor learning. The paper explicitly states that it does not implement literal value-gradient matching in ODE space, but it can be interpreted as an implicit form of Q-guided value alignment in latent and action spaces (Zhang et al., 15 Mar 2026). Relative to that formulation, Q-VGM is more literal about matching a denoising-time correction field rather than only shaping latent startpoints or terminal action objectives.
The method also leaves open questions stated directly in the paper. It relies on critic quality and on reasonably accurate Q-gradients near the behavior distribution; outside this region, 10 can be misleading, with mitigation provided only by gradient clipping, keep-best fallback, and base-anchored look-forward. The authors identify more principled trust-region control, citing TR-QAM as a possible future direction. They also note critic scalability and horizon length as unresolved issues, and suggest that combining world models with Q-functions could shorten effective horizons and provide better gradients. Finally, although Q-VGM is motivated by the VGG-Flow stochastic-optimal-control formulation, it does not solve the exact adjoint PDE; formal convergence or error bounds for its approximate value-gradient propagation remain unexplored (Wang et al., 6 Jun 2026).
A plausible implication is that Q-VGM occupies an intermediate point between exact value-gradient control theory and purely heuristic Q-guided sampling. It is more structured than terminal action distillation or inference-time rescoring, but less explicit than an exact HJB or adjoint solution. That intermediate position appears to be the source of both its practical effectiveness and its remaining theoretical gap.