Papers
Topics
Authors
Recent
Search
2000 character limit reached

Q-Guided Value-Gradient Matching (Q-VGM)

Updated 15 July 2026
  • Q-VGM is an off-policy reinforcement learning method that transforms critic gradients into denoising-time velocity corrections to fine-tune VLA policies.
  • It overcomes limitations of policy-gradient and direct Q-maximization by aligning the policy’s residual velocity field with value gradients without backpropagating through the denoising chain.
  • Empirical results on LIBERO, RoboTwin 2.0, and real-robot tasks show significant improvements, validating Q-VGM's effectiveness over existing approaches.

Searching arXiv for the named method and closely related work to ground the article in current papers. {"query":"Q-Guided Value-Gradient Matching for Flow-Matching VLA Policies arXiv (Wang et al., 6 Jun 2026)", "max_results": 5} Q-Guided Value-Gradient Matching (Q-VGM) is an off-policy reinforcement learning method for fine-tuning flow-matching vision-language-action (VLA) policies against a learned Q-function. Its defining move is to transform end-point action gradients from a critic into a denoising-time value-gradient field, and then to align the policy’s residual velocity field to that target without backpropagating through the denoising chain and without requiring an action likelihood. In the formulation introduced for a few-shot-supervised π0.5\pi_{0.5} VLA, Q-VGM operates on a fixed replay buffer, uses self-generated rollout data rather than additional expert supervision, and reports average success-rate improvements from 75.0%75.0\% to 92.5%92.5\% on LIBERO, from 76.4%76.4\% to 87.2%87.2\% on RoboTwin 2.0, and from 40.0%40.0\% to 67.5%67.5\% on two real-robot tabletop tasks (Wang et al., 6 Jun 2026).

1. Methodological context and problem setting

Q-VGM addresses a specific difficulty in applying RL to flow-matching VLA policies: the policy is an iterative denoising model over action chunks, while the critic evaluates clean, executable chunks at t=0t=0. In the underlying setup, the base policy is a flow-matching VLA model, specifically π0.5\pi_{0.5} from Physical Intelligence. At each decision step kk, given observation 75.0%75.0\%0, language instruction 75.0%75.0\%1, and proprioception 75.0%75.0\%2, it outputs an action chunk

75.0%75.0\%3

with a frozen vision-language backbone and a fine-tuned flow-based action expert (Wang et al., 6 Jun 2026).

The paper isolates two standard RL routes as problematic for such policies. Policy-gradient-style RL requires a tractable action likelihood 75.0%75.0\%4 or a usable surrogate, but flow policies generate actions by iterative denoising and do not expose a simple likelihood. Naïve value-based methods are also unsatisfactory: direct Q-maximization through the denoising chain yields extremely long gradient paths and numerical instability at VLA scale; test-time Q selection or guidance does not amortize improvement into policy parameters; and Q-guided action distillation supervises only the terminal clean action while ignoring the intermediate velocity field that the flow actually learns (Wang et al., 6 Jun 2026).

This positioning makes Q-VGM distinct from nearby approaches. “GoldenStart: Q-Guided Priors and Entropy Control for Distilling Flow Policies” shapes the initial noise distribution and the distillation objective with Q-guided mechanisms, but it does not implement explicit value-gradient matching in ODE space (Zhang et al., 15 Mar 2026). “Guided Action Flow” uses a learned critic only at inference time to guide a frozen SmolVLA reverse-time sampler through action gradients, again without updating the base policy (Yang et al., 2 Jul 2026). By contrast, Q-VGM is a training-time method that converts critic information into supervision for the denoising-time velocity field itself (Wang et al., 6 Jun 2026).

A recurrent misconception is that Q-VGM is merely a variant of Q-improved action distillation. The method is explicitly framed otherwise: actions improved by local 75.0%75.0\%5-ascent are not the sole training targets; instead, the corresponding displacement is converted into a time-dependent velocity correction, and the policy’s residual velocity is trained to match that correction across denoising steps (Wang et al., 6 Jun 2026).

2. Flow formulation and value-gradient view

The action expert is a conditional flow model with time convention 75.0%75.0\%6 for clean actions and 75.0%75.0\%7 for Gaussian noise. For a clean chunk 75.0%75.0\%8 and noise 75.0%75.0\%9, the interpolation path is

92.5%92.5\%0

with 92.5%92.5\%1 and 92.5%92.5\%2. The expert predicts a velocity field

92.5%92.5\%3

where 92.5%92.5\%4 is the conditioning context composed of VLM prefix features and proprioception. Sampling uses an Euler discretization over a schedule 92.5%92.5\%5: 92.5%92.5\%6 yielding 92.5%92.5\%7 (Wang et al., 6 Jun 2026).

Q-VGM adopts the value-gradient interpretation supplied by “Value Gradient Guidance for Flow Matching Alignment,” which formulates flow fine-tuning as a KL-regularized stochastic optimal control problem. In that view, the updated flow is decomposed into a pretrained base field and a residual control field: 92.5%92.5\%8 With terminal reward 92.5%92.5\%9, the KL-regularized improved distribution is written as

76.4%76.4\%0

and the corresponding control objective is

76.4%76.4\%1

VGG-Flow shows that the optimal residual velocity takes the value-gradient form

76.4%76.4\%2

with terminal boundary condition

76.4%76.4\%3

Q-VGM instantiates this by setting

76.4%76.4\%4

so that the terminal value gradient is induced by 76.4%76.4\%5 (Liu et al., 4 Dec 2025).

The result is a precise flow-control reinterpretation of actor improvement. Rather than differentiating the critic through the full denoising chain, Q-VGM approximates the denoising-time value-gradient field constructively from terminal Q-gradients and local Euler geometry, then trains the flow’s residual velocity to match it (Wang et al., 6 Jun 2026).

3. Construction of the denoising-time value-gradient field

The central construction in Q-VGM is the effective denoising-time correction field 76.4%76.4\%6. For each denoising step 76.4%76.4\%7, the method first defines a look-forward clean-action estimate under the frozen base flow: 76.4%76.4\%8 This is a one-step Euler projection of the noisy state 76.4%76.4\%9 to an estimated clean action at 87.2%87.2\%0, computed under the base policy and treated with stop-gradient (Wang et al., 6 Jun 2026).

Starting from that estimate, Q-VGM performs iterative Q-gradient ascent in action space: 87.2%87.2\%1

87.2%87.2\%2

Among the sequence 87.2%87.2\%3, it then applies keep-best selection: 87.2%87.2\%4 Because 87.2%87.2\%5 is the unmodified base action, this mechanism effectively turns off guidance when local ascent is not beneficial (Wang et al., 6 Jun 2026).

The selected action displacement is translated into a velocity correction by dividing by the remaining denoising time: 87.2%87.2\%6 The policy is then trained with a gated local regression objective,

87.2%87.2\%7

where the denoising-time gate is 87.2%87.2\%8. The gate emphasizes times closer to clean actions, where the one-step look-forward is more accurate (Wang et al., 6 Jun 2026).

Several implementation choices are integral to the stability claim. Gradients do not propagate through the denoising chain: 87.2%87.2\%9, 40.0%40.0\%0, the critic, and 40.0%40.0\%1 are all treated with stop-gradient. The denoising state is advanced as

40.0%40.0\%2

so the policy update remains strictly local in time. This is the operative sense in which Q-VGM avoids both action likelihoods and gradient-through-time, while still using first-order critic information 40.0%40.0\%3 (Wang et al., 6 Jun 2026).

4. Critic design and action-sensitive representation

The critic in Q-VGM is a chunk-level Q-function trained entirely offline on a fixed replay buffer. Each transition has the form

40.0%40.0\%4

with chunk reward

40.0%40.0\%5

Each critic head 40.0%40.0\%6 is trained with a Cal-QL objective

40.0%40.0\%7

where

40.0%40.0\%8

and

40.0%40.0\%9

After training, the ensemble mean

67.5%67.5\%0

is frozen and used for action gradients 67.5%67.5\%1 (Wang et al., 6 Jun 2026).

State representation is compressed through compact RLT features rather than raw VLA prefix tokens. The VLM prefix is mapped to a RL Token vector 67.5%67.5\%2 via a small transformer autoencoder; proprioception 67.5%67.5\%3 is projected by a linear map 67.5%67.5\%4; and the critic state is

67.5%67.5\%5

The VLM and the RLT encoder are frozen. This design keeps critic inputs compact while retaining VLA-derived state information (Wang et al., 6 Jun 2026).

A second design feature is per-layer action injection. Flattened action chunks are concatenated not only at the input but at every hidden layer: 67.5%67.5\%6

67.5%67.5\%7

67.5%67.5\%8

The paper’s stated purpose is to preserve local sensitivity of Q to small perturbations in 67.5%67.5\%9, which is essential because the actor update uses t=0t=00 rather than scalar values alone. Ablations report that removing per-layer action injection degrades LIBERO performance from t=0t=01 to t=0t=02, while using a single critic head yields t=0t=03 (Wang et al., 6 Jun 2026).

This architecture places Q-VGM in a broader lineage of value-gradient methods. Classical value-gradient learning treats gradients as the core control signal rather than scalar values alone (0803.3539). More recent “Continuous Q-Score Matching” similarly links a diffusion-policy score to the action gradient of a learned continuous Q-function by dynamic programming, using

t=0t=04

as the target score field in continuous time (Hua et al., 20 Oct 2025). Q-VGM differs in formulation and domain, but the shared emphasis on first-order critic structure is explicit.

5. Empirical performance and comparative behavior

Q-VGM is evaluated on three settings: LIBERO, RoboTwin 2.0, and two real-robot tabletop tasks. In all cases, it starts from a few-shot supervised t=0t=05 policy, uses self-generated rollout data to train the critic, and then performs purely offline Q-VGM fine-tuning of the action expert (Wang et al., 6 Jun 2026).

Setting Baseline Q-VGM
LIBERO average success 75.0% 92.5%
RoboTwin 2.0 average success 76.4% 87.2%
Real-robot average success 40.0% 67.5%

On LIBERO, the suite-level results are Spatial t=0t=06, Object t=0t=07, Goal t=0t=08, and Long t=0t=09, with the largest gain on the longest-horizon suite. On the same backbone and same critic, the reported baselines are Test-time Q Selection at π0.5\pi_{0.5}0, Test-time Q Guidance at π0.5\pi_{0.5}1, Q-Improved Action Distillation at π0.5\pi_{0.5}2, and Diffusion-QL at π0.5\pi_{0.5}3, which is worse than the SFT baseline (Wang et al., 6 Jun 2026).

On RoboTwin 2.0, Q-VGM improves the average success rate from π0.5\pi_{0.5}4 to π0.5\pi_{0.5}5, compared with π0.5\pi_{0.5}6 for Test-time Q Selection, π0.5\pi_{0.5}7 for Test-time Q Guidance, and π0.5\pi_{0.5}8 for Q-Distillation. Large task-level gains are reported for place_shoe π0.5\pi_{0.5}9, handover_mic kk0, and lift_pot kk1 (Wang et al., 6 Jun 2026).

On real robots, Pick Peach improves from kk2 kk3 to kk4 kk5, and Stack Bowls improves from kk6 kk7 to kk8 kk9. The average is therefore 75.0%75.0\%00 (Wang et al., 6 Jun 2026).

The ablation pattern clarifies what the method is using the critic for. On LIBERO, replacing RLT with ResNet drops performance to 75.0%75.0\%01; removing per-layer action injection gives 75.0%75.0\%02; a single critic head yields 75.0%75.0\%03; removing keep-best produces 75.0%75.0\%04; using uniform guidance 75.0%75.0\%05 gives 75.0%75.0\%06; and removing the frozen base anchor for look-forward yields 75.0%75.0\%07 (Wang et al., 6 Jun 2026). These measurements support a narrow conclusion: the empirical advantage is associated not only with having a critic, but with turning its action gradients into a structured denoising-time target field and keeping that target anchored to the behavior policy’s support.

6. Relation to adjacent Q-guided flow methods and open questions

Q-VGM sits at the intersection of several strands of work on value-guided generative control. VGG-Flow provides the most direct theoretical precursor: it derives the optimal residual between a finetuned and pretrained velocity field as a scaled negative value gradient,

75.0%75.0\%08

and trains flow models by matching the residual velocity field to a learned value-gradient oracle (Liu et al., 4 Dec 2025). Q-VGM inherits this velocity-matching perspective, but replaces a terminal reward model with a learned action-value critic over VLA action chunks (Wang et al., 6 Jun 2026).

“Guided Action Flow” shows a different use of the same underlying signal. There, a frozen SmolVLA policy is guided only at inference time by a learned action-chunk critic through

75.0%75.0\%09

which is injected into the reverse-time sampler via a guided velocity update. The method improves validation success on LIBERO but identifies critic generalization and uncertainty-aware guidance as central bottlenecks (Yang et al., 2 Jul 2026). Q-VGM can be read as amortizing this kind of test-time guidance into policy parameters, rather than paying its cost during inference. This suggests a conceptual division: Guided Action Flow is Q-guided sampling for frozen policies, whereas Q-VGM is Q-guided velocity-field learning for trainable policies.

GoldenStart provides another neighboring template. It introduces a Q-guided prior over the initial noise for distilled flow policies and combines it with entropy-regularized actor learning. The paper explicitly states that it does not implement literal value-gradient matching in ODE space, but it can be interpreted as an implicit form of Q-guided value alignment in latent and action spaces (Zhang et al., 15 Mar 2026). Relative to that formulation, Q-VGM is more literal about matching a denoising-time correction field rather than only shaping latent startpoints or terminal action objectives.

The method also leaves open questions stated directly in the paper. It relies on critic quality and on reasonably accurate Q-gradients near the behavior distribution; outside this region, 75.0%75.0\%10 can be misleading, with mitigation provided only by gradient clipping, keep-best fallback, and base-anchored look-forward. The authors identify more principled trust-region control, citing TR-QAM as a possible future direction. They also note critic scalability and horizon length as unresolved issues, and suggest that combining world models with Q-functions could shorten effective horizons and provide better gradients. Finally, although Q-VGM is motivated by the VGG-Flow stochastic-optimal-control formulation, it does not solve the exact adjoint PDE; formal convergence or error bounds for its approximate value-gradient propagation remain unexplored (Wang et al., 6 Jun 2026).

A plausible implication is that Q-VGM occupies an intermediate point between exact value-gradient control theory and purely heuristic Q-guided sampling. It is more structured than terminal action distillation or inference-time rescoring, but less explicit than an exact HJB or adjoint solution. That intermediate position appears to be the source of both its practical effectiveness and its remaining theoretical gap.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Q-Guided Value-Gradient Matching (Q-VGM).