- The paper introduces a Mixture-of-Transformers VLA that integrates reasoning, 3D spatial, and action experts with block-wise causal attention for grounded robot control.
- The paper presents Flow-GSPO, a sequence-level reinforcement learning method for stochastic flow-matching actions, achieving 80.3% on LIBERO-Plus versus 78.7% with PPO and 65.7% with GRPO.
- The paper reports 97.6% average success on LIBERO, while removing the Spatial Expert reduces LIBERO-Plus performance by 8.3 points, highlighting the value of integrated geometric grounding despite simulation-only validation.
OmniVLA-RL is a vision-language-action (VLA) framework that combines two contributions: a Mixture-of-Transformers (MoT) architecture integrating spatial, reasoning, and action experts within shared Transformer layers, and an online reinforcement learning algorithm, Flow-GSPO, that adapts Group Sequence Policy Optimization (GSPO) to flow-matching action generation. The authors report state-of-the-art results on LIBERO, with an average success rate of 97.6%, and substantial improvements over PPO- and GRPO-based fine-tuning on the more challenging LIBERO-Plus benchmark.
Motivation and positioning
The paper identifies two shortcomings in existing VLA systems. First, spatial perception is typically injected either at the encoder (early fusion, e.g., Evo-0, SpatialVLA) or at the action head (late fusion, e.g., FALCON), leaving the core large model untouched; the authors argue this prevents efficient joint integration of linguistic, visual-semantic, spatial, and action representations. Second, existing RL fine-tuning methods for VLAs have structural drawbacks: PPO requires a value network of comparable size to the policy, while GRPO's token-level importance ratios are reported to be unstable and prone to training collapse. OmniVLA-RL addresses both by moving spatial reasoning into the backbone itself and by performing sequence-level (action-block-level) policy optimization.
Architecture
The model is built on the Mixture-of-Transformers backbone with three specialized experts sharing Transformer layers:
- Reasoning Expert: initialized from a pre-trained VLM (PaLiGemma weights), using SigLIP as the vision encoder over multi-view observations, producing semantic and linguistic tokens.
- Spatial Expert: uses VGGT to extract fine-grained 3D features from multi-view scenes. A lightweight Transformer decoder serves as an auxiliary spatial head during training, supervised by reconstruction losses on point clouds, camera parameters, and surface normals (inspired by π3). This head is discarded at inference.
- Action Expert: generates action chunks autoregressively via Conditional Flow Matching, conditioned on the fused spatial-semantic representation.
A key architectural element is Block-wise Causal Attention. Reasoning and spatial tokens form an omni-visible, bidirectionally attended prefix; action chunks form a causal suffix that can attend to the full prefix but only to preceding action tokens. Critically, prefix tokens are masked from attending to action tokens, preventing stochastic denoising noise from contaminating scene understanding. This design preserves perceptual fidelity while enforcing autoregressive causality during action synthesis.
Flow-GSPO
Because flow matching defines a deterministic ODE denoising path, it is incompatible with the stochastic exploration required by policy gradient RL. The authors reformulate the ODE as an SDE via the Fokker-Planck equation, injecting noise with schedule στ=σmax(1−τ). The discretized update (Euler-Maruyama) yields Gaussian transition probabilities, from which sequence likelihoods are computed as products over denoising steps.
GSPO is then applied at the action-block level: the importance ratio is the length-normalized exponential of summed log-likelihood ratios over the entire action chunk (action horizon H=16 times K=10 denoising steps), advantages are computed by group normalization of cumulative task rewards over G=8 sampled blocks per state, and a KL penalty (β=0.01) between old and new policies is added to the clipped objective. The authors provide a full gradient derivation showing that the gradient reduces to a weighted sum of velocity-field gradients, with the weighting involving the advantage, importance ratio, and a θ-dependent SDE drift coefficient [1+στ2(1−τ)δ/2].
The stated rationale is that action-block-level optimization captures temporal dependencies and multimodal action distributions better than token-level GRPO, whose per-token ratios accumulate bias and disrupt action continuity.
Training procedure
Training follows three stages. Stage I jointly trains the Reasoning and Spatial Experts on large-scale 3D datasets with the Action Expert frozen, using point-cloud, camera, and surface-normal reconstruction losses. Stage II unfreezes the Action Expert and trains end-to-end on the full DROID dataset with the CFM loss. Stage III applies Flow-GSPO online RL with all parameters unfrozen, using a binary task-completion reward plus a continuous gripper-alignment reward, AdamW at learning rate 10−5, and 200 RL update steps with a rollout buffer refreshed every 10 steps.
Results
On LIBERO, OmniVLA-RL ranks first in all four task suites, with an average success rate of 97.6% versus 96.9% for π0.5 and 95.7% for στ=σmax(1−τ)0. Gains over στ=σmax(1−τ)1 are modest (0.4% on Spatial, 0.5% on Goal), but the margin on LIBERO-Long is 1.1%, which the authors note is meaningful given compounding errors in long-horizon tasks. Relative to στ=σmax(1−τ)2, the improvement is 21.1% absolute. It should be noted that the strongest baselines are already near saturation on LIBERO, so the headline number should be interpreted as confirming parity-or-better rather than a large capability gap on this benchmark.
The more informative evaluation is LIBERO-Plus, where the SFT-only model achieves only 41.2% success. RL fine-tuning results:
| Configuration |
Success rate |
στ=σmax(1−τ)3 |
| OmniVLA-RL (SFT only) |
41.2% |
— |
| + PPO |
78.7% |
+37.5% |
| + GRPO |
65.7% |
+24.5% |
| + Flow-GSPO |
80.3% |
+39.1% |
| w/o Spatial Expert |
32.9% |
−8.3% |
Flow-GSPO outperforms PPO by 1.6% and GRPO by 14.6% in final success rate, and crosses 70% success within the first 50 training steps. The authors attribute the stability advantage to the sequence-level importance ratio and the action-block KL term, noting that PPO exhibits performance regressions around step 80 while Flow-GSPO improves monotonically.
The ablation identifies the Spatial Expert as the most consequential architectural component: removing it costs 8.3% success rate, which the authors link to degraded localization of small objects and handling of occlusions. This supports the central claim that geometric grounding must be integrated within the backbone rather than at the periphery.
Limitations and open questions
The paper concedes two significant limitations. All results are from simulation; the sim-to-real gap and behavior under physical hardware constraints remain untested, so the LIBERO numbers should not be assumed to transfer. Second, the architecture lacks a world model for structured long-horizon reasoning and transition prediction; action-block optimization improves short-horizon coherence, but the paper does not demonstrate compositional planning beyond the LIBERO-Plus task structure. Additional questions left open include the sensitivity of Flow-GSPO to the noise schedule στ=σmax(1−τ)4, group size, and reward shaping, none of which are ablated, and whether the auxiliary spatial supervision (Stage I) is necessary when large-scale 3D pretraining data are unavailable.
Conclusion
OmniVLA-RL presents a coherent integration of three ideas: tri-expert MoT modeling with block-wise causal attention for noise-free spatial-semantic grounding, stochastic flow matching derived from the Fokker-Planck transformation of the CFM ODE, and sequence-level GSPO optimization at the action-block granularity. The empirical evidence—97.6% average on LIBERO and a 39.1% absolute gain over SFT on LIBERO-Plus, with clear superiority over GRPO and a modest margin over PPO—supports the claim that action-block-level optimization suits continuous, multimodal robot action distributions better than token-level alternatives. The principal caveat is that validation is entirely in simulation, leaving real-world deployment as the most pressing unresolved question.