Papers
Topics
Authors
Recent
Search
2000 character limit reached

LIBERO-Plus: VLA Robustness Benchmark

Updated 6 September 2026
  • LIBERO-Plus is a benchmark that tests the robustness and generalizability of vision-language-action (VLA) models in robotic manipulation by introducing perturbations in various dimensions including visual appearance, language instructions, and robot states.
  • Primary robustness metrics include success rate and degradation from the unperturbed condition, with the benchmark comprising 10,030 tasks across seven dimensions including camera viewpoint, lighting, and sensor noise.
  • The benchmark reveals substantial model-specific variation in performance across different perturbations and underscores factors like language grounding, geometric reasoning, and visual invariance which are essential for robust performance

LIBERO-Plus is a robustness and distribution-shift extension of the LIBERO robotic-manipulation benchmark for evaluating vision-language-action (VLA) models under controlled changes in visual appearance, camera configuration, robot state, language, sensing, and scene geometry. The benchmark contains 10,030 perturbed test tasks spanning seven dimensions: camera viewpoint, robot initial state, language instruction, lighting, background texture, sensor noise, and object layout. It was introduced to test whether high success rates on standard LIBERO reflect transferable perception, language grounding, geometric reasoning, and closed-loop control rather than interpolation or memorization of fixed scenes and trajectories (Fei et al., 15 Oct 2025).

1. Benchmark scope and construction

Standard LIBERO evaluates four task suites: Spatial, Object, Goal, and Long. Its canonical evaluation uses fixed scenes, camera configurations, robot initial states, lighting, backgrounds, and task instructions. LIBERO-Plus preserves the underlying LIBERO task families while introducing systematic perturbations to the evaluation distribution.

The benchmark construction begins with the 40 original LIBERO evaluation tasks. For each of the four suites and seven generalization dimensions, 500 candidate instances are generated, producing an initial pool of 14,000 candidate tasks. Tasks solved by all or most evaluated models are removed to reduce ceiling effects, and the remaining tasks are balanced across perturbation subcategories. The resulting benchmark contains 10,030 test tasks organized around seven dimensions and 21 low-level components.

LIBERO suite Camera Robot Language Light Background Noise Layout Total
Spatial 376 350 354 292 258 351 312 2,293
Object 396 398 390 297 248 422 425 2,576
Goal 408 409 410 279 281 379 403 2,569
Long 419 393 383 274 289 449 385 2,592
Total 1,599 1,550 1,537 1,142 1,076 1,601 1,525 10,030

The benchmark uses five empirically defined difficulty levels based on four representative models—OpenVLA-OFT, π0\pi_0, π0\pi_0-Fast, and UniVLA:

  • L1: all four models succeed.
  • L2: exactly three models succeed.
  • L3: exactly two models succeed.
  • L4: exactly one model succeeds.
  • L5: none of the four models succeeds.

The difficulty levels are model-dependent because they are defined relative to the selected representative systems. Complete numerical success curves for every model and every L1–L5 or perturbation-intensity condition are not provided.

LIBERO-Plus is generally evaluated as an out-of-distribution test rather than as a training set. In several studies, including VLA-JEPA, JEPA-VLA, LaMP, ELAN4D, 3DThinkVLA, LIRA, and OA-WAM, policies are trained on standard LIBERO demonstrations and evaluated directly on perturbed tasks. Other studies report supervised fine-tuning on LIBERO-Plus training data, so results must be interpreted according to the stated protocol.

2. Perturbation dimensions and implementation

LIBERO-Plus separates robustness into seven factors. These factors probe different capabilities: camera and robot perturbations primarily test geometric reasoning and replanning; language tests instruction use; lighting, background, and noise test visual invariance; and layout tests object identity, scene relations, and positional generalization.

Object layout

Object-layout perturbations contain two subdimensions. In the confounding-object condition, unseen distractors are randomly added from a predefined set of 416 object categories. The task BDDL files are modified, and perturbed files are marked with an add suffix. In the target-object-pose condition, the target’s position (x,y,z)(x,y,z) and orientation (pitch,yaw,roll)(\text{pitch},\text{yaw},\text{roll}) are independently perturbed while preserving semantic relations to other objects.

Distractors are often tolerated, whereas target displacement causes substantial degradation. This pattern is interpreted as evidence of positional bias: a policy may identify or ignore additional objects while still relying on the target’s memorized location. Typical failures include target mislocalization, replay of a trajectory appropriate to the original position, recognition confusion, and collision-prone motion.

Camera viewpoint

Camera perturbations modify the third-person camera through distance, spherical position, and orientation:

  1. Camera distance is scaled by 1.01×1.01\times–2.00×2.00\times relative to the scene center.
  2. Azimuth and elevation are changed within 15∘15^\circ–75∘75^\circ cones.
  3. Yaw, pitch, and roll are perturbed by 2∘2^\circ–10∘10^\circ.

These changes are implemented through the LIBERO Problem interface and task filename parameters. Camera variation produces some of the most severe failures because it changes image geometry without changing task semantics. Errors include viewpoint-dependent object localization, incorrect spatial relations, and failure to transfer trajectories to a new camera configuration.

Robot initial state

Robot-initial-state perturbations randomly alter the robot joint positions (qpos) with reported magnitudes from 0.1 to 0.5. The resulting distribution shift requires a policy to re-plan from a different arm configuration and compensate for altered kinematics. Failures include incorrect reaching trajectories, poor end-effector alignment, and replay of trajectory patterns tied to the training initialization.

Language instruction

Language perturbations are generated using LLM-based rewrites in three categories:

  • Distraction: longer conversational instructions containing irrelevant context.
  • Commonsense: replacement of canonical object descriptions with commonsense descriptions.
  • Reasoning-chain: altered complexity or expression of multistep instructions.

For example, “push the plate to the front of the stove” may be rewritten as “before turning on the burner, push the plate to the front of the stove,” or as a commonsense description of the plate and stove.

Small performance drops under language rewrites are not sufficient evidence of language understanding. Blank-instruction tests show that OpenVLA-OFT performance can remain largely unchanged on the Object suite when language is replaced by an empty string. In goal-replacement tests, changing only the target noun can reduce success nearly to zero, with models continuing to execute the original visual-action mapping. These findings indicate that language invariance may reflect language under-utilization rather than robust paraphrase grounding.

Lighting

Lighting perturbations modify diffuse color, light direction, specular intensity, and shadow casting through scene XML files. Third-person-only models can be highly sensitive to these changes because shadows and global appearance affect localization. Models with wrist cameras often retain stronger local geometric and contact cues.

An all-black camera ablation causes performance to collapse to nearly zero, while masking only the third-person view can preserve substantial performance when the wrist view remains available. This supports the interpretation that wrist observations provide relatively stable local geometry under changes in global illumination.

Background texture

Background perturbations modify scene themes and surface appearance. Scene themes use a curated collection of 950 textures, while tabletop or floor textures are altered separately. Background changes are relatively benign for some models but severely damaging for others, indicating that background invariance is architecture- and training-dependent.

A model’s ability to ignore distractor objects or irrelevant background regions should not be equated with complete scene understanding. Policies may still rely on fixed object positions or other scene regularities.

Sensor noise

Sensor-noise perturbations contain five image corruptions with five severity levels:

  • Motion blur, parameterized by radius, Gaussian spread, and angle π0\pi_00.
  • Gaussian blur with standard deviation from π0\pi_01 to π0\pi_02.
  • Zoom blur with weak and strong rescaling ranges.
  • Fog with density π0\pi_03 and decay rate π0\pi_04.
  • Glass blur with Gaussian blur, pixel displacement, and iteration count.

Typical failures involve corrupted visual features, loss of object boundaries, and degraded localization. Robustness varies substantially across VLA architectures.

3. Metrics, evaluation protocols, and interpretation

The primary metric is task or episode success rate, reported as a percentage. For a perturbation π0\pi_05 and model π0\pi_06, the success rate is the proportion of evaluated episodes that satisfy the benchmark’s binary task-completion criterion.

The absolute degradation from the unperturbed condition is defined as

π0\pi_07

where π0\pi_08 is the unperturbed success rate and π0\pi_09 is the success rate under a specified perturbation.

Some studies report an aggregate average over the seven perturbation categories, while others report a task-weighted overall score across all 10,030 tasks. These values are not necessarily identical. For example, LIRA explicitly reports that its Overall score is computed over all 10,030 tasks rather than as the unweighted mean of the seven displayed category values (Zhang et al., 6 Aug 2026).

The benchmark has also been used in compositional experiments involving simultaneous perturbations. OpenVLA-OFT was evaluated over 2,000 independent repeated experiments using six non-language categories: layout, background/environment, illumination, camera, robot initialization, and noise. Combined perturbations were often worse than expected from independent single-factor effects. Statistical dependence was examined using a (x,y,z)(x,y,z)0 chi-square test, and most perturbation pairs showed significant dependence. The findings support the characterization of robustness as non-decomposable: representations entangle visual, geometric, and sensor factors, so joint shifts expose failures not predicted by isolated tests (Fei et al., 15 Oct 2025).

Comparisons require careful protocol control. Relevant differences include:

  • whether the model is trained only on standard LIBERO or fine-tuned on LIBERO-Plus;
  • whether the model is trained jointly across suites or separately for each suite;
  • whether the reported score is task-weighted or factor-averaged;
  • whether results are averaged over multiple seeds;
  • whether the evaluation uses one rollout per task or repeated episodes;
  • whether baseline numbers are reproduced or taken from prior papers.

Many LIBERO-Plus studies report point estimates without confidence intervals, standard deviations, random-seed averages, or hypothesis tests. Consequently, numerical differences should generally be interpreted as empirical benchmark comparisons rather than formally established statistical superiority.

4. Robustness findings across VLA models

The original LIBERO benchmark can produce success rates near 95–98% for contemporary VLAs, whereas modest perturbations can reduce performance below 30%. The main single-dimension results demonstrate strong model-specific variation.

Model Original Camera Robot Language Light Background Noise Layout
OpenVLA 76.5 1.1 4.1 26.8 4.4 25.3 19.3 31.6
OpenVLA-OFT 97.1 59.7 37.2 81.5 85.8 92.4 76.7 77.1
(x,y,z)(x,y,z)1 94.2 15.8 6.6 61.0 79.6 78.5 79.4 70.4
(x,y,z)(x,y,z)2-Fast 85.5 66.4 24.8 63.3 73.0 67.7 75.8 70.3
Nora 87.9 4.0 41.1 67.0 31.0 50.5 17.6 63.9
WorldVLA 79.1 0.3 30.2 44.2 29.4 14.5 12.2 39.4
UniVLA 95.2 4.3 50.3 71.8 59.1 80.0 25.3 34.3
RIPT-VLA 97.5 58.3 36.7 80.1 87.9 90.4 73.8 76.5

Camera viewpoint and robot initialization are generally the most damaging factors. OpenVLA falls from 76.5% to 1.1% under camera perturbations and to 4.1% under robot-state perturbations. (x,y,z)(x,y,z)3 falls from 94.2% to 15.8% and 6.6%, respectively. Even models with strong canonical performance can therefore exhibit severe brittleness under modest geometric changes.

Wrist-camera input substantially improves robustness in some settings. OpenVLA-OFT with wrist and third-person views achieves 59.7% under camera perturbation, compared with 16.8% for the third-person-only version. The difference supports the interpretation that wrist views provide a stable local geometric reference and reduce dependence on a particular global camera pose.

Training and architecture affect robustness unevenly. (x,y,z)(x,y,z)4-Fast is comparatively strong on camera and sensor noise, while OpenVLA-OFT is strong under background and lighting changes. WorldVLA is weak across most categories in the reported table. These differences prevent LIBERO-Plus from being interpreted as measuring a single scalar property of VLA quality.

The central language finding is counterintuitive. Language perturbations generally cause smaller decreases than camera or robot changes, but blank-instruction and goal-replacement tests show that the apparent language robustness can result from weak language use. A model that ignores language may appear invariant to paraphrase while failing when the instruction changes the intended object.

5. Methods developed for LIBERO-Plus robustness

LIBERO-Plus has become a testbed for architectural, representational, and training interventions.

Predictive and latent-state representations

VLA-JEPA uses leakage-free latent state prediction with a frozen V-JEPA2 target encoder. Future frames are used only as supervision targets, never as inputs to the student pathway. Its reported LIBERO-Plus average is 79.5%, compared with 69.6% for OpenVLA-OFT. It leads on Robot, Language, Light, Background, and Layout, but is not best on Camera or Noise (Sun et al., 10 Feb 2026).

JEPA-VLA integrates frozen V-JEPA2 embeddings into VLA models. In a one-tenth-data basic-VLA setting, it improves the aggregate score from 18.9% to 25.6%. The largest gains occur under background, lighting, language, and layout perturbations, while camera and noise performance remain weak (Miao et al., 12 Feb 2026).

Geometric and motion priors

LaMP introduces dense 3D scene flow as a latent motion prior. Its Motion Expert represents image-plane and depth displacement over a (x,y,z)(x,y,z)5 grid and conditions an Action Expert through gated cross-attention. LaMP reaches 79.3% on LIBERO-Plus, 9.7 points above OpenVLA-OFT. Removing the motion pathway reduces performance to 71.6%, with the largest losses under camera and robot perturbations (Wang et al., 26 Mar 2026).

ELAN4D uses future robot keypoint tracks derived from forward kinematics. Seven arm joints and the end-effector provide eight keypoints, and a lightweight control branch predicts future displacements during training. The track decoder is discarded at inference. ELAN4D((x,y,z)(x,y,z)6) obtains 78.2%, compared with 73.6% for (x,y,z)(x,y,z)7, with particularly notable gains under background, noise, and robot-initialization shifts (He et al., 28 May 2026).

3DThinkVLA co-trains latent geometry perception and spatial reasoning. It uses 3D foundation-model supervision during training but retains only lightweight adapters at deployment. The reported LIBERO-Plus average is 81.0%, with especially strong performance on Camera at 73.8% and Light at 98.4%. Its advantage is not uniform: it is below competing systems on Language, Robot, Background, Noise, and Layout (Shi et al., 3 Jun 2026).

Object binding and target grounding

TAG performs inference-time contrastive guidance by evaluating a VLA on the normal observation and on a target-agnostic observation. The residual between the two policy predictions is amplified using a classifier-free-guidance-style rule. With (x,y,z)(x,y,z)8, TAG improves the LIBERO-Plus average from 81.4% to 86.1% using a static background baseline and to 87.2% using a black baseline. The strongest gains occur under Camera, Robot, and Noise perturbations, while language performance can decrease (Zhou et al., 25 Mar 2026).

OA-WAM decomposes each frame into a robot slot and object slots, separating persistent address vectors from time-varying content vectors. Address-only keys and address resetting constrain cross-slot routing. OA-WAM achieves an 84.3% geometric average over Camera, Robot initialization, and Layout, the strongest reported geometric score in its comparison, but its overall seven-axis average of 83.9% remains below (x,y,z)(x,y,z)9 at 85.7%. Its sensor-noise score is substantially weaker, illustrating dependence on the upstream SAM3/DINOv3 slot-extraction pipeline (Liu et al., 7 May 2026).

Action and conditioning interfaces

CAC-VLA trains the VLM to predict compact latent encodings of future action segments and uses a context gate to modulate their influence on a continuous flow-matching action expert. It obtains 83.8% in zero-shot transfer from LIBERO to LIBERO-Plus and 89.5% after supervised fine-tuning on LIBERO-Plus. The latent-action and context-gating ablations are primarily reported on standard LIBERO rather than directly on LIBERO-Plus (Xiong et al., 6 Jul 2026).

LIRA modifies the VLA-Adapter interface by routing local neighborhoods of intermediate VLM layers into action-decoder blocks. In zero-shot LIBERO-Plus transfer, it improves from 59.1% for VLA-Adapter to 78.0%, an 18.9-point absolute gain. The strongest controlled evidence comes from the matched 0.5B comparison; broader comparisons involve models with different backbones, action representations, and training procedures (Zhang et al., 6 Aug 2026).

VLANeXt combines Qwen3-VL-2B-Instruct, a dedicated policy module, soft VLM-policy coupling, wrist and third-person views, VLM-side proprioception, flow matching, and a frequency-domain action loss. It reports 80.1% on LIBERO-Plus, compared with 69.6% for OpenVLA-OFT. Its advantages are largest under Robot and Camera perturbations, although OpenVLA-OFT remains slightly better on Background (Cheng et al., 18 May 2026).

Online reinforcement learning

OmniVLA-RL combines semantic, spatial, and action experts with Flow-GSPO, an action-block-level policy-optimization method for stochastic flow matching. In the reported aggregate LIBERO-Plus ablation, SFT-only training reaches 41.2%, PPO reaches 78.7%, GRPO reaches 65.7%, and Flow-GSPO reaches 80.3%. The evidence is narrower than a full benchmark leaderboard because the manuscript does not provide complete suite-level or task-level results (Jie et al., 20 Apr 2026).

6. Limitations and significance

LIBERO-Plus exposes weaknesses that standard LIBERO can conceal, but it is not a complete measure of robotic competence.

The benchmark is simulation-based and inherits LIBERO’s visual and physical assumptions. Synthetic perturbations may not capture the full distribution of real calibration errors, contact dynamics, occlusions, sensor failures, and embodiment variation. Several papers use simulator pose, privileged object state, or external perception systems during training, which can complicate comparisons.

Perturbation definitions and evaluation details are not uniform across all reports. Some papers provide detailed task counts and generation procedures, while others report only the seven category names. Episode counts range from one rollout per task to repeated episodes or two episodes per perturbation variant, depending on the study. Aggregate scores may be factor-averaged or task-weighted. Baseline values may be reproduced locally or imported from prior work.

The benchmark also does not fully resolve language grounding. A model can perform well under language rewrites while ignoring the instruction, as demonstrated by blank-instruction and target-replacement tests. Conversely, MINERVA maps rephrased instructions to externally supplied task IDs, so its high language score does not establish natural-language understanding (Sendai et al., 3 Sep 2026).

Model capacity alone is insufficient for robust performance. MINERVA reaches 95.1% on standard LIBERO with 0.54 million parameters but falls to 46.7% on LIBERO-Plus. Its performance under light and background changes is near zero even at 4.89 million parameters, indicating that robustness depends on representation and training variation rather than only parameter count (Sendai et al., 3 Sep 2026).

The principal implication is methodological: canonical LIBERO success should not be treated as sufficient evidence of reliable autonomy. Evaluation should include camera pose, robot initialization, object displacement, distractors, lighting, background, noise, language changes, and combinations of these factors. Stronger evaluation should also include explicit language-grounding tests, state-conditioned replanning, view-invariant representations, proprioceptive reasoning, causal object-binding interventions, uncertainty estimates, matched seeds, and task-level failure analyses.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to LIBERO-Plus.