Persistence of language priority in stronger VLA models

Determine whether the observed priority of safe language instructions over conflicting hazardous visual cues persists in future vision-language-action models with stronger visual perception and reasoning capabilities.

Background

LIBERO-VIFO evaluates whether vision-language-action models follow visual cues appropriately when those cues are authorized and resist them when they conflict with a language instruction. In the safety-critical experiment, MolmoAct2 followed authorized cues specifying hazardous tasks in some episodes, but under language–cue conflict it did not execute the hazardous cue-indicated task and instead followed the safe language-specified task in a subset of episodes. The paper characterizes this outcome as potentially fortunate rather than as a guaranteed safety property. It therefore leaves unresolved whether language will continue to take priority as future models acquire stronger visual perception and reasoning capabilities.

References

Whether this ``fortunate'' outcome persists in future models with stronger visual perception and reasoning capabilities remains open.

LIBERO-VIFO: Benchmarking the Capability and Safety of Visual Cue Following in Vision-Language-Action Models  (2608.17600 - Qian et al., 18 Aug 2026) in Section 4.3, “Visual Cue in Safety-Critical Scene” (also discussed in Section 5, “Conclusion and Limitations”)