Persistence of language priority in stronger VLA models
Determine whether the observed priority of safe language instructions over conflicting hazardous visual cues persists in future vision-language-action models with stronger visual perception and reasoning capabilities.
References
Whether this ``fortunate'' outcome persists in future models with stronger visual perception and reasoning capabilities remains open.
— LIBERO-VIFO: Benchmarking the Capability and Safety of Visual Cue Following in Vision-Language-Action Models
(2608.17600 - Qian et al., 18 Aug 2026) in Section 4.3, “Visual Cue in Safety-Critical Scene” (also discussed in Section 5, “Conclusion and Limitations”)