Safe behavior and evaluation of vision-language models controlling real vehicles

Determine how general-purpose vision-language models should behave when given control of a real-world vehicle and how that behavior should be evaluated properly.

Background

DrivingBench evaluates general-purpose vision-LLMs that directly control a Toyota Corolla through camera observations and steering and speed commands. The evaluation shows that model behavior depends strongly on task framing, tool outputs, latency, perception, and planning, with only one evaluated model completing the low-speed cone course.

The paper identifies an unresolved broader problem concerning the appropriate behavior of models entrusted with physical control and the development of evaluation methods that can assess such behavior. This extends beyond course completion to questions of safety, alignment, and the proper assessment of real-world vehicle control.

References

How such models should behave when handed control of a real-world vehicle, and how to evaluate that behavior properly, is a question we leave for future work.

— DrivingBench: Can Vision-Language Models Drive a Toyota Corolla?  (2609.38948 - Ramabadran et al., 30 Sep 2026) in Section 7, Conclusion

Third, safety and reliability are unresolved concerns, since LLMs may hallucinate, misinterpret ambiguous instructions, or generate unsafe control suggestions.

— A Survey on End-to-End Autonomous Driving Training from the Perspectives of Data, Strategy, and Platform  (2610.00926 - Xu et al., 1 Oct 2026) in Section 4.4, “Multimodal Large Language Models”