Cross-platform generalization of contact-level failure detection
Determine whether the contact-level performance gap observed for parallel-jaw tabletop robot arms persists on humanoid, mobile-manipulation, and dexterous-hand platforms.
References
Whether the contact-level gap we measure looks the same on those platforms is untested.
— FailBench: How Reliable are VLMs at Judging Robot Task Success?
(2609.03611 - Navasardyan et al., 3 Sep 2026) in Section 6, Limitations, subsection “Embodiment”