Cross-platform generalization of contact-level failure detection

Determine whether the contact-level performance gap observed for parallel-jaw tabletop robot arms persists on humanoid, mobile-manipulation, and dexterous-hand platforms.

Background

All sources in FailBench primarily involve parallel-jaw arms performing tabletop manipulation, with only a limited number of bimanual examples and no humanoid, mobile-manipulation, or dexterous-hand data.

The benchmark finds that detectors perform poorly when success depends on establishing fine-grained physical contact. The authors explicitly state that it has not been tested whether this contact-level gap has the same character on other embodiments, leaving cross-platform generalization unresolved.

References

Whether the contact-level gap we measure looks the same on those platforms is untested.

FailBench: How Reliable are VLMs at Judging Robot Task Success?  (2609.03611 - Navasardyan et al., 3 Sep 2026) in Section 6, Limitations, subsection “Embodiment”