Automated, reliable task-success evaluation for robotic manipulation
Develop a fully automated and reliable methodology to assess task success in real-world robotic manipulation videos, robust to diverse failure modes and ambiguous semantics, eliminating reliance on proxy metrics and partial human validation in evaluation suites such as the Embodied World Model Benchmark (EWMBench).
References
Fully automated and reliable assessment of task successâparticularly under diverse failure modes and ambiguous semanticsâremains an open challenge.
The open problem is calibration under shift: hidden-state probes lose accuracy under clutter, lighting, novel objects and reworded instructions, and SAFECAST recovers part of it by training the probe on contrast-set perturbations \citep{safecast2026}, while Foresight makes the threshold itself adaptive \citep{foresight2026}.
These results suggest that robot failure detection remains an open problem even for current general-purpose VLMs.