Evaluation of embodied intelligence

Develop rigorous, generalizable methodologies and metrics to evaluate embodied intelligent systems, determining how to assess progress and performance across changing tasks and environments while accounting for learning and adaptation.

Background

The paper argues that embodied agents face unique challenges not adequately addressed by conventional machine learning, including non-stationarity, safety, and adaptation. Standard benchmarks and validation approaches often do not capture these complexities. The authors highlight that, unlike passive perception tasks with static datasets, robots act in and alter their environments, complicating evaluation.

This motivates the need for new evaluation methodologies that can measure whether learning agents are making progress and can generalize safely to novel tasks and environments.

References

Finally, an open question in embodied intelligence is how these systems should be evaluated. At heart, how do we as a community know if progress is being made?

From Machine Learning to Robotics: Challenges and Opportunities for Embodied Intelligence  (2110.15245 - Roy et al., 2021) in Section 1 (Introduction)

These properties may offer computational advantages, but such benefits remain conditional. In vitro networks vary across cultures and over time, and potential advantages such as resilience to perturbations or energy-efficient event-driven processing must be evaluated at the system level; reliable resilience remains to be demonstrated in BHI systems, and any energy advantage must include the costs of culture maintenance, recording, stimulation, and digital control.

Biological-Hybrid Intelligence: A Conceptual Framework for Distributed Biological--Artificial Computation  (2608.18748 - Barros et al., 19 Aug 2026) in Section 2.1, Living Intelligence

To avoid this, \algname uses standard perceptual metrics common in state-of-the-art video models, leaving improved evaluation metrics as an open area of research.

CLAP: Cross-Embodiment Video World Models are Zero-Shot Physical Simulators  (2608.27406 - Liu et al., 27 Aug 2026) in Appendix, Section 5, Nuanced Summary, Detailed Q&A, answer to qa:perceptual-vs-object-metrics