HumanCLAW: Can Vision-Language Models Act Through a Body?
This presentation examines a critical gap in artificial intelligence: while vision-language models excel at recognizing objects and understanding scenes, they fundamentally struggle to act through a physical body. HumanCLAW reveals that state-of-the-art models fail at embodied tasks not because they can't see or move, but because they lack awareness of their own body in space—a bottleneck that exposes the frontier between perception and true embodied intelligence.Script
State-of-the-art vision models can describe a room in exquisite detail, yet most cannot walk across it and sit on a chair. The researchers behind HumanCLAW ask a deceptively simple question: can vision-language models act through a body?
The challenge isn't perception or planning. HumanCLAW decouples high-level decisions from motor control to isolate what the authors call action intelligence: the moment-to-moment bodily decision-making required when your choices have physical consequences.
The benchmark is unforgiving. Agents must complete over 1,200 find-navigate-interact episodes across 41 indoor scenes, issuing atomic skills like walk, turn, and sit at sub-second intervals. The system tracks not just success, but collisions, object disturbances, and whether the agent even knows when it has arrived.
Gemini 3.1, the top performer, recognizes target objects 65 percent of the time but completes the full task in only 17 percent of episodes. The collapse happens after perception succeeds. Agents stop short of targets, misjudge arrival, or attempt to sit while spatially misaligned, revealing a profound deficit in embodied self-awareness.
The bottleneck is not actuation. The motion synthesis system executes commanded skills with near-perfect fidelity. Instead, 81 percent of post-perception failures stem from the inability to maintain an online belief about where the body is relative to the environment—a capacity humans take for granted.
HumanCLAW exposes the frontier between perception and true embodied intelligence. The path forward demands architectures that don't just see the world but inhabit it—integrating proprioceptive signals, spatial memory, and consequence prediction. Explore the full paper and create your own videos at EmergentMind.com.