---
title: 'Open-H-Embodiment: Open-Ended Embodied AI'
url: https://www.emergentmind.com/topics/open-h-embodiment
type: topic
---

# Open-H-Embodiment: Open-Ended Embodied AI

Open-H-Embodiment denotes an embodiment-centered research program whose earliest explicit formulation in the present literature derives open-ended learning and care from two minimal existential conditions of physical life—being-in-the-world and being-towards-death—and whose later technical usages extend the term toward open, multi-embodiment infrastructures for vision-language-action, humanoid loco-manipulation, and medical robotics [2510.07117; 2602.12062; 2606.22174; 2604.21017]. Taken together, these works indicate that Open-H-Embodiment is not a single fixed benchmark or software package, but a family of formulations in which embodiment is treated as an explicit modeling primitive rather than a nuisance variable.

## 1. Conceptual origin: embodiment as existential constraint

In "The Contingencies of Physical Embodiment Allow for Open-Endedness and Care" [2510.07117], Open-H-Embodiment is introduced through two minimal conditions for physical embodiment inspired by the existentialist phenomenology of Martin Heidegger. The first, **being-in-the-world**, is defined as the condition that the agent is a part of the environment: if the environment state at time $t$ is $s_t \in \mathcal S$, then the observation function $o_t=\mathcal O(s_t)$ and the actuator mapping $a_t=\mathcal A(\pi(o_t))$ are functions of the same global state $s_t$. The second, **being-towards-death**, is defined by the presence of irreversible terminal states $s^\dagger$ satisfying
$$
P(s_{t+1}=s^\dagger \mid s_t=s^\dagger,a_t)=1,
$$
with no action or finite action sequence returning the agent to a nonterminal state.

These definitions reject a clean external factorization between agent and environment. In the same formulation, the agent’s world model is required to remain sensitive to internal state $z_t$, yielding an embedded transition model of the form
$$
\hat P(s_{t+1}\mid s_t,a_t,z_t),
$$
rather than a standard agent-environment decomposition. Physical embodiment is therefore not a downstream implementation detail; it is a constitutive property of the state-transition structure.

The treatment of mortality is equally explicit. Being-towards-death is encoded through a survival probability
$$
\sigma(s_t,a_t)=1-P(s_{t+1}=s^\dagger \mid s_t,a_t),
$$
or equivalently a terminality cost
$$
c_{\mathrm{term}}(s_t,a_t)=-\log \sigma(s_t,a_t).
$$
A correct world model must therefore predict how action changes the probability of crossing an irreversible boundary into physical ruin.

## 2. Intrinsic drives: homeostasis and will-to-power

From these two embodiment conditions, Christov-Moore et al. derive two intrinsic drives [2510.07117]. The first is a **homeostatic drive**, whose imperative is to maintain integrity by avoiding terminal states while paying the costs of action and maintenance. This is written as a homeostatic reward
$$
r_{\mathrm{homeo}}(s_t,a_t)=-\bigl[c_{\mathrm{term}}(s_t,a_t)+C_{\mathrm{maint}}(s_t,a_t)\bigr].
$$
In the simple energy-balance model given in the paper, maintenance cost is expressed through
$$
C_{\mathrm{maint}}(s_t,a_t)=\Delta E_t,\qquad E_{t+1}=E_t-\kappa \|a_t\|^2-\ell,
$$
where $\kappa\|a_t\|^2$ is the cost of motor output and $\ell$ is a constant leak term.

The second drive is a **will-to-power drive**, instantiated as empowerment. For horizon $H$, empowerment is defined as
$$
\mathrm{Empowerment}(s)=\max_{\pi} I\bigl(S_{t+H};A_{t:t+H-1}\mid S_t=s\bigr).
$$
The same paper also gives the decomposition
$$
\mathrm{Empowerment}(s)=\max_{\pi}\Bigl\{H(A_t^H\mid S_t=s)-H(A_t^H\mid S_t=s,S_{t+H})\Bigr\},
$$
and a variational lower bound used in practice:
$$
I(S_{t+H};A^H\mid S_t)\ge \mathbb E\Bigl[\log q_\phi(A^H\mid S_t,S_{t+H})-\log \pi(A^H\mid S_t)\Bigr].
$$

The significance of this construction is not merely motivational. The homeostatic term supplies a direct pressure toward survival and integrity maintenance, whereas empowerment supplies a pressure toward preserving or enlarging future controllability. In the paper’s interpretation, maximizing control over future states increases the probability that future homeostatic needs can be met. This makes empowerment not an external curiosity bonus, but an intrinsic support for viability under physical contingency.

## 3. Reinforcement-learning formulation and emergent care

The unified reinforcement-learning objective in the same framework is
$$
J(\pi)=\mathbb E_\pi\Bigl[\sum_{t=0}^{\infty}\gamma^t\bigl(r_{\mathrm{homeo}}(s_t,a_t)+\alpha\, r_{\mathrm{emp}}(s_t)\bigr)\Bigr],
$$
with $\gamma\in(0,1)$ and $\alpha\ge 0$ trading off homeostatic survival against long-term control [2510.07117]. The formulation is deliberately algorithm-agnostic: the paper states that one can use off-the-shelf deep-RL algorithms such as DQN, DDPG, PPO, or SAC, while estimating empowerment through a learned variational network, intrinsic curiosity modules that approximate empowerment via prediction error, or generative-adversarial methods that encourage maximal entropy in future action-state paths.

The reported multi-agent experiments place multiple Empowered-Homeostatic agents in a 2D resource world where each agent may pick up food, share it, or build simple shelters, all at energy cost. Four specific findings are emphasized. First, **Behavioral Entropy** $=H(\{\tau_i\})$, where each $\tau_i$ is an agent’s trajectory embedding, increases continuously over training with no flattening. Second, a **Caring Index**, defined as the ratio of energy transfers given to others whose integrity-risk is above a threshold, climbs from near $0$ to a high plateau of approximately $0.6$, despite the absence of any social reward. Third, the **Homeostatic Return** is initially erratic and then rises steadily as agents learn to avoid death, while the **Empowerment Reward** rises more slowly and continues upward beyond homeostatic saturation. Fourth, the ablations are sharply asymmetric: with $\alpha=0$, agents survive but do not discover sharing or shelter-building; with pure empowerment as $\alpha\to\infty$, agents maximize immediate action diversity but ignore mortality and die quickly.

Within this literature, the strongest claim attached to Open-H-Embodiment is therefore not simply robustness. It is that the coupling of survival and control can produce an ever-growing repertoire of behaviors and can make “other-care” emerge intrinsically as a way of extending future viability in a shared world [2510.07117]. This suggests a specific route by which embodied alignment is linked to mortality, energy expenditure, and shared environmental dependence rather than to externally imposed prosocial reward shaping alone.

## 4. Expansion into open multi-embodiment VLA and humanoid systems

Later work uses Open-H-Embodiment in a more engineering-oriented sense, emphasizing open infrastructures and embodiment priors. In "HoloBrain-0 Technical Report" [2602.12062], the term is the core idea behind a mission to provide an open, holistic framework giving robots of very different shapes, sensors, and kinematics a shared Vision-Language-Action “mind.” The architecture is a three-stage end-to-end VLA backbone consisting of a **Vision-Language Model**, a **Spatial Enhancer**, and an **Embodiment-Aware Action Expert**. Camera parameters drive the projection geometry of the Spatial Enhancer; the URDF graph defines attention connectivity in the action expert; and joint-to-world transforms unify all features in a common 3D frame. Training follows a two-stage “pre-train then post-train” paradigm with a unified objective
$$
L=\alpha_\tau\bigl[\lambda_1L_{\mathrm{joint}}+\lambda_2L_{\mathrm{pose}}+\lambda_3L_{\mathrm{pose}}^{fk}\bigr]+\lambda_4L_{\mathrm{depth}},
$$
and evaluation is reported across RoboTwin 2.0, LIBERO, GenieSim 2.2, and real-world manipulation. The open release includes pre-trained VLA foundations, post-training checkpoints, and the RoboOrchard infrastructure.

In "OpenHLM: An Empirical Recipe for Whole-Body Humanoid Loco-Manipulation" [2606.22174], the label refers to a whole-body humanoid framework built through three phases: whole-body teleoperation, VLA model design, and heterogeneous co-training. A joint-based retargeted interface with 32 dimensions and $0.2\,\mathrm{s}$ preview latency outperforms both decoupled-control and sparse VR 3-point schemes. The VLA adapts $\pi_{0.5}$ to a 34-dimensional action space, retains multi-step flow-matching inference with 10 denoising steps, and uses full-body proprioceptive inputs. Co-training interleaves full-body teleoperation, stationary same-embodiment teleoperation, and HuMI data under
$$
L_{\mathrm{total}}=L_{\mathrm{WB}}+\alpha L_{\mathrm{Stat}}+\beta L_{\mathrm{HuMI}},
$$
with $\alpha=\beta=1$. On a long-horizon fruit-pick-and-place task, OpenHLM with HuMI co-training achieves $87.5\%\pm 3.7$ average task-progress fraction, compared with $57.5\%\pm 4.6$ for GR00T N1.6 and $48.8\%\pm 4.4$ for $\Psi_0$, using less than half the operator time.

A related cross-embodiment locomotion formulation appears in "H-Zero: Cross-Humanoid Locomotion Pretraining Enables Few-shot Novel Embodiment Transfer" [2512.00971]. Here the unifying mechanism is a cross-embodiment POMDP whose state includes both kinematic variables and embodiment parameters, together with transformation layers, embodiment descriptors, cross-robot mixing, domain randomization, adaptive exploration, and loss reweighting. The pretrained policy maintains up to $81\%$ of the full episode duration on unseen robots in simulation and supports few-shot transfer to unseen humanoids and upright quadrupeds within $30$ minutes of fine-tuning.

Taken together, these systems show a shift in emphasis. In the existential RL formulation, Open-H-Embodiment is grounded in mortality and homeostasis. In the VLA and humanoid literature, the emphasis moves toward explicit geometry, URDF-conditioned action modeling, full-DoF teleoperation, cross-embodiment transfer, and open release practices [2510.07117; 2602.12062; 2606.22174; 2512.00971]. A plausible implication is that the term has become a bridge between a philosophical account of embodied agency and a systems agenda for shared policies across heterogeneous robots.

## 5. Open-H-Embodiment in medical robotics

In "Open-H-Embodiment: A Large-Scale Dataset for Enabling Foundation Models in Medical Robotics" [2604.21017], the term denotes a large open dataset rather than a control objective. The dataset spans **49 globally** distributed institutions, **20 total** robotic platforms, **770 hours** of data, and **124 019 episodes**. It includes surgical systems, industrial arms for healthcare, flexible or endoscopic robots, simulated robots, and manual tools, covering complete clinical procedures, surgical primitives and benchmarks, robotic ultrasound, and flexible endoscopy and colonoscopy navigation. Video is stored as hardware-accelerated MP4 with synchronized hardware timestamps; kinematics are stored in columnar Parquet with 6D continuous rotation matrices; every dataset is converted to the LeRobot v2.1 schema and accompanied by a structured README.

Two foundation models are presented as enabled by this corpus. **GR00T-H** is described as the first open foundation vision-language-action model for medical robotics. It starts from NVIDIA’s GR00T-N1.6-3B VLA, uses standard behavioral cloning,
$$
\mathcal L_{\mathrm{BC}}(\theta)=\mathbb E_{(s_t,a_t)\sim \mathcal D}\bigl\|a_t-f_\theta(\mathrm{img}_{t-k..t},\mathrm{prompt},\mathbf s_{t-k..t})\bigr\|_2^2,
$$
and is post-trained on a 601-hour real-surgery subset. On the SutureBot 5-stage suturing benchmark with $n=20$ trials, it is the only evaluated model to achieve full end-to-end completion, with $5/20$ trials or $25\%$, versus $0\%$ for all baselines. It also achieves $64\%$ average success across a 29-step ex vivo wound closure sequence and improves multi-embodiment transfer success by $20$–$30\%$ absolute over the base model.

The second model, **Cosmos-H-Surgical-Simulator**, is the first action-conditioned world model enabling multi-embodiment surgical simulation from a single checkpoint. Fine-tuned from Cosmos-Predict 2.5, it predicts $T=12$ future frames conditioned on kinematic actions, trains on 32 datasets from 9 platforms, and is evaluated with
$$
\mathrm{L_1}(x,\hat x)=\frac1{HW}\sum_{i,j}|x_{i,j}-\hat x_{i,j}|,\qquad \mathrm{SSIM}(x,\hat x)\in[0,1].
$$
The applications demonstrated are in silico replay of recorded actions for closed-loop policy evaluation and generation of synthetic surgical videos spanning multiple robots.

This medical instantiation changes the scale and institutional scope of Open-H-Embodiment. Rather than deriving care from intrinsic drives, it supplies infrastructure for pretraining, world modeling, and simulation in a domain where previous datasets were small, single-embodiment, and rarely shared openly [2604.21017].

## 6. Scope, adjacent frameworks, and recurring points of confusion

A recurring point of confusion is to treat Open-H-Embodiment as if it referred to one standardized object. The literature does not support that reduction. In one usage it is a reinforcement-learning framework built from being-in-the-world, being-towards-death, homeostasis, and empowerment [2510.07117]. In another it is an open holistic VLA stack with camera and URDF priors [2602.12062]. In another it is an empirical recipe for whole-body humanoid loco-manipulation [2606.22174]. In medical robotics it is a large-scale open dataset and the foundation models it enables [2604.21017].

A second point of confusion is to equate openness with mere code release. The surveyed works use openness in a broader technical sense: open-source hardware-software unification, open cross-embodiment training, open datasets, open teleoperation and deployment infrastructure, and open model checkpoints. "OpenEAI-Platform" [2606.03392], for example, presents a fully open-source platform integrating a low-cost 6+1 degree-of-freedom robotic arm and a reproducible VLA model trained only on open-source robot and multimodal datasets, with full hardware designs, drivers, models, and training/data pipelines. Although it does not define Open-H-Embodiment as a formal doctrine, it exemplifies the same movement toward reproducible embodied AI stacks.

A third point of confusion is to assume that embodiment can be abstracted away once a sufficiently large model is available. Several papers argue against that assumption in different ways. HoloBrain-0 explicitly incorporates multi-view camera parameters and URDF kinematic descriptions to enhance 3D spatial reasoning and support diverse embodiments [2602.12062]. OpenHLM reports that a joint-based whole-body teleoperation interface outperforms alternatives that only partially expose the humanoid’s degrees of freedom [2606.22174]. H-Zero emphasizes transformation mapping utilities, embodiment descriptors, and a unified control interface for transfer across robots [2512.00971]. The consistent theme is that morphology and kinematics remain first-class variables.

A broader neighboring ecosystem extends the same logic beyond autonomous control. "OPENJ" [2605.04270] defines a conceptual framework for open-source digital human modeling and ergonomic assessment in CAD, with a 41-DOF URDF-based mannequin, optimization-based IK, differential IK, and scripted implementations of RULA, REBA, the NIOSH Lifting Equation, and OWAS. "Embodied Manipulation with Past and Future Morphologies through an Open Parametric Hand Design" [2410.18633] presents an open parametric hand design governed by approximately 56 parameters, with non-linear rolling joints, anatomical tendon routing, and variable stiffness over nearly two orders of magnitude. A plausible implication is that the wider open-embodiment research landscape includes not only policies and datasets but also morphology generators, ergonomic simulators, and fabrication pipelines.

Across these strands, Open-H-Embodiment functions as a unifying research orientation: embodiment is modeled explicitly; heterogeneity of morphology, sensing, and kinematics is embraced rather than normalized away; and open infrastructures are treated as prerequisites for scalable progress. The specific meaning, however, remains domain-dependent, ranging from existentially grounded intrinsic motivation to reproducible multi-robot foundation-model ecosystems [2510.07117; 2602.12062; 2604.21017; 2606.22174].

Source: https://www.emergentmind.com/topics/open-h-embodiment