---
title: Humanoid Hanoi Benchmark
url: https://www.emergentmind.com/topics/humanoid-hanoi
type: topic
---

# Humanoid Hanoi Benchmark

Humanoid Hanoi defines a long-horizon whole-body manipulation benchmark for locomoting humanoid robots, centered on the Physical Tower-of-Hanoi box rearrangement challenge. The system employs a task-agnostic, shared whole-body control (WBC) architecture and a skill-based hierarchical framework—enabling robust skill sequencing and compositional generalization. Evaluation is conducted both in simulation and on the Digit V3 humanoid platform, highlighting challenges in long-horizon robustness, skill chaining, and the interplay between locomotion and manipulation under evolving state and command distributions [2602.13850].

## 1. Overview of System Architecture

The framework consists of a modular skill library instantiated by four reusable loco-manipulation skills: **GoTo**, **GoTo-with-Box**, **Pickup**, and **Place**. Each skill is implemented as a two-layer LSTM policy $\pi_i$, operating at 50 Hz, with inputs comprising robot full-body proprioception $(q, \dot q)$, SE(3) box state and attributes, and skill-specific target poses. Rather than outputting direct joint torques, these skill policies produce *masked motion directives* $\mathcal{M}_t$—encapsulating subsets of desired position, velocity, and mask-encoded degrees of freedom.

A task-agnostic shared WBC receives the masked directive and the robot's current state, generating joint-level PD setpoints $\tau_t$ for a 2 kHz sub-millisecond low-level controller. At each 50 Hz WBC timestep, an optimization problem is solved to best track the masked directive while enforcing full-body dynamic consistency and contact constraints. This is formalized as a quadratic program:

\[
\tau_t^* = \underset{\tau}{\arg\min}
\left\|W_q(q_{des} - q_t)\right\|^2 +
\left\|W_{\dot q}(\dot q_{des} - \dot q_t)\right\|^2 +
\left\|W_f(f_{des} - f_t)\right\|^2
\]

subject to

\[
M(q_t)\ddot q + C(q_t, \dot q_t)\dot q_t + g(q_t)
= S^\top \tau + J_c^\top \lambda, \qquad \lambda \ge 0,
\]

where $M, C, g$ denote the robot’s inertia, Coriolis, and gravity terms, $J_c$ and $\lambda$ model contact complementarity, and $W_q, W_{\dot q}, W_f$ are tracking weights.

The system contrasts with non-shared, skill-specific low-level controllers, pursuing a unified control interface that persists across skill boundaries.

## 2. Data Aggregation and Domain Randomization

The introduction of new skills or novel skill compositions alters both state and command distributions encountered by the WBC, degrading naive reuse robustness over extended horizons. To address this, a *WBC_Coverage_Expansion* procedure is adopted: upon the addition of each new skill $\pi_i$, the current shared WBC is used to execute $\pi_i$ in $K$ randomized scenes, accumulating successful closed-loop masked-directive trajectories $\mathcal{S}_i$. The reference set $\mathcal{R}$ of directives is then augmented, and the shared WBC retrained over the expanded $\mathcal{D}[\mathcal{R}]$ distribution.

Training incorporates extensive domain randomization: e.g., body mass $\in [0.75,1.25] \times$ nominal, joint damping $\in [0.5,3.5] \times$ nominal, box mass, friction, and communication delay ($2$–$4$ ms) are all randomized per Table 1 [2602.13850]. Directive sampling involves selecting random past directives, time-shifting, random masking, and injecting noise consistent with original command distributions.

The learning objective for parameters $\theta$ is

\[
\min_\theta\; \mathbb{E}_{(q,\dot q,\mathcal{M})\sim \mathcal{D}[\mathcal{R}_i],\;
p_{\rm dyn} \sim DR}
\left[
\|W_q(q_{des}-q)\|^2 +
\|W_{\dot q}(\dot q_{des}-\dot q)\|^2 +
\|W_f(f_{des}-f)\|^2
\right],
\]

subject to sampled dynamics parameters $p_{\rm dyn}$.

This iterative aggregation and retraining sequence expands the WBC’s operational coverage, improving robustness to skill-induced distributional shifts and long-horizon execution error accumulation.

## 3. The Humanoid Hanoi Benchmark

The Humanoid Hanoi benchmark formalizes a long-horizon, multi-object stacking scenario with the following specifications:

- **Task**: Rearrangement of three rigid boxes (“disks”) of strictly increasing size among three fixed tower locations $T_1, T_2, T_3$ on the circumference of a 1.5–2.5 m radius circle. The objective is to move all disks from $T_1$ to re-establish the ordered stack at $T_3$, obeying classical Tower-of-Hanoi rules (only one box moved at a time; no larger box atop a smaller).
- **Minimal Solution**: Requires seven sequential moves, each decomposable into the four core skills (GoTo $\rightarrow$ Pickup $\rightarrow$ GoTo-with-Box $\rightarrow$ Place).
- **State and Action Spaces**: Full robot proprioception $(q_t, \dot q_t) \in \mathbb{R}^{20}$, box SE(3) poses, and tower SE(2) locations; actions correspond to masked motion directives $\mathcal{M}_t$.
- **Collision and Stacking**: Tower pegs are implemented as fixed SE(2) targets, with stacking

Source: https://www.emergentmind.com/topics/humanoid-hanoi