Papers
Topics
Authors
Recent
Search
2000 character limit reached

HANDOFF: Humanoid Agentic Task-Space Whole-Body Control via Distilled Complementary Teachers

Published 4 Jun 2026 in cs.RO, cs.AI, and cs.LG | (2606.06493v1)

Abstract: For a humanoid robot to be deployed in the real world, the choice of command space (i.e., the interface between task planning and whole-body control) is crucial. Existing whole-body controllers typically demand dense kinematic or spatial references that planners struggle to synthesize from task semantics. We instead propose a compact, explicit interface that is intuitive, general, modular, and expressive enough for diverse manipulation skills. To this end, we introduce HANDOFF, a single humanoid whole-body controller that follows this interface and is distilled via multi-teacher KL distillation under a context-conditioned gating scheme into a mixture-of-experts student from three complementary specialists: whole-body motion tracking with safety-filtered data, locomotion, and fall-recovery. On the Unitree G1, HANDOFF matches state-of-the-art velocity tracking and offers one of the largest robust manipulation workspaces. We further demonstrate hardware feasibility through multiple natural-language-driven task roll-outs, powered by a VLM-driven agentic planner with no task-specific data or controller fine-tuning.

Summary

  • The paper introduces a 10-D command interface and context-conditioned mixture-of-experts distillation that combines motion tracking, locomotion, and fall-recovery teachers into one 29-DoF G1 controller.
  • The resulting policy achieves a 0.31 m³ robust bilateral wrist workspace, 97.7% feasibility, and 0.06 m/s forward-velocity error in its strongest variant while remaining competitive with state-of-the-art tracking.
  • The planner-agnostic interface enables VLM-driven pick-and-place, transport, squatting, and bimanual hand-off tasks on an untethered humanoid without task-specific data collection or controller fine-tuning.

HANDOFF addresses a persistent bottleneck in humanoid loco-manipulation: the interface between task-level planning and whole-body control. The authors observe that state-of-the-art whole-body controllers (WBCs) overwhelmingly operate in a motion-tracking regime, requiring the planner to emit dense full-body kinematic streams at controller rate — a requirement that forces planners into data replay of curated demonstration libraries and makes every new skill a retargeting and filtering exercise. HANDOFF instead proposes a compact, explicit 10-D command vector,

ct=[vx, vy, ωz, z, pLP, pRP],c_t = [v_x,\ v_y,\ \omega_z,\ z,\ p_L^P,\ p_R^P],

comprising planar base velocity, yaw rate, commanded root height, and bilateral pelvis-frame wrist targets. Each component maps onto an existing planner family (locomotion stacks, grasp planners, squat heuristics), and end-effector actions from VLAs map directly onto the wrist-target slots without retargeting or fine-tuning.

Method

The central technical contribution is a context-conditioned multi-teacher distillation scheme that composes three independently trained specialists into one deployable student on the 29-DoF Unitree G1:

  • Whole-body motion-tracking teacher: trained with asymmetric actor-critic PPO on retargeted human motion clips (BONES-SEED dataset), providing posture, reach, squat, and bilateral coordination priors. Retargeted clips contain dynamically infeasible squat frames, which are corrected by a closed-form control barrier function (CBF) projection on the static center-of-pressure margin within a 7-D joint-correction subspace (bilateral hip pitch, ankle pitch/roll, waist pitch); the same barrier powers a deployment-time velocity-space filter.
  • Locomotion teacher: a 15-DoF body-slice policy trained on flat terrain with curriculum-blended arm perturbations to tolerate arm-induced CoM shifts during distillation.
  • Fall-recovery teacher: a 29-DoF Adversarial Motion Prior (AMP) policy trained on locomotion plus paired fall-and-recovery sequences, with up to 40% of environments spawned in delayed fallen states.

The student is a soft mixture-of-experts (MoE) head with one expert per teacher, gated over a shared 64-D proprioception latent. Distillation supervision is split by action slice and conditioned on a regime signal xt=(∥ctvel∥,recovert)\mathbf{x}_t = (\|c_t^{\mathrm{vel}}\|, \mathrm{recover}_t): the body slice follows a sigmoid-gated convex blend of WBC and locomotion KL (transition at 0.1 m/s commanded velocity), the arm slice is anchored to the WBC teacher throughout, and the AMP teacher assumes full-action supervision under a binary recovery mask. Soft routing plus load-balancing and recovery-routing losses keep expert usage well-structured; KL coefficients are cosine-annealed (λB\lambda_B: 0.4→0.2, λA\lambda_A: 0.1→0.05, λAMP\lambda_{\mathrm{AMP}}: 0.4→0.2). Notably, the arm coefficient is deliberately smaller than the body coefficient so the body teachers retain authority over locomotion stability. An optional stability reward stack (CoM-in-support-polygon, LIPM capture point, ankle/hip/step strategy hierarchy, momentum-change penalties) can be shared across teachers and student.

The framework is extensible by design: a new specialist requires only one additional expert head and one gated-KL term, leaving existing teachers and the command interface untouched. This contrasts with GMT, which gates over clusters of a single motion-tracking manifold under dense references, and TeleGate, which retains experts and trains inference-time gating rather than collapsing them into one student.

Evaluation

Experiments run in mjlab/MuJoCo on the Unitree G1 with an external Jetson Thor + Dex1-1 gripper payload included in domain randomization. Two axes are quantified: mean absolute velocity-tracking error over [−1,1][-1,1] per-axis sweeps, and a robust bilateral wrist workspace defined as hull volume times feasibility fraction in the forward half-space, where feasibility requires both wrists within 15 cm of target, no fall, and pelvis drift under 25 cm.

Ablations show each specialist is necessary. The motion-tracking teacher alone ("Direct") is weakest on every axis (∣Δvx∣=0.29|\Delta v_x| = 0.29 m/s, robust WS 0.20 m³). Adding the standalone locomotion teacher produces the largest single improvement (∣Δvx∣|\Delta v_x| drops to 0.14), randomized commands close the lateral gap, and split-KL + MoE reaches ∣Δvx∣=0.07|\Delta v_x| = 0.07. Stability rewards push robust workspace to 0.31 m³ while preserving tracking.

SOTA comparison uses an adapted interface: baselines lacking native wrist-target inputs are equipped with a differential-IK head (mink) mapping (pLP,pRP)(p_L^P, p_R^P) to arm joints while freezing non-arm DoFs — a favorable adapter, since differential IK is precise whenever a solution exists. Even so, HANDOFF variants deliver the largest robust workspace (0.31 m³ vs. SONIC's 0.26 and AMO's 0.22), with Ours+Stab. achieving the best feasibility at 97.7%, and Ours+Stab.+Rec. attaining the best xt=(∥ctvel∥,recovert)\mathbf{x}_t = (\|c_t^{\mathrm{vel}}\|, \mathrm{recover}_t)0 among the paper's variants (0.06 m/s). Velocity tracking sits within the SOTA cluster on all three axes, though SONIC retains lower errors on xt=(∥ctvel∥,recovert)\mathbf{x}_t = (\|c_t^{\mathrm{vel}}\|, \mathrm{recover}_t)1 (0.03) and xt=(∥ctvel∥,recovert)\mathbf{x}_t = (\|c_t^{\mathrm{vel}}\|, \mathrm{recover}_t)2 (0.02) — the paper does not claim uniform superiority in tracking, only parity within the SOTA cluster combined with superior reachable workspace.

Agentic deployment demonstrates the interface's planner-agnosticism. A VLM-driven stack decomposes natural-language instructions into atomic tasks, projects 2D detections onto RGB-D point clouds for pelvis-frame waypoints, and emits the 10-D stream at rates from 0.001 Hz (instruction) down to 50 Hz (controller), tracked at 500 Hz on hardware. On a fully untethered G1 (single 140 W powerbank), the same controller executes pick-and-place, pick-transport-place, squat-pick, bimanual hand-off, and bilateral pick-and-place, plus simulated task continuation after fall recovery — with zero task-specific data collection or controller fine-tuning. The recovery continuation is only possible because the recovery specialist was distilled into the same policy rather than switched in externally.

Limitations

The paper is candid about several constraints. The interface exposes 3-D wrist positions rather than full 6-D gripper poses; orientation residuals are handled by a runtime kinematic correction, and direct 6-D tracking remains open. Hardware perception relies on a single fixed-pose head-mounted RGB-D camera, restricting sensing to the forward field of view. The specialist set omits terrain adaptation, contact-rich skills, and heavy-load handling, though the architecture is designed to admit such teachers incrementally. Additionally, the SOTA comparison depends on the differential-IK adapter for baselines, which, while described as favorable, introduces a translation layer whose interaction with each baseline's internal assumptions is not exhaustively characterized. Whether the 0.1 m/s gating threshold and its 0.02 width generalize across embodiments is also unexamined.

Conclusion

HANDOFF argues that the command space, not controller capacity alone, determines whether a WBC is usable by real planners, and substantiates this with a distilled MoE student that reconciles expressive motion priors, reliable velocity tracking, and fall recovery under a single 10-D interface. Its quantitative results — SOTA-cluster velocity tracking, the largest reported robust manipulation workspace at 0.31 m³, 97.7% feasibility, and untethered multi-task hardware rollouts driven by a VLM planner without any task-specific retraining — support the claim that compact explicit interfaces and multi-specialist distillation are compatible. The open questions it leaves are concrete: extending the interface to 6-D end-effector poses, broadening perception beyond a fixed forward camera, and validating the context-gating scheme across additional specialists and embodiments.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 1 tweet with 0 likes about this paper.