Papers
Topics
Authors
Recent
Search
2000 character limit reached

GigaBrain-WBC-0.5: A Behavior World Model for Robust Whole-Body Control with Environment Interaction

Published 18 Aug 2026 in cs.RO, cs.AI, and cs.LG | (2608.18234v1)

Abstract: Whole-body motion tracking policies turn a humanoid into a robust control interface: the teleoperator---or an upstream model---only supplies a coarse movement intent, while the low-level policy keeps the robot balanced and physically feasible. Existing trackers deliver this interface only on flat ground: trained in empty scenes, they never learn how contact with terrain and objects reshapes their dynamics, and they attempt to teach the policy to balance under any command by continually enlarging the reference-motion corpus, which stops working once feasible behaviors become environment-dependent. We present GigaBrain-WBC-0.5, the first Behavior World Model (BWM) for humanoid whole-body control. Rather than a purely reactive tracker, we train a causal Transformer to jointly predict its next action, next state, and the distribution over its next latent behavior command, so the network that acts also models how the environment shapes what it can do next. An automatic terrain-annotation pipeline recovers full 3D contact geometry from retargeted motion, enabling terrain annotation at the scale of existing motion datasets. The predicted distribution is reused at deployment to detect implausible commands online and retract them onto learned behaviors, so the robot attempts tasks in a "best-effort" manner. The result is a unified policy that takes real-time command, interacts with environment, and stays robust to implausible commands, falls, and disturbances. GigaBrain-WBC-0.5 achieves the highest success rate across all four regimes among three large-scale tracker baselines: 81.3% on terrain interaction (4.3x the strongest baseline), 83.1% under implausible commands, and 99.3% fall recovery (16.8x the strongest baseline). Hardware trials show robust interaction under missing supports and disturbances; the Unitree G1 checkpoint transfers to the Maker L01 robot with simple fine-tuning.

Summary

  • The paper develops GigaBrain-WBC-0.5, a behavior world model that enables robust whole-body control for humanoid robots by learning both proprioceptive state and latent behavior commands, leading to a 22% improvement in terrain interaction performance.
  • Under the proposed framework, the policy accurately reconstructs terrain geometry from retargeted motion data at 92% audited accuracy, addressing the data challenge that previously limited interactive trackers.
  • The model demonstrates nativre fatler recovery with 99.3% success without hardware limitations needing significant retraining.

Motivation and problem setting

Large-scale whole-body motion trackers have matured into general low-level control interfaces for humanoid robots: a teleoperator or upstream policy supplies only coarse movement intent, while the tracker absorbs balance and physical feasibility. The paper identifies two structural limitations of this paradigm. First, essentially all such trackers (SONIC, HoloMotion-1, Humanoid-GPT, GMT) are trained and evaluated in empty scenes on flat ground, so the policy never learns how contact with terrain and objects reshapes its own dynamics. Second, the standard recipe for command robustness—enlarging the reference-motion corpus until the policy stays balanced under nearly any command—fails once feasibility becomes environment-conditional: a lunge feasible in free space may be impossible on a narrow step, and no unconditional data coverage covers a conditional feasible set.

The authors argue these two gaps are two faces of one problem, and address both by training the controller as what they call a Behavior World Model (BWM): rather than only reproducing an action, the network predicts its own next proprioceptive state and the distribution over its next latent behavior command, forcing it to internalize contact dynamics and to model which behaviors the current environment admits. The claimed contribution is, to the authors' knowledge, the first behavior world model for humanoid whole-body control (2608.18234).

Method

Architecture and objective

The system controls a Unitree G1 (29 DoF) at 50 Hz from proprioception st∈R67s_t \in \mathbb{R}^{67}, the previous action, and a 10-frame reference window ct∈R440c_t \in \mathbb{R}^{440} assembled from reference joint positions, root-relative rotations and translations, and gravity alignment expressed in the reference frame. An MLP encoder maps the window to a 64-dimensional latent behavior command as two 32-d tokens. A finite scalar quantizer with a kinematic decoder is attached only during training as a cycle-consistency regularizer; reconstruction is routed through this auxiliary branch rather than through the policy network, which the authors report preserves continuous control resolution while shaping a smooth, well-spread latent space—an empirical contrast with SONIC's design, where quantization sits on the path to the action head.

The core network is a 6-layer causal Transformer (4 heads, rotary embeddings, 32-frame local attention, per-layer KV cache) over [st,at−1,zt][s_t, a_{t-1}, z_t], with three linear heads emitting the action ata_t, next-state prediction s^t+1\hat{s}_{t+1}, and a 4-component diagonal Gaussian Mixture Model GtG_t over the next latent command. Training combines PPO with reconstruction, cycle-consistency, next-state regression against parallel-simulator rollouts, and a negative log-likelihood term on the mixture.

A notable input-design decision concerns the reference-root translation block, which corrects global drift but is unavailable on hardware. It is masked two-thirds of the time during training so the policy learns to run without absolute root feedback and re-anchors whenever the signal briefly reappears; at deployment it is zeroed throughout. This choice costs velocity-profile fidelity (discussed below) but keeps the controller driveable without global localization—a prerequisite for real teleoperation.

Automatic spatial terrain annotation

Because ordinary retargeted motion corpora carry no scene geometry, the paper introduces an unsupervised pipeline that reconstructs the terrain a motion must have been performed on directly from its own kinematics. Sampled points on collision geometry of contact-relevant links are replayed kinematically in MuJoCo; a state machine marks contacts via a deceleration signature (normal speed below 0.2 m/s, tangential below 0.3 m/s, preceded by deceleration above 1 m/s²), which distinguishes support events from momentarily slow limbs. Candidate points are filtered by a whole-body penetration test that provably cannot invent geometry the robot would intersect, grouped by normal-consistent DBSCAN affinity into planar clusters, fitted with oriented box primitives, and laterally expanded by penetration-safe binary search. Unlike SceneBot's 2.5D elevation maps (Chen et al., 25 Jun 2026), the output is genuine 3D geometry—chairs with clearance, table edges, stair treads—which is what allows terrain-paired training at the scale of existing datasets (~72 h of terrain-interaction motion identified across Bones-Seed, MotionMillion, and MotionDecode).

Online out-of-distribution filtering

At deployment, the mixture Gt−1G_{t-1} emitted one step earlier serves as a one-step approximation of the conditional training distribution. Following CMP's single-step reduction of the infinite-horizon constraint (Cheng et al., 8 Apr 2026), each incoming latent is checked against a Mahalanobis ellipsoid Zk={z:χk2(z)≤r2}\mathcal{Z}_k = \{z : \chi_k^2(z) \le r^2\} of the MAP-selected component k⋆k^\star. Crucially, the ordering is strict—filtering with the mixture produced at the current step would leak the command into its own admissibility test—and letting an arbitrary component serve instead of k⋆k^\star was found to be a major failure mode allowing unrelated modes to silently bypass correction. When the test fires, the command is not rejected or replaced by the component mean but radially retracted onto the ellipsoid boundary in closed form, yielding best-effort behavior that still points toward the operator's intent. The filter is stateless, O(1), well under 1 ms per step, and exposes a single runtime-tunable safety radius ct∈R440c_t \in \mathbb{R}^{440}0 trading precision against survival without retraining.

Training details

Training uses PPO in Isaac Lab under a sequence-level update that recomputes rollout segments with detached KV prefixes attached, ensuring gradient context matches acting context. Robustness ingredients include fallen-pose initialization under a curriculum up to fully prone (so recovery is a native behavior rather than an exit to a specialist get-up controller), a recovery gate shielding tracking rewards while standing, persistent wrist/torso forces emulating payloads, and extensive domain randomization. Fine-tuning the G1 checkpoint transfers tracking to the Maker L01 embodiment on a single 8-GPU node, whereas from-scratch L01 training converges slowly—evidence the world model carries embodiment-general interaction structure.

Experimental results

Evaluation is sim-to-sim in MuJoCo against SONIC, HoloMotion-1, and Humanoid-GPT under identical failure criteria, full-length rollouts (no early termination shrinking error metrics), and no privileged signals unavailable on hardware. The filter runs at its deployed radius throughout, so results reflect the shipped system rather than an ablation-friendly variant.

Regime Metric Best baseline GigaBrain-WBC-0.5
Standard (AMASS) MPKPE / SR 82.3 mm / 94.1% (SONIC) 76.6 mm / 96.3%
Terrain MPKPE / SR 283.3 mm / 18.7% 93.3 mm / 81.3%
OOD (implausible refs) MPKPE / SR 208.0 mm / 70.6% 158.0 mm / 83.1%
Fall recovery SR† / Jerk 5.9% / 1295.5 rad/s³ 99.3% / 1050.6 rad/s³

Three findings stand out. First, the world-model formulation costs nothing in free space: the policy leads all three baselines on the flat-ground regime they were specialized for, though it trails HoloMotion-1 on root velocity (211.1 vs. 121.3 mm/s), which the authors attribute honestly to intermittent root-position masking—the price of hardware-driveable commands—and not to any keypoint degradation. Second, terrain interaction separates categorically: baselines collapse to 14–19% success with MPKPE rising three- to fourfold relative to their flat-ground numbers, whereas the proposed policy's MPKPE rises only 22% (76.6 → 93.3 mm). Third, fall recovery is a qualitative rather than incremental gap: 99.3% versus single-digit baseline rates, with the lowest joint jerk of the four policies, making the behavior deployable on hardware rather than merely successful in simulation.

The annotation pipeline itself achieves 92% overall correctness in a human audit (98% for boxes/platforms, 92% stairs, 94% seats, 84% for hand-supported motions where upstream retargeting noise corrupts the contact signature). Because failures are conservative—only deletions or displacements, never spurious obstacles—and audited failures are removed before training, the audit supports the claim that unsupervised annotation is reliable enough as training supervision. Hardware trials show matched-command superiority over SONIC on sitting, climbing, lifting, and carrying tasks, plus recovery after kick-induced falls and graceful degradation when expected supports are removed mid-task. Sweeping the safety radius traces an explicit precision–survival frontier: tightening from unfiltered buys 2.9 points of OOD survival for 12.7 mm of Standard MPKPE, with diminishing returns thereafter.

Limitations and open questions

The authors are explicit about several caveats. The filter's interception of infeasible commands is high-probability rather than guaranteed: it flags inputs unlike the training experience rather than reasoning about physical risk directly, and the safety radius must be recalibrated per checkpoint and platform before being relied upon on hardware. The terrain annotation recovers only geometry a motion actually touches, so reconstructed scenes are collections of supports rather than complete environments; surfaces present but unused are invisible to the pipeline. Baselines use different corpora and retargeting pipelines, so comparisons speak to generalization and environment-aware training effects, not a data-matched ablation. Component ablations and real-robot latency measurements are deferred to future releases. Two natural extensions named in the paper are extending reconstruction to non-supporting geometry and connecting the state-prediction head to a more direct notion of physical risk.

Conclusion

GigaBrain-WBC-0.5 demonstrates that training a humanoid tracker to predict its own next state and next-behavior distribution—rather than only its next action—yields a single policy that simultaneously tracks precisely on flat ground, exploits environmental contacts, filters implausible commands online via a closed-form projection, and recovers natively from falls. The supporting infrastructure, particularly the automatic spatial terrain annotation achieving dataset-scale terrain pairing at 92% audited accuracy, addresses the data bottleneck that previously confined interactive trackers to small specialized corpora. Whether the world-model formulation scales beyond the ~72-hour identified terrain subset, and whether the Mahalanobis filter can be grounded in physical risk rather than distributional novelty, remain the questions this work leaves open.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We found no open problems mentioned in this paper.

Tweets

Sign up for free to view the 2 tweets with 34 likes about this paper.