Papers
Topics
Authors
Recent
Search
2000 character limit reached

Scaling Behavior Foundation Model for Humanoid Robots

Published 16 Jul 2026 in cs.RO and cs.AI | (2607.15163v1)

Abstract: Humanoid control requires natural whole-body coordination, precise real-time responses to control signals, and robust generalization across diverse environmental contexts, making it a cornerstone for generalist embodied agents. Behavior Foundation Models (BFMs) have recently emerged as a promising solution to address these challenges by leveraging large-scale behavioral data to achieve superior expressiveness, versatility and generalization. However, despite growing interest in scaling BFMs to further improve their capabilities, it remains unclear how key factors, including the learning paradigm, behavioral data and model architecture should be coordinated to enable effective scaling. In this work, we revisit the scaling recipe for BFMs and demonstrate that substantial performance gains can be achieved through the coordination of three core components: 1) the learning paradigm of motion tracking that reformulates diverse humanoid control problems as the reproduction of integrated whole-body behaviors in the global frame; 2) the strategic synergy between on-policy rollout quantity and reference motion diversity; and 3) the expressive and scalable model architecture termed Humanoid Transformer that facilitates the natural emergence of structured behavioral representations. Through extensive experiments in both simulation and real-world deployment, we demonstrate that our approach yields significant improvements in control fidelity and task generalization, reducing Mean Per-Keypoint Position Error (MPKPE) on the test set by over 10% in local mode and 82% in global mode compared with existing humanoid controllers. These results establish BFM as a principled and effective foundation for scalable and general-purpose humanoid control.

Summary

  • The paper presents a coordinated scaling recipe combining global-frame motion tracking, behaviorally diverse reference data, and Transformer architecture to improve humanoid whole-body control.
  • The proposed 3M-parameter Transformer achieves 0.9677 success on the BONES benchmark and 0.9776 on cross-source tests, with global-frame G-MPKPE values of 0.0798 m and 0.0915 m, respectively.
  • The results show that data diversity matters more than raw motion count, while moderate latent-space robustness, 50 Hz onboard inference, and eight masked-pose control modes support practical deployment on the Unitree G1.

This paper presents a systematic study of the scaling behavior of Behavior Foundation Models (BFMs) for humanoid whole-body control, identifying a coordinated recipe across three axes: the learning paradigm, the composition of training data, and the model architecture. The work is grounded in the observation that prior BFM efforts—such as distillation-based approaches like BFM4Humanoid and BeyondMimic, unsupervised RL methods like BFM-Zero, and motion-tracking-based pretraining as in SONIC—have explored scaling only in fragmented ways, without clarifying how these factors interact. The authors' central contribution is to make this coupling explicit and to demonstrate that coordinating it yields large gains in tracking fidelity and cross-domain generalization on the Unitree G1 platform.

Unified learning paradigm: global-frame motion tracking

The paper formulates BFM pretraining as goal-conditioned RL in which behavior is defined as a trajectory of proprioceptive states and actions, excluding goal states, which are treated as external specifications. Motion tracking serves purely as a proxy task for behavior learning rather than as the deployment-time control objective, which is what distinguishes a BFM from a conventional motion tracker. Training uses PPO with asymmetric actor-critic observations, where the critic receives privileged simulator state.

The most consequential design choice is the reward formulation: unlike BeyondMimic or SONIC, which either drop root-position tracking or decouple root following from pose tracking, the proposed model must reproduce reference motions as integrated whole-body trajectories in the global frame. The authors argue that removing root-position tracking makes behaviors with distinct global semantics (e.g., walking forward versus marching in place) nearly indistinguishable in the learning signal, while decoupling root and pose objectives compromises their coordination. This claim is supported by an ablation baseline (BFM-Bym) trained identically but with the BeyondMimic reward: the global-frame reward consistently reduces G-MPKPE relative to the ablation, indicating more coherent behavioral guidance. Because control signals are re-anchorable to the robot's current root state, the same pretrained model supports both global control (with root localization) and local control at deployment.

The control interface consists of masked whole-body target poses in root-relative Cartesian space, sampled from eight curated control modes spanning granularities from root-only to full 14-link whole-body specification. Unspecified links are naturally inpainted, allowing sparse commands such as end-effector targets to induce plausible whole-body behavior.

Data scaling: quantity versus diversity

A key conceptual clarification is that under PPO, the effective training data are the on-policy rollouts, whose quantity is governed by environment parallelism and rollout horizon—not by the number of reference motions, which instead shapes the behavioral distribution. Scaling experiments vary both dimensions across three levels (32/48/64 GPUs; rollout horizons 32/48/64), with 8192 environments per GPU. Jointly scaling width and depth produces consistent improvements, with the largest configuration best in nearly all settings, but scaling either dimension alone does not reliably help—suggesting an unresolved interplay between how much experience is collected and how it is accumulated per update.

For reference-motion scaling, the 102M-frame corpus (aggregated from LAFAN, AMASS, OMOMO, GRAB, SnapMoGen, FineDance, BONES-SEED, and Embody3D) is partitioned into five nested subsets. A K-Means occupancy analysis over a 64-dimensional behavioral feature space reveals two regimes: homogeneous scaling (XXS→S, all within BONES-SEED) leaves cluster occupancy essentially flat (~0.935–0.937), whereas heterogeneous scaling (S→L, adding OMOMO, GRAB, dance data, Embody3D) raises occupancy from 0.9365 to 0.9995. The empirical results align sharply with this analysis: homogeneous scaling yields only marginal gains even on the source-aligned BONES test set, while heterogeneous scaling produces little benefit on BONES but substantial gains on the cross-source test set (Xsens captures plus 100Style). The implication is direct: gains from more reference motions materialize only when scaling measurably expands behavioral coverage relevant to the target distribution—an important corrective to the common practice of equating dataset size with capability. The authors acknowledge that this coverage metric depends on the clustering support being defined by the full corpus, an assumption they flag as inherently untestable against truly long-tail behaviors.

Supporting infrastructure includes a two-stage retargeting pipeline (skeleton alignment via SMPL shape or BVH offset optimization, then sequential frame-wise IK), termination at 0.5 m global deviation, reference state initialization at the termination point, and clamped adaptive sampling (β=0.999\beta=0.999, weights clipped to [0.03,1.0][0.03, 1.0]).

Architecture: the Humanoid Transformer

The proposed Humanoid Transformer tokenizes temporal windows of proprioception and actions into an interleaved context sequence processed by self-attention with RoPE, injects future goal tokens through cross-attention, and predicts actions or values from a learnable query token that attends over—but is not attended to by—the context. RMSNorm projects goal embeddings onto a hypersphere, inducing a structured latent space shaped solely by the tracking objective, without auxiliary regularization losses. The actor uses five consecutive future frames plus one stochastically sampled offset in [5,32][5,32] (dynamically adjustable at deployment to absorb latency); the critic uses fixed exponentially spaced offsets {0,1,2,4,8,16,32}\{0,1,2,4,8,16,32\}.

Scaling experiments compare MLP backbones (3.05M and 11.86M parameters) against Transformer variants from 0.41M to 9.91M parameters. The medium Transformer (3M parameters) matches or exceeds the substantially larger MLP, and further capacity growth yields diminishing returns—evidence that architectural expressiveness, not raw parameter count, drives effective capacity utilization here. However, scaling does not uniformly improve all control modes; some saturate early, which the authors attribute to inter-mode learning-difficulty imbalance and optimization trade-offs within the shared latent space.

Latent-space structure

Qualitative visualization shows that latent trajectories for individual motions are locally smooth on the unit hypersphere, and globally organized: opposing intentions (forward/backward walking, left/right crouching) occupy separated regions preserving directional relationships. Robustness is quantified by rotating latent vectors along random directions: success rate degrades only modestly up to 20° perturbations (0.9717→0.9646 on BONES; 0.9836→0.9400 on Ours), supporting the claim that behavioral interpretation tolerates moderate latent noise. Notably, increasing model capacity drives convergence of latent representations across different control modes, offering a mechanistic explanation for the observed trade-offs among modes during architecture scaling—that improved alignment in one mode can come at slight cost to others sharing the latent space.

Benchmark results

Against off-the-shelf controllers GMT, TWIST, and SONIC, plus the BFM-Bym ablation, the 3M-parameter model achieves strong margins. On the source-aligned BONES test set it reaches a 0.9677 success rate with G-MPKPE of 0.0798 m, versus 0.9239 for SONIC (whose training set may overlap this benchmark, a caveat the authors note) and below 0.45 for GMT and TWIST. Generalization gaps are larger on the cross-source test set: GMT and TWIST collapse to success rates near 0.08, SONIC reaches 0.5937, while the proposed model attains 0.9776 with G-MPKPE of 0.0915 m. Relative to existing humanoid controllers overall, the paper reports MPKPE reductions exceeding 10% in local mode and 82% in global mode. Real-world deployment runs onboard inference at 50 Hz via TensorRT atop a 200 Hz PD loop, with timestamp-adjusted future-frame indexing compensating communication latency, and supports online switching among all eight control modes.

Limitations and open questions

The authors are explicit about several constraints. First, whether the eight-mode masked-pose interface is the right abstraction—and how it should integrate with future high-level policies—remains open. Second, the scaling study is limited relative to LLM-scale investigations: distributed training infrastructure for humanoid pretraining is fragile, and onboard compute caps model size if 50 Hz inference must be preserved alongside high-level policy headroom. Third, the behavioral-coverage metric presupposes that the full corpus provides adequate clustering support, an assumption that cannot be verified against unobserved long-tail behaviors. Fourth, the mechanism behind width–depth interactions in on-policy data collection, and the mode-coupling dynamics in the shared latent space, are described but not fully characterized.

Conclusion

This paper reframes BFM scaling as a coordination problem among learning paradigm, data quantity–diversity synergy, and architecture, rather than a matter of enlarging any single factor. Its principal empirical findings—that global-frame integrated tracking yields more coherent guidance than decoupled rewards, that reference-motion scaling pays off only when it expands measured behavioral coverage on relevant distributions, and that a moderately sized Transformer outperforms larger MLPs while inducing structured, robust latents without auxiliary objectives—together constitute a concrete, reproducible recipe for general-purpose humanoid foundation controllers.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Explain it Like I'm 14

Overview

This paper is about teaching humanoid robots (robots shaped like people) to move naturally and solve many different tasks using one powerful, general “brain.” The authors build a Behavior Foundation Model (BFM) that learns from a huge amount of human movement data and from its own practice in simulation. They also study how to “scale up” this kind of model so it keeps getting better as you give it more data, more practice, and a smarter neural network design.

What are the main goals?

The paper tries to answer three simple questions:

  • How should we set up the learning problem so one model can learn many whole‑body skills (like walking, reaching, and manipulating objects) without being tied to one specific task?
  • What kind of training data matters most: more practice time, more variety in the motions, or both?
  • What model design helps the robot naturally learn a clean “language” of behaviors it can reuse across tasks?

How did they do it?

The authors combine three key ideas. Think of it like teaching a student dancer:

1) A unified way to learn: track full motions in the global world

  • The robot learns by “motion tracking,” which means watching a reference motion (like a dance or a walk) and trying to copy it.
  • Importantly, it copies the whole motion in the world, not just the pose of the limbs. That includes where the robot’s body moves over the ground (the “root” motion). Why this matters: Walking forward and marching in place can look similar if you ignore where you move. By including global movement, the robot learns the real intention of the motion.
  • The same model can handle “global” control (follow a path in the world) and “local” control (move relative to your current position) by switching how targets are defined.

2) The right kind of data: quantity and diversity work together

  • Quantity: The robot learns with trial-and-error (a method called PPO). Each “on-policy rollout” is like a practice session using the robot’s current skills. Scaling quantity means collecting more of these sessions by:
    • Running more simulations in parallel (more GPUs).
    • Letting each simulation run longer (longer rollout horizon).
  • Diversity: They also feed the robot a huge, varied library of human motions (over 102 million frames from many datasets: walking, dancing, manipulating objects, etc.). These are retargeted to the robot’s body so it can imitate them.
  • They add smart tricks:
    • Early termination and quick reset near where it failed, so it doesn’t waste time practicing badly off-track.
    • Adaptive sampling, which focuses training more on the motions the robot finds hardest, while still keeping variety.

3) A stronger model: the “Humanoid Transformer”

  • Instead of a simple MLP (a basic neural network), they use a Transformer tailored to robot control.
  • It looks at short windows of recent body states and actions (context), plus future motion goals, and learns to predict the next action.
  • The goals are fed through special layers (with RMSNorm) that help form a tidy, continuous “latent space” of intentions—like an internal map of behaviors the robot can reuse.
  • The actor (which picks actions) and critic (which judges how good things are going) share the backbone design but see different sets of future frames to balance quick reactions and longer-term planning.

What did they find?

Here are the main results, explained simply:

  • Scaling works—when done right. Increasing both the number of parallel simulations (width) and how long each simulation runs (depth) clearly improves performance. Doing only one or the other helps less consistently; the best gains come from growing both together.
  • More varied motions help the robot generalize. A bigger, more diverse motion library teaches the robot a wider set of natural, whole‑body skills.
  • The new Transformer architecture leads to better control and cleaner internal behavior representations than typical MLPs.
  • Big accuracy gains. On standard tests, their BFM reduced joint-position errors a lot compared to strong baselines:
    • Over 10% lower error in local mode.
    • About 82% lower error in global mode.
    • In plain words: the robot’s body parts end up much closer to where they’re supposed to be.
  • It works in simulation and on a real humanoid robot, showing the method transfers beyond a computer setting.

Note on metrics: “MPKPE” is like the average distance between where each important body point should be and where it actually ended up. Lower is better.

Why does this matter?

  • One model, many skills: This approach is a step toward a single, general controller that can handle walking, balancing, using hands, and coordinating the whole body—without rewriting a new controller for every task.
  • Better scaling recipe: The paper offers a clear guide for making humanoid behavior models stronger:
    • Learn by tracking full motions in the world (not just poses).
    • Scale both practice time (on-policy rollouts) and motion diversity.
    • Use an expressive, scalable architecture (the Humanoid Transformer) that naturally organizes behavior intentions.
  • Broader impact: Stronger, general-purpose humanoid control could speed up real-world robots that help in homes, hospitals, warehouses, and disaster zones—anywhere human-shaped movement helps.

In short, the paper shows how to systematically scale data, practice, and model design so humanoid robots learn natural, reusable behaviors that transfer across many tasks.

Knowledge Gaps

Knowledge Gaps, Limitations, and Open Questions

Below is a single, actionable list of what remains missing, uncertain, or unexplored in the paper.

  • Lack of a formal analysis explaining why global-frame tracking reduces behavioral ambiguity and improves learning (e.g., credit assignment, optimization landscape) compared to local/root-decoupled variants.
  • No quantitative study of how global-frame tracking affects invariances (translation/rotation) and whether it harms transfer to local control modes or downstream tasks that only need relative motion.
  • The synergy between on-policy rollout quantity (width/depth) and reference motion diversity is asserted but not characterized: no metrics of diversity, no causal analysis, and no scaling laws linking performance to these two axes independently.
  • The width vs depth trade-off in on-policy scaling remains unresolved; guidance on how to allocate compute between parallelism and horizon for best sample efficiency is absent.
  • Reliance on PPO is justified by engineering maturity, but comparative evidence vs off-policy (e.g., SAC/TD3), model-based, or actor-critic variants at scale is missing; sample efficiency and stability under larger scales are unquantified.
  • Adaptive sampling introduces non-stationarity; there is no convergence analysis, ablation on hyperparameters (β, w_min, w_max, T_eval), or safeguards against catastrophic forgetting of rare skills.
  • Early termination (0.5 m global deviation) and RSI may bias the data distribution; the sensitivity of performance to the termination threshold and initialization policy is not studied.
  • No ablation on reward weights or components; robustness of results to reward shaping choices and whether automatic reward tuning could further improve scaling is unknown.
  • The large human-motion corpus is treated as diverse, but diversity is not measured (e.g., coverage of contact-rich skills, crawls, fast agility, load carrying); underrepresented skills and their impact on generalization remain unidentified.
  • Retargeting via frame-wise IK lacks dynamics/contact consistency; effects of foot sliding, contact mismatch, or force infeasible poses on learning quality and final control robustness are not quantified.
  • Object-interaction motions (GRAB/OMOMO) are retargeted without explicit object dynamics; how much object-contact information is lost and how it limits loco-manipulation remains unclear.
  • The control interface is restricted to masked whole-body Cartesian targets with eight predefined masks; automatic discovery of masks, continuous partial specifications, and conflict resolution between overlapping intents are unexplored.
  • Generalization to other command modalities (language, vision, force/EMG) via the same interface is not demonstrated; procedures to map multimodal inputs into the goal space are unspecified.
  • The proposed Humanoid Transformer’s scalability envelope on embedded hardware (latency, memory, real-time FPS) is not reported; trade-offs between token budget, temporal window size, and control delay are unquantified.
  • No ablation isolating the benefit of cross-attention conditioning vs simple concatenation, or comparing Transformers against large MLPs, RNNs, and state space models at equal compute.
  • The RMSNorm-induced latent hypersphere is claimed to encourage structure, but there is no quantitative probing (linearity, disentanglement, smoothness, controllability) or evaluation of downstream editability/steerability.
  • Delay compensation via a stochastically sampled future offset lacks a principled design; sensitivity to sensor/actuation latency, timestamp jitter, and time-warping is not analyzed.
  • The critic uses privileged information; the impact of removing or constraining privilege (for real-world-only training) on performance and stability is unknown.
  • Domain randomization is limited in scope (friction, CoM, hand mass, velocity perturbations); robustness to broader disturbances (large external pushes, uneven terrain, stairs, slopes, compliance, payload variations) is not evaluated.
  • Only a single morphology (Unitree G1) is studied; cross-morphology transfer, scaling to different link lengths/masses/joint limits, or multi-embodiment pretraining is not examined.
  • Action space is limited to PD joint targets; how performance changes under torque control, impedance/admittance control, or hybrid force–position control (especially for manipulation) is untested.
  • Evaluation focuses on tracking metrics (succ/G-MPKPE/L-MPKPE/rotations); task-level outcomes (manipulation success, locomotion over obstacles, energy efficiency, stability margins, comfort/human-likeness) are not reported.
  • Global-mode improvements are highlighted, but sensitivity to root localization drift/noise in real deployment and failure modes under degraded global pose estimates are not analyzed.
  • The sim-to-real pipeline is deferred to the appendix; quantitative transfer rates, calibration procedures, sensor noise models, latency budgets, and failure analysis on hardware are missing from the main text.
  • Training compute, wall-clock, and carbon/energy cost vs performance are not disclosed; no empirical scaling laws (e.g., returns vs FLOPs/data) or diminishing-returns diagnostics are provided.
  • Failure cases are not broken down by skill category or motion attributes (speed, amplitude, contact complexity); per-skill performance diagnostics that would guide dataset curation are absent.
  • The effect of engine mismatch (training in IsaacLab, testing in MuJoCo) on learned dynamics, contact behaviors, and evaluation fairness is not quantified.
  • Safety, stability guarantees, and recovery behaviors (fall detection/recovery, safe shutdown) are not modeled; constraints-based RL or certified control integration is unexplored.
  • No evaluation of multi-agent or human–robot interaction behaviors (e.g., social navigation, co-manipulation), which are central to humanoid deployment.
  • Open question: can a BFM be trained to discover behaviors without any reference motions (unsupervised RL/skill discovery) while still supporting a promptable control interface?
  • Open question: how to incorporate scene/object state and constraints directly into the goal space to enable context-aware loco-manipulation with closed-loop perception.
  • Open question: how to measure and enforce dataset diversity in a principled way (e.g., coverage metrics on pose/velocity/contact manifolds) and design curricula that balance exploration with coverage.
  • Open question: what theoretical guarantees (if any) can be established for generalization across behavior specifications, environments, and body morphologies under the proposed scaling recipe.

Practical Applications

Immediate Applications

The following applications can be piloted or deployed using the paper’s current methods (motion-tracking BFM with global/local control, Humanoid Transformer backbone, large-scale retargeted motion corpus, PPO with on-policy scaling, latency compensation, and domain randomization).

  • Humanoid whole-body controller for R&D and prototyping (Robotics, Software)
    • Use the pretrained BFM as a drop-in whole-body controller for Unitree G1–class robots to rapidly test locomotion, dexterous manipulation, and loco-manipulation in simulation and on hardware.
    • Potential tools/workflows: BFM Control API (masked whole-body target poses), IsaacLab/MuJoCo sim harnesses, PD-controlled low-level execution, global/local control modes, MPKPE-based acceptance tests.
    • Assumptions/dependencies: Accurate state estimation; tuned PD gains; IsaacLab/MuJoCo parity; compute for training (up to 64 GPUs); availability of diverse reference motions.
  • Teleoperation via local control and behavior inpainting (Robotics, Entertainment)
    • Map human motion (e.g., Xsens, BVH) to the humanoid in local mode for low-latency teleop with natural, whole-body coordination; masked targets can inpaint unspecified links.
    • Potential tools/workflows: Retargeting pipeline, stochastic future offset for latency compensation, safety bounds using global error thresholds.
    • Assumptions/dependencies: Reliable MoCap; network QoS; operator safety framework; coverage of target behaviors in the motion corpus.
  • Warehouse/factory pilots for navigation and simple pick/place (Logistics, Manufacturing)
    • Execute planner-specified base trajectories in global mode while controlling arms via masked targets for box carrying, handoff, or staging.
    • Potential tools/workflows: Motion-centric task authoring (reference or partial targets), global MPKPE-based monitoring, ROS 2/Isaac bridge.
    • Assumptions/dependencies: External perception and grasping stack; suitable end-effectors; floor friction variability; safety compliance.
  • Robust sim-to-real evaluation framework (Academia, Industry)
    • Adopt the paper’s global-frame tracking, early termination, and success criteria to standardize evaluation across sims (IsaacLab→MuJoCo) and hardware.
    • Potential tools/workflows: BONES-style test sets; G-/L-MPKPE and rotation errors; disturbance tests via domain randomization.
    • Assumptions/dependencies: Community acceptance of metrics; alignment on thresholds.
  • Motion retargeting for animation and digital human labs (Media, XR, Academia)
    • Use the two-stage retargeting and RSI to convert heterogeneous BVH/SMPL motion libraries to humanoids for previsualization, XR experiences, or training datasets.
    • Potential tools/workflows: Automated skeleton alignment, IK-based framewise retargeting, adaptive sampling to prioritize hard sequences.
    • Assumptions/dependencies: Data usage rights; quality of source motions; morphology differences.
  • Curriculum and scaling-law experiments in embodied AI (Academia)
    • Reproduce and extend the paper’s findings on the synergy between on-policy quantity (envs×horizon) and reference diversity; study architecture scaling of the Humanoid Transformer vs MLPs.
    • Potential tools/workflows: On-policy collection scaling scripts; adaptive sampling scheduler; ablations for mask sets and reward terms.
    • Assumptions/dependencies: Multi-GPU access; large test suites; reproducible seeds.
  • Safety and QA test harness for humanoids (Policy, Industry)
    • Apply global-frame error bounds and success thresholds as objective acceptance tests for behavior fidelity and recovery.
    • Potential tools/workflows: Automated acceptance pipelines; “red lines” for link deviation; disturbance-injection tests.
    • Assumptions/dependencies: Legal/safety frameworks; standardized lab environments.
  • ROS 2 real-time control integration (Software, Robotics)
    • Wrap the BFM control interface in ROS 2 for command multiplexing (planners, teleop, scripted partial targets) and logging of tracking metrics.
    • Potential tools/workflows: State estimator bridge, real-time scheduler, tokenized goal injection nodelets.
    • Assumptions/dependencies: Deterministic timing; on-board compute or edge inference server.
  • Educational labs on general-purpose humanoid control (Education)
    • Hands-on coursework using the BFM to illustrate whole-body coordination, goal-conditioned RL, and sim-to-real techniques.
    • Potential tools/workflows: Prebuilt scenarios (walking, reaching, loco-manip), mask-mode “what-if” experiments.
    • Assumptions/dependencies: Access to simulation GPUs; safe demo hardware.
  • Benchmark curation and sharing of diverse motion corpora (Academia, Data)
    • Aggregate open datasets (AMASS, GRAB, LAFAN, etc.) with documented diversity metrics to drive cross-lab comparability in BFM training.
    • Potential tools/workflows: Dataset metadata tools; coverage/difficulty scoring; standardized retargeting configs.
    • Assumptions/dependencies: Dataset licensing and redistribution rights.

Long-Term Applications

These require additional research, scaling, integration with perception/manipulation stacks, safety certification, and/or broader ecosystem maturation.

  • General-purpose household humanoids with multi-modal prompting (Robotics, Consumer)
    • Control daily tasks via language/gesture to latent-behavior prompts that the BFM executes in global/local modes with whole-body coordination.
    • Potential tools/products: “Behavior Prompting” SDK; task planners that output masked targets; home-safe controllers.
    • Assumptions/dependencies: Robust perception, grasp planning, failure recovery, and certification for in-home operation.
  • Human-aware industrial co-bots in shared spaces (Manufacturing, Logistics)
    • Safe, natural whole-body motions around people (e.g., handovers, co-carrying) leveraging the BFM’s coordinated loco-manipulation.
    • Potential tools/workflows: Proximity and compliance controllers on top of BFM; risk-aware global error bounds.
    • Assumptions/dependencies: Tactile/force sensing; legal frameworks; contact-rich safety guarantees beyond PD control.
  • Hospital and retail service robots (Healthcare, Retail)
    • Aisle navigation and shelf restocking or delivery with planners feeding global trajectories and arm targets to the BFM.
    • Potential tools/workflows: Hospital/retail scene priors; reliability monitors; fallback policies.
    • Assumptions/dependencies: Robustness to clutter, crowds, and narrow aisles; hygienic/end-effector constraints.
  • Foundation controller marketplace across morphologies (Software, Robotics)
    • Package and fine-tune BFMs for different humanoids using the paper’s retargeting and latent-space transfer without per-task reward engineering.
    • Potential tools/products: Model Zoo with behavior coverage indices; adapter layers for morphology gaps.
    • Assumptions/dependencies: Consistent kinematic descriptions; scalable retargeting; IP licensing.
  • Autonomous mobile manipulation with task planners (Robotics, Software)
    • High-level planners output temporally indexed whole-body targets; the BFM provides robust execution and inpainting for underspecified links.
    • Potential tools/workflows: Temporal goal tokenization APIs; plan-to-mask compilers.
    • Assumptions/dependencies: Reliable long-horizon planning; perception for constraints; closed-loop correction.
  • Standardization and certification suites for humanoids (Policy, Standards)
    • Industry-wide benchmarks based on G-/L-MPKPE, success rates, and disturbance recovery across canonical control modes.
    • Potential tools/workflows: Third-party test beds; reporting templates; conformance levels.
    • Assumptions/dependencies: Standards bodies’ engagement; consensus on thresholds and scenarios.
  • Healthcare rehab and eldercare assist (Healthcare)
    • Transfer BFM principles to exoskeleton/humanoid assist, enabling natural, patient-aligned whole-body support.
    • Potential tools/workflows: Safety-critical controller variants; clinician-in-the-loop teleop.
    • Assumptions/dependencies: Medical device certification; redundancy and fail-safes; biomechanical personalization.
  • Telepresence and remote operations at scale (Enterprise, Infrastructure)
    • Long-duration telepresence workers operating humanoids with latency-aware control and behavior inpainting.
    • Potential tools/products: Network QoS-aware future-offset tuning; ergonomic operator interfaces.
    • Assumptions/dependencies: Reliable high-bandwidth networks; ergonomic input devices; liability frameworks.
  • Continuous learning from web-scale motion libraries (Data, Software)
    • Periodic retraining/fine-tuning as larger, more diverse motion datasets emerge; adaptive sampling to focus on gaps.
    • Potential tools/workflows: Automated retargeting farms; diversity scoring; data governance.
    • Assumptions/dependencies: Data rights and privacy; compute/energy budgets; reproducibility.
  • Contact-rich dexterity and tactile-informed behaviors (Robotics)
    • Extend BFM to integrate tactile/force signals for fine manipulation and robust contact transitions.
    • Potential tools/workflows: Multimodal tokenizers (proprioception + tactile); reward shaping that respects contact physics.
    • Assumptions/dependencies: High-fidelity sensors; improved simulators; safety under unexpected contacts.
  • Energy-/compute-efficient deployment via model compression (Software, Embedded)
    • Distill or quantize the Humanoid Transformer for on-board, low-latency inference without server dependence.
    • Potential tools/workflows: Policy distillation pipelines; mixed-precision runtimes; scheduler co-design.
    • Assumptions/dependencies: Maintain fidelity post-compression; embedded GPU/TPU availability.
  • Automated sim-to-real pipelines (Robotics, Tools)
    • Self-tuning domain randomization and automatic acceptance testing to continuously push new behaviors to fleets.
    • Potential tools/workflows: DR parameter search; hardware-in-the-loop validation; roll-back safeguards.
    • Assumptions/dependencies: High-fidelity physics; telemetry; CI/CD for robots.

Cross-cutting assumptions and risks

  • Coverage of behaviors in the reference motion corpus strongly affects generalization; rare or safety-critical behaviors require curated data.
  • Sim-to-real success depends on actuator quality, latency, sensing, contact modeling, and PD tuning; heavy domain randomization helps but is not sufficient for all tasks.
  • Training scale is compute-intensive; reproducibility and cost may limit adoption without shared checkpoints.
  • The current method does not integrate perception or task planning; end-to-end autonomy needs additional stacks.
  • Data licensing and governance for motion datasets may constrain commercial use.

Glossary

  • Adaptive sampling: A training data strategy that increases the frequency of difficult examples while maintaining coverage of the dataset. Example: "Adaptive sampling intends to bias the data distribution toward challenging motion sequences"
  • Asymmetric actor-critic: An RL design where the actor and critic receive different observations (e.g., the critic gets privileged information) to stabilize learning. Example: "We adopt the asymmetric actor-critic design in PPO"
  • Behavior Foundation Models (BFMs): Large pretrained humanoid controllers that learn general behaviors and can be conditioned by diverse specifications. Example: "Behavior Foundation Models (BFMs) have recently emerged as a promising solution"
  • Classifier guidance: A technique to steer diffusion models toward desired outputs using an auxiliary classifier signal. Example: "it applies classifier guidance~\cite{dhariwal2021diffusion} to the diffusion-based BFM"
  • Cross-attention: An attention mechanism that conditions one sequence (e.g., queries) on another sequence (e.g., goals). Example: "tokenized, and injected into the backbone through cross-attention"
  • DAgger: A dataset aggregation framework for imitation learning that iteratively queries an expert to label states visited by the learned policy. Example: "within the DAgger framework~\cite{ross2011reduction}"
  • Discount factor: The scalar in RL that exponentially downweights future rewards relative to immediate rewards. Example: "and γ\gamma the discount factor"
  • Domain randomization: Training-time randomization of environment and physical parameters to improve real-world robustness. Example: "we apply domain randomization during training."
  • Forward-backward representations: A representation learning approach that models dynamics in both forward and backward temporal directions. Example: "based on forward-backward representations~\cite{touati2021learning,tirinzoni2025zero}"
  • Goal-conditioned reinforcement learning (GCRL): RL where policies are conditioned on goal states specifying desired outcomes. Example: "we structure the pretraining of BFMs as a goal-conditioned reinforcement learning (GCRL) problem,"
  • Global frame: A world-fixed coordinate frame used to express absolute positions and orientations. Example: "in the global frame"
  • Heading frame: A body-centric coordinate frame aligned with the humanoid’s heading direction. Example: "All these whole-body states of the critic are expressed in humanoid's heading frame."
  • Humanoid Transformer: The proposed transformer-based backbone tailored for scalable humanoid behavior learning. Example: "We finally introduce the Humanoid Transformer, an expressive and scalable architecture"
  • Inverse kinematics: The process of computing joint configurations that realize desired end-effector or link poses. Example: "by solving an inverse kinematics problem"
  • IsaacLab: A physics simulation platform used for training robotic policies. Example: "All models are trained in IsaacLab~\cite{mittal2025isaac}"
  • Latent space: A learned, usually lower-dimensional space that encodes behavioral intentions or features. Example: "within a coherent latent space."
  • Loco-manipulation: Coordinated behaviors that combine locomotion and manipulation with the whole body. Example: "whole-body coordinated loco-manipulation"
  • Markov Decision Process (MDP): A formal RL framework defined by states, actions, transitions, rewards, and a discount factor. Example: "We formulate humanoid control as a Markov Decision Process (MDP)"
  • Mean Per-Keypoint Position Error (MPKPE): An evaluation metric averaging positional errors over body keypoints. Example: "reducing Mean Per-Keypoint Position Error (MPKPE) on the test set by over 10\% in local mode and 82\% in global mode compared with existing humanoid controllers."
  • Modality-specific tokenizers: Separate encoders that convert inputs of different types (e.g., proprioception, goals) into token embeddings. Example: "These inputs are then encoded by modality-specific tokenizers"
  • Motion retargeting: Adapting human motion data to a humanoid robot’s morphology and kinematics. Example: "Motion Retargeting. We adopt a two-stage retargeting pipeline"
  • Motion tracking: A learning paradigm that imitates reference motions for whole-body behavior acquisition. Example: "Our framework identifies motion tracking~\cite{peng2018deepmimic} as a unified and scalable paradigm"
  • MuJoCo: A physics engine commonly used for robotic control and evaluation. Example: "and evaluated in MuJoCo~\cite{todorov2012mujoco}"
  • On-policy data collection: Gathering training data from trajectories generated by the current policy during interaction. Example: "quantity is obtained by scaling on-policy data collection"
  • On-policy rollouts: Trajectories sampled by the current policy used for policy updates in on-policy RL. Example: "effective training data are the on-policy rollouts collected through environment interaction."
  • Proportional-Derivative (PD) controller: A low-level controller that tracks desired joint angles using proportional and derivative terms. Example: "executed by a low-level proportional-derivative (PD) controller"
  • Proximal Policy Optimization (PPO): A popular on-policy RL algorithm that stabilizes updates via clipped objectives. Example: "We instantiate motion tracking with Proximal Policy Optimization (PPO)~\cite{schulman2017proximal}"
  • Proprioceptive state: Internal sensory state of the robot, such as joint angles and velocities. Example: "the agent’s proprioceptive state"
  • Query token: A learnable token in transformers used to aggregate context for prediction via attention. Example: "followed by a learnable query token for action or value prediction."
  • Reference State Initialization (RSI): Resetting episode initial states from reference motion frames to keep rollouts on-manifold. Example: "reference state initialization (RSI)~\cite{peng2018deepmimic}"
  • Reward engineering: Manually designing reward functions tailored to specific tasks. Example: "often relying on extensive reward engineering tailored to individual tasks and contextual settings."
  • RMSNorm: A normalization technique that scales activations by their root-mean-square. Example: "we employ RMSNorm~\cite{zhang2019root} to normalize the goal embeddings"
  • Rollout horizon: The number of timesteps collected per environment per policy update. Example: "the rollout horizon"
  • Root localization: Estimating or specifying the robot’s root (base) pose in global coordinates for control. Example: "support both global control with root localization and local control"
  • Root-relative Cartesian space: A coordinate system expressing targets relative to the robot’s root frame. Example: "based on masked whole-body target poses in the root-relative Cartesian space."
  • Self-attention: An attention mechanism where tokens attend to each other within the same sequence. Example: "During self-attention, the query token is prevented from being attended to by the context tokens"
  • Success rate (Succ): The fraction of motion sequences tracked without violating predefined error thresholds. Example: "we report the success rate (Succ)"
  • Value estimation: Predicting the expected return from a state, used by the critic to guide learning. Example: "to provide more accurate value estimation."

Open Problems

We're still in the process of identifying open problems mentioned in this paper. Please check back in a few minutes.

Tweets

Sign up for free to view the 4 tweets with 83 likes about this paper.