Scaling Behavior Foundation Model for Humanoid Robots
Abstract: Humanoid control requires natural whole-body coordination, precise real-time responses to control signals, and robust generalization across diverse environmental contexts, making it a cornerstone for generalist embodied agents. Behavior Foundation Models (BFMs) have recently emerged as a promising solution to address these challenges by leveraging large-scale behavioral data to achieve superior expressiveness, versatility and generalization. However, despite growing interest in scaling BFMs to further improve their capabilities, it remains unclear how key factors, including the learning paradigm, behavioral data and model architecture should be coordinated to enable effective scaling. In this work, we revisit the scaling recipe for BFMs and demonstrate that substantial performance gains can be achieved through the coordination of three core components: 1) the learning paradigm of motion tracking that reformulates diverse humanoid control problems as the reproduction of integrated whole-body behaviors in the global frame; 2) the strategic synergy between on-policy rollout quantity and reference motion diversity; and 3) the expressive and scalable model architecture termed Humanoid Transformer that facilitates the natural emergence of structured behavioral representations. Through extensive experiments in both simulation and real-world deployment, we demonstrate that our approach yields significant improvements in control fidelity and task generalization, reducing Mean Per-Keypoint Position Error (MPKPE) on the test set by over 10% in local mode and 82% in global mode compared with existing humanoid controllers. These results establish BFM as a principled and effective foundation for scalable and general-purpose humanoid control.
First 10 authors:
Paper Prompts
Sign up for free to create and run prompts on this paper.
Top Community Prompts
Explain it Like I'm 14
Overview
This paper is about teaching humanoid robots (robots shaped like people) to move naturally and solve many different tasks using one powerful, general “brain.” The authors build a Behavior Foundation Model (BFM) that learns from a huge amount of human movement data and from its own practice in simulation. They also study how to “scale up” this kind of model so it keeps getting better as you give it more data, more practice, and a smarter neural network design.
What are the main goals?
The paper tries to answer three simple questions:
- How should we set up the learning problem so one model can learn many whole‑body skills (like walking, reaching, and manipulating objects) without being tied to one specific task?
- What kind of training data matters most: more practice time, more variety in the motions, or both?
- What model design helps the robot naturally learn a clean “language” of behaviors it can reuse across tasks?
How did they do it?
The authors combine three key ideas. Think of it like teaching a student dancer:
1) A unified way to learn: track full motions in the global world
- The robot learns by “motion tracking,” which means watching a reference motion (like a dance or a walk) and trying to copy it.
- Importantly, it copies the whole motion in the world, not just the pose of the limbs. That includes where the robot’s body moves over the ground (the “root” motion). Why this matters: Walking forward and marching in place can look similar if you ignore where you move. By including global movement, the robot learns the real intention of the motion.
- The same model can handle “global” control (follow a path in the world) and “local” control (move relative to your current position) by switching how targets are defined.
2) The right kind of data: quantity and diversity work together
- Quantity: The robot learns with trial-and-error (a method called PPO). Each “on-policy rollout” is like a practice session using the robot’s current skills. Scaling quantity means collecting more of these sessions by:
- Running more simulations in parallel (more GPUs).
- Letting each simulation run longer (longer rollout horizon).
- Diversity: They also feed the robot a huge, varied library of human motions (over 102 million frames from many datasets: walking, dancing, manipulating objects, etc.). These are retargeted to the robot’s body so it can imitate them.
- They add smart tricks:
- Early termination and quick reset near where it failed, so it doesn’t waste time practicing badly off-track.
- Adaptive sampling, which focuses training more on the motions the robot finds hardest, while still keeping variety.
3) A stronger model: the “Humanoid Transformer”
- Instead of a simple MLP (a basic neural network), they use a Transformer tailored to robot control.
- It looks at short windows of recent body states and actions (context), plus future motion goals, and learns to predict the next action.
- The goals are fed through special layers (with RMSNorm) that help form a tidy, continuous “latent space” of intentions—like an internal map of behaviors the robot can reuse.
- The actor (which picks actions) and critic (which judges how good things are going) share the backbone design but see different sets of future frames to balance quick reactions and longer-term planning.
What did they find?
Here are the main results, explained simply:
- Scaling works—when done right. Increasing both the number of parallel simulations (width) and how long each simulation runs (depth) clearly improves performance. Doing only one or the other helps less consistently; the best gains come from growing both together.
- More varied motions help the robot generalize. A bigger, more diverse motion library teaches the robot a wider set of natural, whole‑body skills.
- The new Transformer architecture leads to better control and cleaner internal behavior representations than typical MLPs.
- Big accuracy gains. On standard tests, their BFM reduced joint-position errors a lot compared to strong baselines:
- Over 10% lower error in local mode.
- About 82% lower error in global mode.
- In plain words: the robot’s body parts end up much closer to where they’re supposed to be.
- It works in simulation and on a real humanoid robot, showing the method transfers beyond a computer setting.
Note on metrics: “MPKPE” is like the average distance between where each important body point should be and where it actually ended up. Lower is better.
Why does this matter?
- One model, many skills: This approach is a step toward a single, general controller that can handle walking, balancing, using hands, and coordinating the whole body—without rewriting a new controller for every task.
- Better scaling recipe: The paper offers a clear guide for making humanoid behavior models stronger:
- Learn by tracking full motions in the world (not just poses).
- Scale both practice time (on-policy rollouts) and motion diversity.
- Use an expressive, scalable architecture (the Humanoid Transformer) that naturally organizes behavior intentions.
- Broader impact: Stronger, general-purpose humanoid control could speed up real-world robots that help in homes, hospitals, warehouses, and disaster zones—anywhere human-shaped movement helps.
In short, the paper shows how to systematically scale data, practice, and model design so humanoid robots learn natural, reusable behaviors that transfer across many tasks.
Knowledge Gaps
Knowledge Gaps, Limitations, and Open Questions
Below is a single, actionable list of what remains missing, uncertain, or unexplored in the paper.
- Lack of a formal analysis explaining why global-frame tracking reduces behavioral ambiguity and improves learning (e.g., credit assignment, optimization landscape) compared to local/root-decoupled variants.
- No quantitative study of how global-frame tracking affects invariances (translation/rotation) and whether it harms transfer to local control modes or downstream tasks that only need relative motion.
- The synergy between on-policy rollout quantity (width/depth) and reference motion diversity is asserted but not characterized: no metrics of diversity, no causal analysis, and no scaling laws linking performance to these two axes independently.
- The width vs depth trade-off in on-policy scaling remains unresolved; guidance on how to allocate compute between parallelism and horizon for best sample efficiency is absent.
- Reliance on PPO is justified by engineering maturity, but comparative evidence vs off-policy (e.g., SAC/TD3), model-based, or actor-critic variants at scale is missing; sample efficiency and stability under larger scales are unquantified.
- Adaptive sampling introduces non-stationarity; there is no convergence analysis, ablation on hyperparameters (β, w_min, w_max, T_eval), or safeguards against catastrophic forgetting of rare skills.
- Early termination (0.5 m global deviation) and RSI may bias the data distribution; the sensitivity of performance to the termination threshold and initialization policy is not studied.
- No ablation on reward weights or components; robustness of results to reward shaping choices and whether automatic reward tuning could further improve scaling is unknown.
- The large human-motion corpus is treated as diverse, but diversity is not measured (e.g., coverage of contact-rich skills, crawls, fast agility, load carrying); underrepresented skills and their impact on generalization remain unidentified.
- Retargeting via frame-wise IK lacks dynamics/contact consistency; effects of foot sliding, contact mismatch, or force infeasible poses on learning quality and final control robustness are not quantified.
- Object-interaction motions (GRAB/OMOMO) are retargeted without explicit object dynamics; how much object-contact information is lost and how it limits loco-manipulation remains unclear.
- The control interface is restricted to masked whole-body Cartesian targets with eight predefined masks; automatic discovery of masks, continuous partial specifications, and conflict resolution between overlapping intents are unexplored.
- Generalization to other command modalities (language, vision, force/EMG) via the same interface is not demonstrated; procedures to map multimodal inputs into the goal space are unspecified.
- The proposed Humanoid Transformer’s scalability envelope on embedded hardware (latency, memory, real-time FPS) is not reported; trade-offs between token budget, temporal window size, and control delay are unquantified.
- No ablation isolating the benefit of cross-attention conditioning vs simple concatenation, or comparing Transformers against large MLPs, RNNs, and state space models at equal compute.
- The RMSNorm-induced latent hypersphere is claimed to encourage structure, but there is no quantitative probing (linearity, disentanglement, smoothness, controllability) or evaluation of downstream editability/steerability.
- Delay compensation via a stochastically sampled future offset lacks a principled design; sensitivity to sensor/actuation latency, timestamp jitter, and time-warping is not analyzed.
- The critic uses privileged information; the impact of removing or constraining privilege (for real-world-only training) on performance and stability is unknown.
- Domain randomization is limited in scope (friction, CoM, hand mass, velocity perturbations); robustness to broader disturbances (large external pushes, uneven terrain, stairs, slopes, compliance, payload variations) is not evaluated.
- Only a single morphology (Unitree G1) is studied; cross-morphology transfer, scaling to different link lengths/masses/joint limits, or multi-embodiment pretraining is not examined.
- Action space is limited to PD joint targets; how performance changes under torque control, impedance/admittance control, or hybrid force–position control (especially for manipulation) is untested.
- Evaluation focuses on tracking metrics (succ/G-MPKPE/L-MPKPE/rotations); task-level outcomes (manipulation success, locomotion over obstacles, energy efficiency, stability margins, comfort/human-likeness) are not reported.
- Global-mode improvements are highlighted, but sensitivity to root localization drift/noise in real deployment and failure modes under degraded global pose estimates are not analyzed.
- The sim-to-real pipeline is deferred to the appendix; quantitative transfer rates, calibration procedures, sensor noise models, latency budgets, and failure analysis on hardware are missing from the main text.
- Training compute, wall-clock, and carbon/energy cost vs performance are not disclosed; no empirical scaling laws (e.g., returns vs FLOPs/data) or diminishing-returns diagnostics are provided.
- Failure cases are not broken down by skill category or motion attributes (speed, amplitude, contact complexity); per-skill performance diagnostics that would guide dataset curation are absent.
- The effect of engine mismatch (training in IsaacLab, testing in MuJoCo) on learned dynamics, contact behaviors, and evaluation fairness is not quantified.
- Safety, stability guarantees, and recovery behaviors (fall detection/recovery, safe shutdown) are not modeled; constraints-based RL or certified control integration is unexplored.
- No evaluation of multi-agent or human–robot interaction behaviors (e.g., social navigation, co-manipulation), which are central to humanoid deployment.
- Open question: can a BFM be trained to discover behaviors without any reference motions (unsupervised RL/skill discovery) while still supporting a promptable control interface?
- Open question: how to incorporate scene/object state and constraints directly into the goal space to enable context-aware loco-manipulation with closed-loop perception.
- Open question: how to measure and enforce dataset diversity in a principled way (e.g., coverage metrics on pose/velocity/contact manifolds) and design curricula that balance exploration with coverage.
- Open question: what theoretical guarantees (if any) can be established for generalization across behavior specifications, environments, and body morphologies under the proposed scaling recipe.
Practical Applications
Immediate Applications
The following applications can be piloted or deployed using the paper’s current methods (motion-tracking BFM with global/local control, Humanoid Transformer backbone, large-scale retargeted motion corpus, PPO with on-policy scaling, latency compensation, and domain randomization).
- Humanoid whole-body controller for R&D and prototyping (Robotics, Software)
- Use the pretrained BFM as a drop-in whole-body controller for Unitree G1–class robots to rapidly test locomotion, dexterous manipulation, and loco-manipulation in simulation and on hardware.
- Potential tools/workflows: BFM Control API (masked whole-body target poses), IsaacLab/MuJoCo sim harnesses, PD-controlled low-level execution, global/local control modes, MPKPE-based acceptance tests.
- Assumptions/dependencies: Accurate state estimation; tuned PD gains; IsaacLab/MuJoCo parity; compute for training (up to 64 GPUs); availability of diverse reference motions.
- Teleoperation via local control and behavior inpainting (Robotics, Entertainment)
- Map human motion (e.g., Xsens, BVH) to the humanoid in local mode for low-latency teleop with natural, whole-body coordination; masked targets can inpaint unspecified links.
- Potential tools/workflows: Retargeting pipeline, stochastic future offset for latency compensation, safety bounds using global error thresholds.
- Assumptions/dependencies: Reliable MoCap; network QoS; operator safety framework; coverage of target behaviors in the motion corpus.
- Warehouse/factory pilots for navigation and simple pick/place (Logistics, Manufacturing)
- Execute planner-specified base trajectories in global mode while controlling arms via masked targets for box carrying, handoff, or staging.
- Potential tools/workflows: Motion-centric task authoring (reference or partial targets), global MPKPE-based monitoring, ROS 2/Isaac bridge.
- Assumptions/dependencies: External perception and grasping stack; suitable end-effectors; floor friction variability; safety compliance.
- Robust sim-to-real evaluation framework (Academia, Industry)
- Adopt the paper’s global-frame tracking, early termination, and success criteria to standardize evaluation across sims (IsaacLab→MuJoCo) and hardware.
- Potential tools/workflows: BONES-style test sets; G-/L-MPKPE and rotation errors; disturbance tests via domain randomization.
- Assumptions/dependencies: Community acceptance of metrics; alignment on thresholds.
- Motion retargeting for animation and digital human labs (Media, XR, Academia)
- Use the two-stage retargeting and RSI to convert heterogeneous BVH/SMPL motion libraries to humanoids for previsualization, XR experiences, or training datasets.
- Potential tools/workflows: Automated skeleton alignment, IK-based framewise retargeting, adaptive sampling to prioritize hard sequences.
- Assumptions/dependencies: Data usage rights; quality of source motions; morphology differences.
- Curriculum and scaling-law experiments in embodied AI (Academia)
- Reproduce and extend the paper’s findings on the synergy between on-policy quantity (envs×horizon) and reference diversity; study architecture scaling of the Humanoid Transformer vs MLPs.
- Potential tools/workflows: On-policy collection scaling scripts; adaptive sampling scheduler; ablations for mask sets and reward terms.
- Assumptions/dependencies: Multi-GPU access; large test suites; reproducible seeds.
- Safety and QA test harness for humanoids (Policy, Industry)
- Apply global-frame error bounds and success thresholds as objective acceptance tests for behavior fidelity and recovery.
- Potential tools/workflows: Automated acceptance pipelines; “red lines” for link deviation; disturbance-injection tests.
- Assumptions/dependencies: Legal/safety frameworks; standardized lab environments.
- ROS 2 real-time control integration (Software, Robotics)
- Wrap the BFM control interface in ROS 2 for command multiplexing (planners, teleop, scripted partial targets) and logging of tracking metrics.
- Potential tools/workflows: State estimator bridge, real-time scheduler, tokenized goal injection nodelets.
- Assumptions/dependencies: Deterministic timing; on-board compute or edge inference server.
- Educational labs on general-purpose humanoid control (Education)
- Hands-on coursework using the BFM to illustrate whole-body coordination, goal-conditioned RL, and sim-to-real techniques.
- Potential tools/workflows: Prebuilt scenarios (walking, reaching, loco-manip), mask-mode “what-if” experiments.
- Assumptions/dependencies: Access to simulation GPUs; safe demo hardware.
- Benchmark curation and sharing of diverse motion corpora (Academia, Data)
- Aggregate open datasets (AMASS, GRAB, LAFAN, etc.) with documented diversity metrics to drive cross-lab comparability in BFM training.
- Potential tools/workflows: Dataset metadata tools; coverage/difficulty scoring; standardized retargeting configs.
- Assumptions/dependencies: Dataset licensing and redistribution rights.
Long-Term Applications
These require additional research, scaling, integration with perception/manipulation stacks, safety certification, and/or broader ecosystem maturation.
- General-purpose household humanoids with multi-modal prompting (Robotics, Consumer)
- Control daily tasks via language/gesture to latent-behavior prompts that the BFM executes in global/local modes with whole-body coordination.
- Potential tools/products: “Behavior Prompting” SDK; task planners that output masked targets; home-safe controllers.
- Assumptions/dependencies: Robust perception, grasp planning, failure recovery, and certification for in-home operation.
- Human-aware industrial co-bots in shared spaces (Manufacturing, Logistics)
- Safe, natural whole-body motions around people (e.g., handovers, co-carrying) leveraging the BFM’s coordinated loco-manipulation.
- Potential tools/workflows: Proximity and compliance controllers on top of BFM; risk-aware global error bounds.
- Assumptions/dependencies: Tactile/force sensing; legal frameworks; contact-rich safety guarantees beyond PD control.
- Hospital and retail service robots (Healthcare, Retail)
- Aisle navigation and shelf restocking or delivery with planners feeding global trajectories and arm targets to the BFM.
- Potential tools/workflows: Hospital/retail scene priors; reliability monitors; fallback policies.
- Assumptions/dependencies: Robustness to clutter, crowds, and narrow aisles; hygienic/end-effector constraints.
- Foundation controller marketplace across morphologies (Software, Robotics)
- Package and fine-tune BFMs for different humanoids using the paper’s retargeting and latent-space transfer without per-task reward engineering.
- Potential tools/products: Model Zoo with behavior coverage indices; adapter layers for morphology gaps.
- Assumptions/dependencies: Consistent kinematic descriptions; scalable retargeting; IP licensing.
- Autonomous mobile manipulation with task planners (Robotics, Software)
- High-level planners output temporally indexed whole-body targets; the BFM provides robust execution and inpainting for underspecified links.
- Potential tools/workflows: Temporal goal tokenization APIs; plan-to-mask compilers.
- Assumptions/dependencies: Reliable long-horizon planning; perception for constraints; closed-loop correction.
- Standardization and certification suites for humanoids (Policy, Standards)
- Industry-wide benchmarks based on G-/L-MPKPE, success rates, and disturbance recovery across canonical control modes.
- Potential tools/workflows: Third-party test beds; reporting templates; conformance levels.
- Assumptions/dependencies: Standards bodies’ engagement; consensus on thresholds and scenarios.
- Healthcare rehab and eldercare assist (Healthcare)
- Transfer BFM principles to exoskeleton/humanoid assist, enabling natural, patient-aligned whole-body support.
- Potential tools/workflows: Safety-critical controller variants; clinician-in-the-loop teleop.
- Assumptions/dependencies: Medical device certification; redundancy and fail-safes; biomechanical personalization.
- Telepresence and remote operations at scale (Enterprise, Infrastructure)
- Long-duration telepresence workers operating humanoids with latency-aware control and behavior inpainting.
- Potential tools/products: Network QoS-aware future-offset tuning; ergonomic operator interfaces.
- Assumptions/dependencies: Reliable high-bandwidth networks; ergonomic input devices; liability frameworks.
- Continuous learning from web-scale motion libraries (Data, Software)
- Periodic retraining/fine-tuning as larger, more diverse motion datasets emerge; adaptive sampling to focus on gaps.
- Potential tools/workflows: Automated retargeting farms; diversity scoring; data governance.
- Assumptions/dependencies: Data rights and privacy; compute/energy budgets; reproducibility.
- Contact-rich dexterity and tactile-informed behaviors (Robotics)
- Extend BFM to integrate tactile/force signals for fine manipulation and robust contact transitions.
- Potential tools/workflows: Multimodal tokenizers (proprioception + tactile); reward shaping that respects contact physics.
- Assumptions/dependencies: High-fidelity sensors; improved simulators; safety under unexpected contacts.
- Energy-/compute-efficient deployment via model compression (Software, Embedded)
- Distill or quantize the Humanoid Transformer for on-board, low-latency inference without server dependence.
- Potential tools/workflows: Policy distillation pipelines; mixed-precision runtimes; scheduler co-design.
- Assumptions/dependencies: Maintain fidelity post-compression; embedded GPU/TPU availability.
- Automated sim-to-real pipelines (Robotics, Tools)
- Self-tuning domain randomization and automatic acceptance testing to continuously push new behaviors to fleets.
- Potential tools/workflows: DR parameter search; hardware-in-the-loop validation; roll-back safeguards.
- Assumptions/dependencies: High-fidelity physics; telemetry; CI/CD for robots.
Cross-cutting assumptions and risks
- Coverage of behaviors in the reference motion corpus strongly affects generalization; rare or safety-critical behaviors require curated data.
- Sim-to-real success depends on actuator quality, latency, sensing, contact modeling, and PD tuning; heavy domain randomization helps but is not sufficient for all tasks.
- Training scale is compute-intensive; reproducibility and cost may limit adoption without shared checkpoints.
- The current method does not integrate perception or task planning; end-to-end autonomy needs additional stacks.
- Data licensing and governance for motion datasets may constrain commercial use.
Glossary
- Adaptive sampling: A training data strategy that increases the frequency of difficult examples while maintaining coverage of the dataset. Example: "Adaptive sampling intends to bias the data distribution toward challenging motion sequences"
- Asymmetric actor-critic: An RL design where the actor and critic receive different observations (e.g., the critic gets privileged information) to stabilize learning. Example: "We adopt the asymmetric actor-critic design in PPO"
- Behavior Foundation Models (BFMs): Large pretrained humanoid controllers that learn general behaviors and can be conditioned by diverse specifications. Example: "Behavior Foundation Models (BFMs) have recently emerged as a promising solution"
- Classifier guidance: A technique to steer diffusion models toward desired outputs using an auxiliary classifier signal. Example: "it applies classifier guidance~\cite{dhariwal2021diffusion} to the diffusion-based BFM"
- Cross-attention: An attention mechanism that conditions one sequence (e.g., queries) on another sequence (e.g., goals). Example: "tokenized, and injected into the backbone through cross-attention"
- DAgger: A dataset aggregation framework for imitation learning that iteratively queries an expert to label states visited by the learned policy. Example: "within the DAgger framework~\cite{ross2011reduction}"
- Discount factor: The scalar in RL that exponentially downweights future rewards relative to immediate rewards. Example: "and the discount factor"
- Domain randomization: Training-time randomization of environment and physical parameters to improve real-world robustness. Example: "we apply domain randomization during training."
- Forward-backward representations: A representation learning approach that models dynamics in both forward and backward temporal directions. Example: "based on forward-backward representations~\cite{touati2021learning,tirinzoni2025zero}"
- Goal-conditioned reinforcement learning (GCRL): RL where policies are conditioned on goal states specifying desired outcomes. Example: "we structure the pretraining of BFMs as a goal-conditioned reinforcement learning (GCRL) problem,"
- Global frame: A world-fixed coordinate frame used to express absolute positions and orientations. Example: "in the global frame"
- Heading frame: A body-centric coordinate frame aligned with the humanoid’s heading direction. Example: "All these whole-body states of the critic are expressed in humanoid's heading frame."
- Humanoid Transformer: The proposed transformer-based backbone tailored for scalable humanoid behavior learning. Example: "We finally introduce the Humanoid Transformer, an expressive and scalable architecture"
- Inverse kinematics: The process of computing joint configurations that realize desired end-effector or link poses. Example: "by solving an inverse kinematics problem"
- IsaacLab: A physics simulation platform used for training robotic policies. Example: "All models are trained in IsaacLab~\cite{mittal2025isaac}"
- Latent space: A learned, usually lower-dimensional space that encodes behavioral intentions or features. Example: "within a coherent latent space."
- Loco-manipulation: Coordinated behaviors that combine locomotion and manipulation with the whole body. Example: "whole-body coordinated loco-manipulation"
- Markov Decision Process (MDP): A formal RL framework defined by states, actions, transitions, rewards, and a discount factor. Example: "We formulate humanoid control as a Markov Decision Process (MDP)"
- Mean Per-Keypoint Position Error (MPKPE): An evaluation metric averaging positional errors over body keypoints. Example: "reducing Mean Per-Keypoint Position Error (MPKPE) on the test set by over 10\% in local mode and 82\% in global mode compared with existing humanoid controllers."
- Modality-specific tokenizers: Separate encoders that convert inputs of different types (e.g., proprioception, goals) into token embeddings. Example: "These inputs are then encoded by modality-specific tokenizers"
- Motion retargeting: Adapting human motion data to a humanoid robot’s morphology and kinematics. Example: "Motion Retargeting. We adopt a two-stage retargeting pipeline"
- Motion tracking: A learning paradigm that imitates reference motions for whole-body behavior acquisition. Example: "Our framework identifies motion tracking~\cite{peng2018deepmimic} as a unified and scalable paradigm"
- MuJoCo: A physics engine commonly used for robotic control and evaluation. Example: "and evaluated in MuJoCo~\cite{todorov2012mujoco}"
- On-policy data collection: Gathering training data from trajectories generated by the current policy during interaction. Example: "quantity is obtained by scaling on-policy data collection"
- On-policy rollouts: Trajectories sampled by the current policy used for policy updates in on-policy RL. Example: "effective training data are the on-policy rollouts collected through environment interaction."
- Proportional-Derivative (PD) controller: A low-level controller that tracks desired joint angles using proportional and derivative terms. Example: "executed by a low-level proportional-derivative (PD) controller"
- Proximal Policy Optimization (PPO): A popular on-policy RL algorithm that stabilizes updates via clipped objectives. Example: "We instantiate motion tracking with Proximal Policy Optimization (PPO)~\cite{schulman2017proximal}"
- Proprioceptive state: Internal sensory state of the robot, such as joint angles and velocities. Example: "the agent’s proprioceptive state"
- Query token: A learnable token in transformers used to aggregate context for prediction via attention. Example: "followed by a learnable query token for action or value prediction."
- Reference State Initialization (RSI): Resetting episode initial states from reference motion frames to keep rollouts on-manifold. Example: "reference state initialization (RSI)~\cite{peng2018deepmimic}"
- Reward engineering: Manually designing reward functions tailored to specific tasks. Example: "often relying on extensive reward engineering tailored to individual tasks and contextual settings."
- RMSNorm: A normalization technique that scales activations by their root-mean-square. Example: "we employ RMSNorm~\cite{zhang2019root} to normalize the goal embeddings"
- Rollout horizon: The number of timesteps collected per environment per policy update. Example: "the rollout horizon"
- Root localization: Estimating or specifying the robot’s root (base) pose in global coordinates for control. Example: "support both global control with root localization and local control"
- Root-relative Cartesian space: A coordinate system expressing targets relative to the robot’s root frame. Example: "based on masked whole-body target poses in the root-relative Cartesian space."
- Self-attention: An attention mechanism where tokens attend to each other within the same sequence. Example: "During self-attention, the query token is prevented from being attended to by the context tokens"
- Success rate (Succ): The fraction of motion sequences tracked without violating predefined error thresholds. Example: "we report the success rate (Succ)"
- Value estimation: Predicting the expected return from a state, used by the critic to guide learning. Example: "to provide more accurate value estimation."