Papers
Topics
Authors
Recent
Search
2000 character limit reached

Athena-WBC: Capability-Aligned Policy Experts for Long-Tail Humanoid Whole-Body Control

Published 6 Jul 2026 in cs.RO | (2607.04837v2)

Abstract: Large-scale humanoid motion-tracking controllers are commonly improved by reallocating training effort: difficult motions are sampled more often, isolated into smaller subsets, or assigned to specialized experts. We show that this view is incomplete. In strong whole-body-control baselines, a residual set of feasible training clips remains unsolved even under targeted training, especially for high-dynamic transitions and balance-critical motions. These failures arise not only from insufficient exposure, but from a mismatch between the motion demands and the effective capability induced by the default training recipe. We propose Athena-WBC, a compact teacher-student pipeline with capability-aligned policy experts for long-tail humanoid whole-body control. Dynamic experts use a tracking-focused, constraint-aware objective that removes conservative effort and temporal-control penalties while preserving physical feasibility constraints; balance experts use a gravity curriculum to improve early-training survivability. The resulting privileged teachers are motion-routed for DAgger distillation and then compressed into a single controller with deployable observations followed by RL fine-tuning. Experiments on a full-size humanoid show improved recovery of training-set long-tail motions and better held-out tracking than a strong SONIC-recipe baseline, using only a small number of experts.

Summary

  • The paper introduces Athena-WBC, a teacher-student pipeline that recovers long-tail failures in humanoid whole-body control by aligning policy capabilities.
  • The method integrates dynamic and balance experts with adaptive motion sampling to handle high-dynamic and balance-critical motion challenges.
  • The approach enhances tracking accuracy and control smoothness, validated on an 80kg humanoid using advanced metrics and RL fine-tuning.

Capability-Aligned Policy Experts for Long-Tail Humanoid Whole-Body Control

Motivation and Problem Statement

The pursuit of robust humanoid whole-body control (WBC) has advanced significantly, leveraging vast human motion corpora for policy training. While modern motion-conditioned RL pipelines attain high aggregate tracking performance, persistent long-tail failures remain—motions within the training corpus that are unsolved by controllers despite targeted retraining efforts. These unsolved clips, particularly in high-dynamic and balance-critical regimes, expose a fundamental capability bottleneck: they defy acquisition not merely due to under-sampling, but because of a systematic mismatch between task demands and the inductive biases of the standard training recipe.

The paper proposes Athena-WBC, a teacher-student pipeline designed to explicitly address the bottleneck by aligning policy acquisition capabilities to the residual failure modes of training data. This approach transcends mere data allocation or curriculum modifications, targeting fundamental changes in policy inductive bias and training objectives.

Figure 1

Figure 1: Overview of Athena-WBC. The pipeline mines residual failures, trains dynamic and balance experts in parallel, routes frozen teachers per motion, distills expert behaviors into a unified student, then fine-tunes with RL.

Method: Compact Capability-Aligned Expert Pipeline

Residual Failure Mining and Expert Training

The pipeline commences with a general privileged teacher trained on the full motion set. Residual failures (Rgen\mathcal{R}_{\mathrm{gen}}) are identified via rollout-based evaluation, defining a subset of clips that remain unsolved despite nominal training. Importantly, these are not dominated by artifacts or clear physical infeasibility, but expose limitations inherent to the acquisition protocol.

Two capability-aligned experts are trained on Rgen\mathcal{R}_{\mathrm{gen}}:

  • Dynamic Expert: Trained with tracking rewards and physical-constraint penalties, but explicitly removing effort and temporal-control penalties. This encourages acquisition of aggressive, high-momentum motions without conservative bias. Smoothness is enforced via auxiliary policy regularization (Grad-CAPS), allowing structured action changes.
  • Balance Expert: Employs a gravity curriculum that relaxes gravity early in training to facilitate survival and stabilization in balance-critical clips. Normal gravity is restored after initial learning phases.

Adaptive motion sampling allocates rollouts based on temporal-bin difficulty scores derived from smoothed tracking errors, ensuring efficient progressive refinement.

Figure 2

Figure 2: Adaptive motion sampling transforms rollout errors into temporal-bin difficulty scores, yielding dynamic sampling probabilities for clips.

Motion-Routed Teacher Selection and Distillation

Post-training, each motion is empirically assigned to its best-performing frozen teacher via rollout-based routing. DAgger-style distillation transfers teacher behavior to a student policy using deployable observations while preserving representation structure.

RL Fine-Tuning

The distilled student undergoes PPO fine-tuning with a critic warm-up phase. Gaussian action noise covers the student neighborhood, enabling closed-loop tracking improvement and hardware deployment quality enhancements. RL fine-tuning does not further increase training-set coverage, but substantially improves held-out tracking and overall control smoothness.

Evaluation: Metrics and Empirical Findings

Athena-WBC is evaluated on a full-size 80 kg humanoid with planetary-roller-screw actuation, across diverse datasets (AMASS, Bones-Seed, BEAT, and curated mocap). The regime is high-coverage; both training and held-out distributions are scrutinized for residual failures. The evaluation employs standard metrics (SR, MPJPE), but further introduces threshold-robust Success--Tolerance Curves (STC), Threshold-Integrated Success (TIS), and Motion-Salience Weighted MPJPE (MPJPE-W) for fine-grained diagnostic insight.

Figure 3

Figure 3: SONIC baseline leaves residual failures concentrated in high-dynamic and balance-critical regimes, not explained by exposure alone.

Long-Tail Recovery

Athena-WBC's capability-aligned experts recover structured long-tail motions unsolved by SONIC-Base and exposure-only variants. Removing reward-level regularization (no-smoothness) improves tracking but dramatically degrades action smoothness; Grad-CAPS and balance expert policies recover both tracking and deployable control, evidenced by superior MPJPE, MPJPE-W, and action rate metrics.

Tracking–Smoothness Trade-off

Ablation reveals the locus of smoothness loss in reward design: conservative penalties suppress feasible high-dynamic behaviors. Auxiliary policy regularization (CAPS, Grad-CAPS) recovers smooth action trajectories, as shown by frequency-domain spectral analysis and time-domain rollouts.

Figure 4

Figure 4: High-frequency action jitter is broadly increased by removing reward-level smoothness; Grad-CAPS suppresses jitter closest to the reward-smooth baseline.

Figure 5

Figure 5: Action power spectral densities—NoSmooth variant concentrates energy in the high-frequency band, whereas CAPS and Grad-CAPS suppress rapid oscillations.

Figure 6

Figure 6: Time-series from successful backward-running clip; Grad-CAPS and CAPS reduce action-rate and jerk spikes under successful tracking.

Qualitative Case Studies

Capability-aligned policies correct failures in distinctive long-tail motions: crouch-and-walk, high-kick, single-leg recovery, and one-leg stretching. SONIC-Base fails in phases with critical support or rapid transitions, where capability experts maintain dynamic robustness and balance through tailored acquisition.

Figure 7

Figure 7

Figure 7: SONIC-Base vs. capability-aligned policy rollouts on crouch-and-walk-forward motion—the latter tracks the reference more faithfully during critical phases.

Figure 8

Figure 8

Figure 8

Figure 8: Successful single-leg pose maintained by capability expert.

Figure 9

Figure 9

Figure 9

Figure 9

Figure 9

Figure 9: Fast left-hand waving accurately tracked using motion-salience weighting.

Figure 10

Figure 10

Figure 10: Robust tracking on AMASS-eval held-out set, demonstrating improved generalization.

Implications and Future Directions

Athena-WBC demonstrates that residual training-set failures in humanoid WBC reflect fundamental capability bottlenecks, not mere exposure deficits. Aligning acquisition recipes to motion demands (policy reward decomposition, curriculum interventions, auxiliary regularization) is essential for absorbing the long-tail.

Threshold-integrated evaluative tools (STC, TIS, MPJPE-W) enhance diagnostics, exposing robustness across tolerance and highlighting failure modes not captured by conventional metrics. They are poised to become the standard in high-coverage WBC policy benchmarking.

Practical implications abound. Capability-aligned expert pipelines, though more complex, enable systematic recovery of diverse agile and balance-critical behaviors for hardware deployment. However, pipeline complexity, residual low-frequency sway artifacts, and coverage–generalization trade-offs remain unsolved. Further research is needed in frequency-domain regularization, reference correction, scalable expert distillation, and real-robot evaluations.

Conclusion

Athena-WBC establishes that structured capability-aligned expert pipelines, coupled with adaptive motion sampling and robust distillation, recover long-tail failures in humanoid WBC that are unreachable via standard data allocation. The integration of threshold-sensitive and motion-salience-aware diagnostics advances the evaluation of whole-body control policies in challenging high-coverage settings. This paradigm will inform future developments in scalable, robust, and generalizable humanoid control architectures, especially as real-robot deployment fidelity becomes a critical bottleneck.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.