Papers
Topics
Authors
Recent
Search
2000 character limit reached

RecoverFormer: End-to-End Contact-Aware Recovery for Humanoid Robots

Published 24 Apr 2026 in cs.RO | (2604.22911v1)

Abstract: Humanoid robots operating in unstructured environments must recover from unexpected disturbances-a capability that remains challenging for end-to-end control policies. We present RECOVERFORMER, a fully end-to-end humanoid recovery policy that learns when and how to switch among recovery behaviors-including compensatory stepping, hand-environment contact, and center-of-mass reshaping-while maintaining robust performance under model mismatch. The architecture combines a causal transformer over a 50-step observation history with two novel heads: a latent recovery mode that enables smooth transitions among distinct recovery strategies, and a contact affordance head that predicts which environmental surfaces (walls, railings, table edges) are beneficial for stabilization. We evaluate RECOVERFORMER on the Unitree G1 humanoid in MuJoCo. Trained only on open floor, RECOVERFORMER transfers zero shot to walled environments, achieving 100% recovery success across 100-300 N pushes and across wall distances from 0.25-1.4m. Under zero-shot dynamics mismatch, RECOVERFORMER reaches 75.5% at plus +25% mass, 89% under 30 ms latency, 91.5% at low friction, and 99% under compound friction, latency and mass perturbation. The learned latent modes specialize across force regimes without mode-level supervision, validated by t-SNE analysis of 300 episodes. Taken together, these results show that a single end-to-end policy can deliver multi-modal, contact aware humanoid recovery that generalizes across perturbation magnitude, contact geometry, and dynamics shift.

Authors (1)

Summary

  • The paper introduces a causal-transformer reinforcement-learning policy that jointly selects compensatory stepping, hand bracing, and center-of-mass reshaping without explicit mode or contact supervision.
  • RecoverFormer achieves 100% simulated recovery across 100–300 N pushes in unseen walled environments and reaches 68% success at 250–300 N on open floor after training with only 5 million steps.
  • The results suggest latent recovery modes and contact affordances improve robustness, but simulation-only testing, privileged distance inputs, absent baselines, and unexplained compound-perturbation gains limit deployment conclusions.

Overview

RecoverFormer is a fully end-to-end reinforcement learning policy for humanoid push recovery that unifies three capabilities previously addressed separately: compensatory stepping, hand-environment bracing, and center-of-mass reshaping. The system is trained on the Unitree G1 (29 DoF, 1.27 m, 35 kg) in MuJoCo and evaluated exclusively in simulation. Its central claim is that a single policy, trained only on open floor, can learn to select among qualitatively distinct recovery strategies and exploit environmental contact opportunities without any explicit contact planning, mode supervision, or online adaptation module (2604.22911).

Architecture

The policy is a causal transformer over a 50-step observation history (1 s at 50 Hz). Each 106-dimensional observation frame—joint positions and velocities, projected gravity, torso velocities, foot contacts, distances from each hand to Kc=8K_c=8 candidate contact regions, and the previous action—is linearly embedded at d=256d=256 with learned positional codes and processed by L=4L=4 causal self-attention blocks with 4 heads. The final-timestep encoding ete_t feeds two heads:

  • Latent recovery mode head: a temperature-scaled softmax over K=4K=4 discrete modes. Training uses a differentiable soft mixture of learned mode embeddings; inference uses the argmax mode. No mode-level supervision is applied; instead, an entropy term plus a batch-level utilization penalty (umin⁡=0.4/Ku_{\min}=0.4/K) prevents collapse.
  • Contact affordance head: a sigmoid output over the 8 candidate regions predicting each surface's stabilization value. It receives no auxiliary loss and is trained purely through PPO gradients via a sparse contact reward that fires positively when a hand touches geometry whose contact normal opposes the fall direction, and negatively for wrist-over-torso impacts.

The action decoder concatenates encoder state, mode embedding, and affordance vector into 29 joint targets tracked by PD control at 50 Hz. Training uses PPO with GAE (γ=0.99\gamma=0.99, λ=0.95\lambda=0.95), domain-randomized pushes of 50–200 N with random direction and timing, and friction randomization μ∈[0.5,1.2]\mu\in[0.5,1.2]. Mass, latency, and torque limits are held out entirely for zero-shot evaluation. Total training is 5×1065\times10^6 steps across 64 parallel environments, roughly 2 hours on one RTX 5080—a notably modest compute budget.

Open-floor performance

On open floor, recovery success rate (RSR: remaining within 45° tilt and returning to stable standing within 10 s) degrades gracefully with force:

Force (N) 50 100 150 200 250 300
RSR (%) 100.0 100.0 85.0 79.0 68.0 68.0

Balance analysis shows peak torso tilt scaling sub-linearly with force (7° at 100 N vs. 20° at 250 N), which the authors attribute to learned active damping rather than passive impulse response. Base-height dips grow monotonically from 1.5 cm at 100 N to 6.5 cm at 250 N and recover within ~1 s. The implication is that the policy absorbs low-force perturbations in place and switches to multi-step strategies at high force—an emergent behavioral split examined further through the latent modes.

Zero-shot transfer to walled environments

The strongest result in the paper: trained exclusively on open floor, RecoverFormer achieves 100% RSR in walled environments at every force level from 100–300 N, and maintains 100% RSR across wall distances from 0.25–1.40 m and all four tested push directions. Since the affordance head was never exposed to walls during training, this indicates the representation captures a general notion of brace availability rather than memorized layouts. The claim rests on the affordance vector conditioning the action decoder even when no contacts occur during training rollouts on open floor; the mechanism by which an unsupervised affordance signal generalizes spatially to unseen geometry is plausible but not analyzed mechanistically, and the paper does not ablate the affordance head's contribution against a version without it.

Robustness under dynamics mismatch

At a fixed 150 N push, zero-shot perturbations yield:

Condition Nominal Low friction (d=256d=2560) Latency 30 ms +25% mass Compound
RSR (%) 93.5 91.5 89.0 75.5 99.0

Two observations stand out. First, single-axis mismatches cost little (≥89%), while +25% upper-body mass is the hardest single perturbation at 75.5%—the paper asserts memoryless policies collapse in this regime but provides no baseline comparison to substantiate it. Second, the compound condition outperforms every individual mismatch (99%), which the authors attribute to richer multi-source signal enabling faster implicit system identification through the 50-step history. This inversion is counterintuitive and under-explained; it could also reflect favorable interaction effects between randomized conditions rather than improved identification, and no variance estimates are reported for any condition.

Latent mode interpretability

t-SNE over episode-level mean mode vectors from 300 episodes (six force levels, 50 each) shows low-force episodes clustering tightly where Mode 3 ≈ 1, while high-force episodes spread as Modes 1 and 2 gain probability mass. Critically, failure episodes concentrate in the single-mode region regardless of force: high-force episodes responding with Mode 3 near 1 tend to fail, whereas multi-mode responses succeed. This supports the paper's claim that successful high-force recovery requires a multi-modal response and that interpretable strategies emerge from reward alone. The analysis is correlational, however—it establishes association between mode distribution and outcome, not causation, and does not verify that the four modes correspond to the three intended strategy categories (stepping, bracing, COM lowering).

Limitations and open questions

The paper is candid that all results are simulation-only; sim-to-real transfer on the physical G1 remains untested, and the causal transformer's streaming deployment advantage has not been validated against real sensor noise, actuator dynamics, or state-estimation error. The affordance head currently consumes ground-truth distances to candidate contact regions rather than visual perception, so deployment in genuinely unstructured environments requires perception work the paper defers. Baselines are absent throughout: the claims that memoryless policies collapse under mass mismatch and that prior methods cannot jointly reason about contact opportunities are asserted rather than measured. The compound-perturbation result exceeding nominal performance lacks a satisfying explanation. Finally, whether the utilization regularizer's choice of d=256d=2561 modes is necessary—or whether mode count materially affects robustness—is not explored.

Conclusion

RecoverFormer demonstrates that a compact transformer policy (~2 h of training) can unify multi-strategy humanoid recovery with contact-aware decision-making, achieving perfect zero-shot walled-environment success across wide force and distance ranges and ≥75% success under unseen dynamics mismatch, with emergent force-specialized latent modes confirmed by embedding analysis. The principal caveats are the absence of baselines, simulation-only validation, and reliance on privileged contact-region observations—all of which define the concrete questions the work leaves open.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 1 tweet with 1 like about this paper.