- The paper introduces a causal-transformer reinforcement-learning policy that jointly selects compensatory stepping, hand bracing, and center-of-mass reshaping without explicit mode or contact supervision.
- RecoverFormer achieves 100% simulated recovery across 100–300 N pushes in unseen walled environments and reaches 68% success at 250–300 N on open floor after training with only 5 million steps.
- The results suggest latent recovery modes and contact affordances improve robustness, but simulation-only testing, privileged distance inputs, absent baselines, and unexplained compound-perturbation gains limit deployment conclusions.
Overview
RecoverFormer is a fully end-to-end reinforcement learning policy for humanoid push recovery that unifies three capabilities previously addressed separately: compensatory stepping, hand-environment bracing, and center-of-mass reshaping. The system is trained on the Unitree G1 (29 DoF, 1.27 m, 35 kg) in MuJoCo and evaluated exclusively in simulation. Its central claim is that a single policy, trained only on open floor, can learn to select among qualitatively distinct recovery strategies and exploit environmental contact opportunities without any explicit contact planning, mode supervision, or online adaptation module (2604.22911).
Architecture
The policy is a causal transformer over a 50-step observation history (1 s at 50 Hz). Each 106-dimensional observation frame—joint positions and velocities, projected gravity, torso velocities, foot contacts, distances from each hand to Kc=8 candidate contact regions, and the previous action—is linearly embedded at d=256 with learned positional codes and processed by L=4 causal self-attention blocks with 4 heads. The final-timestep encoding et feeds two heads:
- Latent recovery mode head: a temperature-scaled softmax over K=4 discrete modes. Training uses a differentiable soft mixture of learned mode embeddings; inference uses the argmax mode. No mode-level supervision is applied; instead, an entropy term plus a batch-level utilization penalty (umin=0.4/K) prevents collapse.
- Contact affordance head: a sigmoid output over the 8 candidate regions predicting each surface's stabilization value. It receives no auxiliary loss and is trained purely through PPO gradients via a sparse contact reward that fires positively when a hand touches geometry whose contact normal opposes the fall direction, and negatively for wrist-over-torso impacts.
The action decoder concatenates encoder state, mode embedding, and affordance vector into 29 joint targets tracked by PD control at 50 Hz. Training uses PPO with GAE (γ=0.99, λ=0.95), domain-randomized pushes of 50–200 N with random direction and timing, and friction randomization μ∈[0.5,1.2]. Mass, latency, and torque limits are held out entirely for zero-shot evaluation. Total training is 5×106 steps across 64 parallel environments, roughly 2 hours on one RTX 5080—a notably modest compute budget.
On open floor, recovery success rate (RSR: remaining within 45° tilt and returning to stable standing within 10 s) degrades gracefully with force:
| Force (N) |
50 |
100 |
150 |
200 |
250 |
300 |
| RSR (%) |
100.0 |
100.0 |
85.0 |
79.0 |
68.0 |
68.0 |
Balance analysis shows peak torso tilt scaling sub-linearly with force (7° at 100 N vs. 20° at 250 N), which the authors attribute to learned active damping rather than passive impulse response. Base-height dips grow monotonically from 1.5 cm at 100 N to 6.5 cm at 250 N and recover within ~1 s. The implication is that the policy absorbs low-force perturbations in place and switches to multi-step strategies at high force—an emergent behavioral split examined further through the latent modes.
Zero-shot transfer to walled environments
The strongest result in the paper: trained exclusively on open floor, RecoverFormer achieves 100% RSR in walled environments at every force level from 100–300 N, and maintains 100% RSR across wall distances from 0.25–1.40 m and all four tested push directions. Since the affordance head was never exposed to walls during training, this indicates the representation captures a general notion of brace availability rather than memorized layouts. The claim rests on the affordance vector conditioning the action decoder even when no contacts occur during training rollouts on open floor; the mechanism by which an unsupervised affordance signal generalizes spatially to unseen geometry is plausible but not analyzed mechanistically, and the paper does not ablate the affordance head's contribution against a version without it.
Robustness under dynamics mismatch
At a fixed 150 N push, zero-shot perturbations yield:
| Condition |
Nominal |
Low friction (d=2560) |
Latency 30 ms |
+25% mass |
Compound |
| RSR (%) |
93.5 |
91.5 |
89.0 |
75.5 |
99.0 |
Two observations stand out. First, single-axis mismatches cost little (≥89%), while +25% upper-body mass is the hardest single perturbation at 75.5%—the paper asserts memoryless policies collapse in this regime but provides no baseline comparison to substantiate it. Second, the compound condition outperforms every individual mismatch (99%), which the authors attribute to richer multi-source signal enabling faster implicit system identification through the 50-step history. This inversion is counterintuitive and under-explained; it could also reflect favorable interaction effects between randomized conditions rather than improved identification, and no variance estimates are reported for any condition.
Latent mode interpretability
t-SNE over episode-level mean mode vectors from 300 episodes (six force levels, 50 each) shows low-force episodes clustering tightly where Mode 3 ≈ 1, while high-force episodes spread as Modes 1 and 2 gain probability mass. Critically, failure episodes concentrate in the single-mode region regardless of force: high-force episodes responding with Mode 3 near 1 tend to fail, whereas multi-mode responses succeed. This supports the paper's claim that successful high-force recovery requires a multi-modal response and that interpretable strategies emerge from reward alone. The analysis is correlational, however—it establishes association between mode distribution and outcome, not causation, and does not verify that the four modes correspond to the three intended strategy categories (stepping, bracing, COM lowering).
Limitations and open questions
The paper is candid that all results are simulation-only; sim-to-real transfer on the physical G1 remains untested, and the causal transformer's streaming deployment advantage has not been validated against real sensor noise, actuator dynamics, or state-estimation error. The affordance head currently consumes ground-truth distances to candidate contact regions rather than visual perception, so deployment in genuinely unstructured environments requires perception work the paper defers. Baselines are absent throughout: the claims that memoryless policies collapse under mass mismatch and that prior methods cannot jointly reason about contact opportunities are asserted rather than measured. The compound-perturbation result exceeding nominal performance lacks a satisfying explanation. Finally, whether the utilization regularizer's choice of d=2561 modes is necessary—or whether mode count materially affects robustness—is not explored.
Conclusion
RecoverFormer demonstrates that a compact transformer policy (~2 h of training) can unify multi-strategy humanoid recovery with contact-aware decision-making, achieving perfect zero-shot walled-environment success across wide force and distance ranges and ≥75% success under unseen dynamics mismatch, with emergent force-specialized latent modes confirmed by embedding analysis. The principal caveats are the absence of baselines, simulation-only validation, and reliance on privileged contact-region observations—all of which define the concrete questions the work leaves open.