Projected Universal Policy (PUP)
- Projected Universal Policy (PUP) is a sim-to-real transfer framework that leverages two-stage system identification and a low-dimensional policy embedding to adapt control policies for dynamic robotic systems.
- The framework integrates model parameter randomization, latent space projection, and Bayesian optimization to effectively bridge the gap between simulation and real hardware performance.
- Experimental evaluations on the Darwin OP2 robot demonstrate that PUP outperforms nominal and robust baselines, achieving robust multi-gait biped locomotion with minimal real-world trials.
The Projected Universal Policy (PUP) is an algorithmic framework for sim-to-real transfer in dynamic robotic control, designed to address domain discrepancy between simulators and real hardware by leveraging two-stage system identification and a low-dimensional task-adaptive policy embedding. Developed in the context of bipedal locomotion for the Darwin OP2 robot, PUP integrates model parameter randomization, projection into a latent space, and Bayesian optimization to enable efficient policy adaptation and robust real-world deployment (Yu et al., 2019).
1. Mathematical Formulation
Let denote the vector of simulation model parameters, including friction, center-of-mass (COM), and actuator/PD control gains. The main components of PUP are:
- Parameter Embedding: A projection network maps to a low-dimensional latent variable (, with in experiments).
- Policy Network: outputs control actions given system state and embedding .
- Uniform Sampling: Parameters are sampled from 0, the range identified in pre-sysID.
- Training Objective: The expected simulated return,
1
is maximized using PPO with the clipped surrogate loss 2 over the joint network 3.
2. Two-Stage System Identification
2.1 Pre-sysID (Generic Data Collection)
Pre-sysID is performed by collecting hardware trajectories 4 using generic joint exercises and standing/falling sequences. The loss function for parameter fitting is:
5
where 6 are simulated joint positions and torso orientations, 7 is the standing+falling subset, and 8 is the standing subset.
- Nominal parameters 9 are found via
0
- 1 subsets 2 are sampled; for each,
3
- The parameter bounds are set per-dimension as:
4
CMA-ES is utilized to minimize 5, and bounds 6 are used to characterize simulator uncertainty for domain randomization.
2.2 Post-sysID (Task-Relevant Optimization)
Following PUP training, the latent embedding 7 is optimized for real-hardware performance by maximizing the real-world return 8, typically distance walked or time to fall, by running 9 directly on hardware.
- A Gaussian process surrogate for 0 is initialized with 1 uniformly sampled trials, then refined over 2 Bayesian optimization steps (e.g., Expected Improvement). Each hardware trial runs the controller for a full episode.
- The best 3 observed is selected for deployment.
The total post-sysID adaptation is completed within 25 hardware trials (≤ 15 min).
3. Neural Architectures and Training
3.1 Architecture Details
| Component | Architecture | Dimensions |
|---|---|---|
| Projection 4 | 2-layer FC (64→32) with tanh, final linear to 5 | 6 |
| Policy 7 | 2-layer FC (256→256) with tanh, input 8, output 20-d adjustments | 9 |
| State 0 | Motor positions (20), torso Euler angles (3) 1 2 time-steps | 249 |
| Action 3 | Joint target adjustments, 11 bins per joint | 4 |
PPO hyperparameters used: learning rate 5, discount 6, GAE 7, batch size 64 trajectories, horizon 8 steps. The reward combines forward velocity (9), torque regularization (0), torque-velocity (1), reference tracking (2), and a survival bonus (3).
Simulation randomization includes control frequency 4 Hz and sensor noise (position 5 rad, orientation bias 6 rad).
3.2 Pipeline Pseudocode
- Pre-sysID:
- Collect data 7.
- Find 8.
- For each of 9 subsets 0, solve for 1 with regularization.
- Compute per-dimension bounds 2.
- PUP Training:
- Initialize 3.
- Repeat: sample 4, compute 5, run 6 in simulation, optimize 7.
- Post-sysID:
- Initialize GP with 8 random 9.
- For 0 iterations: propose new 1 via acquisition, run on hardware, update GP.
- Deploy best 2.
4. Experimental Evaluation on Darwin OP2
4.1 Gait Results
After post-sysID, stable forward, backward, and sideways walking were achieved on the Darwin OP2. For a trip-wire at 3 m:
| Gait | Mean Distance to Fall (m) | Elapsed Time (s) |
|---|---|---|
| Forward | 4 | 5 |
| Backward | 6 | 7 |
| Sideways | 8 | 9 |
Walking speed was 0 m/s.
4.2 Baseline Comparison
Three approaches were compared:
| Method | Distance Traveled (m) Mean 1 Std (5 trials) |
|---|---|
| Nominal | 2 |
| Robust | 3 |
| PUP+BO | 4 |
PUP with Bayesian optimization substantially outperformed both nominal and robust (domain randomization only) baselines.
4.3 Ablation: NN-PD Actuator Model
Pre-sysID loss 5 was compared for two actuator models:
- PD Only: Fixed 6.
- NN-PD: 7 predicted by a 1-layer NN with 5 units (8).
NN-PD achieved approximately 9 lower 0 compared to PD only over 500 CMA-ES iterations. Hip position step responses and torso pitch were closer to real data for NN-PD.
4.4 Parameter Range Analysis
Some parameters (e.g., ankle damping 1) had tight pre-sysID bounds, indicating high confidence. Others (e.g., torque limit 2, COM3) exhibited wide bounds, reflecting greater simulation–real mismatch. The low-dimensional latent 4 allows PUP to compensate for these ambiguities.
5. Key Principles and Significance
Projected Universal Policy enables robust sim-to-real transfer by decoupling the modeling and control challenges into:
- Broad system identification (pre-sysID) to span plausible simulation variations.
- Training a policy family sensitive to a low-dimensional projection of the high-dimensional model uncertainty.
- Efficient real-world specialization (post-sysID) via Bayesian optimization in the latent space.
A plausible implication is that PUP serves as a general template for sim-to-real adaptation wherever model uncertainty can be characterized and projected, potentially benefiting domains beyond biped locomotion.
6. Conclusions
Projected Universal Policy combines principled domain randomization, differentiable model-to-embedding projection, and sample-efficient hardware adaptation, yielding robust and adaptive controllers on real robotic systems. On low-cost hardware, PUP enabled successful multi-gait biped locomotion using fewer than 25 real-world trials and demonstrated substantially improved performance over nominal or robust baselines (Yu et al., 2019).