Papers
Topics
Authors
Recent
Search
2000 character limit reached

Projected Universal Policy (PUP)

Updated 28 November 2025
  • Projected Universal Policy (PUP) is a sim-to-real transfer framework that leverages two-stage system identification and a low-dimensional policy embedding to adapt control policies for dynamic robotic systems.
  • The framework integrates model parameter randomization, latent space projection, and Bayesian optimization to effectively bridge the gap between simulation and real hardware performance.
  • Experimental evaluations on the Darwin OP2 robot demonstrate that PUP outperforms nominal and robust baselines, achieving robust multi-gait biped locomotion with minimal real-world trials.

The Projected Universal Policy (PUP) is an algorithmic framework for sim-to-real transfer in dynamic robotic control, designed to address domain discrepancy between simulators and real hardware by leveraging two-stage system identification and a low-dimensional task-adaptive policy embedding. Developed in the context of bipedal locomotion for the Darwin OP2 robot, PUP integrates model parameter randomization, projection into a latent space, and Bayesian optimization to enable efficient policy adaptation and robust real-world deployment (Yu et al., 2019).

1. Mathematical Formulation

Let μ∈Rn\mu \in \mathbb{R}^n denote the vector of simulation model parameters, including friction, center-of-mass (COM), and actuator/PD control gains. The main components of PUP are:

  • Parameter Embedding: A projection network fϕ:Rn→Rdf_\phi: \mathbb{R}^n \rightarrow \mathbb{R}^d maps μ\mu to a low-dimensional latent variable η\eta (d≪nd \ll n, with d=3d=3 in experiments).
  • Policy Network: πθ(a∣s,η)\pi_\theta(a|s, \eta) outputs control actions a∈Rma \in \mathbb{R}^m given system state ss and embedding η\eta.
  • Uniform Sampling: Parameters are sampled from fϕ:Rn→Rdf_\phi: \mathbb{R}^n \rightarrow \mathbb{R}^d0, the range identified in pre-sysID.
  • Training Objective: The expected simulated return,

fϕ:Rn→Rdf_\phi: \mathbb{R}^n \rightarrow \mathbb{R}^d1

is maximized using PPO with the clipped surrogate loss fϕ:Rn→Rdf_\phi: \mathbb{R}^n \rightarrow \mathbb{R}^d2 over the joint network fϕ:Rn→Rdf_\phi: \mathbb{R}^n \rightarrow \mathbb{R}^d3.

2. Two-Stage System Identification

2.1 Pre-sysID (Generic Data Collection)

Pre-sysID is performed by collecting hardware trajectories fϕ:Rn→Rdf_\phi: \mathbb{R}^n \rightarrow \mathbb{R}^d4 using generic joint exercises and standing/falling sequences. The loss function for parameter fitting is:

fϕ:Rn→Rdf_\phi: \mathbb{R}^n \rightarrow \mathbb{R}^d5

where fϕ:Rn→Rdf_\phi: \mathbb{R}^n \rightarrow \mathbb{R}^d6 are simulated joint positions and torso orientations, fϕ:Rn→Rdf_\phi: \mathbb{R}^n \rightarrow \mathbb{R}^d7 is the standing+falling subset, and fϕ:Rn→Rdf_\phi: \mathbb{R}^n \rightarrow \mathbb{R}^d8 is the standing subset.

  • Nominal parameters fϕ:Rn→Rdf_\phi: \mathbb{R}^n \rightarrow \mathbb{R}^d9 are found via

μ\mu0

  • μ\mu1 subsets μ\mu2 are sampled; for each,

μ\mu3

  • The parameter bounds are set per-dimension as:

μ\mu4

CMA-ES is utilized to minimize μ\mu5, and bounds μ\mu6 are used to characterize simulator uncertainty for domain randomization.

2.2 Post-sysID (Task-Relevant Optimization)

Following PUP training, the latent embedding μ\mu7 is optimized for real-hardware performance by maximizing the real-world return μ\mu8, typically distance walked or time to fall, by running μ\mu9 directly on hardware.

  • A Gaussian process surrogate for η\eta0 is initialized with η\eta1 uniformly sampled trials, then refined over η\eta2 Bayesian optimization steps (e.g., Expected Improvement). Each hardware trial runs the controller for a full episode.
  • The best η\eta3 observed is selected for deployment.

The total post-sysID adaptation is completed within 25 hardware trials (≤ 15 min).

3. Neural Architectures and Training

3.1 Architecture Details

Component Architecture Dimensions
Projection η\eta4 2-layer FC (64→32) with tanh, final linear to η\eta5 η\eta6
Policy η\eta7 2-layer FC (256→256) with tanh, input η\eta8, output 20-d adjustments η\eta9
State d≪nd \ll n0 Motor positions (20), torso Euler angles (3) d≪nd \ll n1 2 time-steps d≪nd \ll n249
Action d≪nd \ll n3 Joint target adjustments, 11 bins per joint d≪nd \ll n4

PPO hyperparameters used: learning rate d≪nd \ll n5, discount d≪nd \ll n6, GAE d≪nd \ll n7, batch size 64 trajectories, horizon d≪nd \ll n8 steps. The reward combines forward velocity (d≪nd \ll n9), torque regularization (d=3d=30), torque-velocity (d=3d=31), reference tracking (d=3d=32), and a survival bonus (d=3d=33).

Simulation randomization includes control frequency d=3d=34 Hz and sensor noise (position d=3d=35 rad, orientation bias d=3d=36 rad).

3.2 Pipeline Pseudocode

  1. Pre-sysID:
    • Collect data d=3d=37.
    • Find d=3d=38.
    • For each of d=3d=39 subsets πθ(a∣s,η)\pi_\theta(a|s, \eta)0, solve for πθ(a∣s,η)\pi_\theta(a|s, \eta)1 with regularization.
    • Compute per-dimension bounds πθ(a∣s,η)\pi_\theta(a|s, \eta)2.
  2. PUP Training:
    • Initialize πθ(a∣s,η)\pi_\theta(a|s, \eta)3.
    • Repeat: sample πθ(a∣s,η)\pi_\theta(a|s, \eta)4, compute πθ(a∣s,η)\pi_\theta(a|s, \eta)5, run πθ(a∣s,η)\pi_\theta(a|s, \eta)6 in simulation, optimize πθ(a∣s,η)\pi_\theta(a|s, \eta)7.
  3. Post-sysID:
    • Initialize GP with πθ(a∣s,η)\pi_\theta(a|s, \eta)8 random πθ(a∣s,η)\pi_\theta(a|s, \eta)9.
    • For a∈Rma \in \mathbb{R}^m0 iterations: propose new a∈Rma \in \mathbb{R}^m1 via acquisition, run on hardware, update GP.
    • Deploy best a∈Rma \in \mathbb{R}^m2.

4. Experimental Evaluation on Darwin OP2

4.1 Gait Results

After post-sysID, stable forward, backward, and sideways walking were achieved on the Darwin OP2. For a trip-wire at a∈Rma \in \mathbb{R}^m3 m:

Gait Mean Distance to Fall (m) Elapsed Time (s)
Forward a∈Rma \in \mathbb{R}^m4 a∈Rma \in \mathbb{R}^m5
Backward a∈Rma \in \mathbb{R}^m6 a∈Rma \in \mathbb{R}^m7
Sideways a∈Rma \in \mathbb{R}^m8 a∈Rma \in \mathbb{R}^m9

Walking speed was ss0 m/s.

4.2 Baseline Comparison

Three approaches were compared:

Method Distance Traveled (m) Mean ss1 Std (5 trials)
Nominal ss2
Robust ss3
PUP+BO ss4

PUP with Bayesian optimization substantially outperformed both nominal and robust (domain randomization only) baselines.

4.3 Ablation: NN-PD Actuator Model

Pre-sysID loss ss5 was compared for two actuator models:

  • PD Only: Fixed ss6.
  • NN-PD: ss7 predicted by a 1-layer NN with 5 units (ss8).

NN-PD achieved approximately ss9 lower η\eta0 compared to PD only over 500 CMA-ES iterations. Hip position step responses and torso pitch were closer to real data for NN-PD.

4.4 Parameter Range Analysis

Some parameters (e.g., ankle damping η\eta1) had tight pre-sysID bounds, indicating high confidence. Others (e.g., torque limit η\eta2, COMη\eta3) exhibited wide bounds, reflecting greater simulation–real mismatch. The low-dimensional latent η\eta4 allows PUP to compensate for these ambiguities.

5. Key Principles and Significance

Projected Universal Policy enables robust sim-to-real transfer by decoupling the modeling and control challenges into:

  • Broad system identification (pre-sysID) to span plausible simulation variations.
  • Training a policy family sensitive to a low-dimensional projection of the high-dimensional model uncertainty.
  • Efficient real-world specialization (post-sysID) via Bayesian optimization in the latent space.

A plausible implication is that PUP serves as a general template for sim-to-real adaptation wherever model uncertainty can be characterized and projected, potentially benefiting domains beyond biped locomotion.

6. Conclusions

Projected Universal Policy combines principled domain randomization, differentiable model-to-embedding projection, and sample-efficient hardware adaptation, yielding robust and adaptive controllers on real robotic systems. On low-cost hardware, PUP enabled successful multi-gait biped locomotion using fewer than 25 real-world trials and demonstrated substantially improved performance over nominal or robust baselines (Yu et al., 2019).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Projected Universal Policy (PUP).