Papers
Topics
Authors
Recent
Search
2000 character limit reached

PolySim: Multi-Simulator Humanoid Control

Updated 14 July 2026
  • PolySim is a multi-simulator framework for humanoid control that reduces sim-to-real gaps by leveraging diverse physics engines.
  • It employs dynamics-level domain randomization by training a single policy across simulators, enhancing robustness and transferability.
  • Empirical tests show improved motion-tracking accuracy and successful zero-shot deployment on a real Unitree G1 humanoid.

PolySim is a multi-simulator whole-body control training framework for humanoids that targets the sim-to-real gap arising from simulator inductive bias, defined as the structural assumptions baked into individual physics engines. Its central mechanism is to train a single policy jointly across multiple heterogeneous simulators within one loop, thereby realizing dynamics-level domain randomization rather than perturbing parameters inside only one engine. In the formulation reported in "PolySim: Bridging the Sim-to-Real Gap for Humanoid Control via Multi-Simulator Dynamics Randomization" (Lei et al., 2 Oct 2025), PolySim integrates IsaacGym, IsaacSim, Genesis, and MuJoCo, proves a tighter upper bound on simulator inductive bias than single-simulator training under its stated assumptions, and reports both reduced sim-to-sim discrepancy and zero-shot deployment on a real Unitree G1 without additional fine-tuning.

1. Conceptual basis and problem formulation

PolySim is motivated by the observation that humanoid WBC policies trained in a single simulator often overfit to that simulator’s contact, integration, and actuator assumptions. In the paper’s terminology, the sim-to-real gap “fundamentally arises from simulator inductive bias,” and these biases create discrepancies both across simulators and between simulation and hardware (Lei et al., 2 Oct 2025). Parameter-level domain randomization can broaden robustness within one engine, but it does not alter the underlying transition model; PolySim instead exposes the policy to a mixture of distinct engines during training.

The framework therefore treats simulator diversity itself as the randomization axis. A simulator index ss is sampled from a mixture p(s)p(s) with weights ww on the probability simplex, and within-engine parameters θ∼p(θ∣s)\theta \sim p(\theta \mid s) may optionally randomize friction, restitution, solver iterations, dt\mathrm{d}t, or actuator gains. The expected return optimized during training is

J(π)=Es∼p(s)Eθ∼p(θ∣s)Eτ∼Ps,θ(τ∣π)[∑t=0Tγtr(st,at)].J(\pi) = \mathbb{E}_{s \sim p(s)} \mathbb{E}_{\theta \sim p(\theta \mid s)} \mathbb{E}_{\tau \sim P_{s,\theta}(\tau \mid \pi)} \left[\sum_{t=0}^{T} \gamma^{t} r(s_t, a_t)\right].

A useful way to read this objective is that PolySim does not seek a policy specialized to any one transition kernel. Instead, it optimizes against a family of kernels induced by multiple engines and, in sim-to-real deployment, optionally by parameter-based DR as well. This suggests that PolySim operationalizes sim-to-real transfer as robustness to engine-class variation, not merely robustness to parameter perturbation.

2. System architecture and execution model

PolySim’s implementation is organized around training–simulation isolation and a unified routing layer that presents heterogeneous engines through one contract (Lei et al., 2 Oct 2025). The architecture is explicitly decomposed into a TrainClient, multiple SimServers, and a Simulator Router.

Component Function Details
TrainClient Simulator-agnostic RL loop Maintains the policy, computes observations and rewards, performs optimization
SimServers Engine hosting Advance physics and return typed tensors (o,r,d,info)(o, r, d, \text{info})
Simulator Router Engine virtualization Physics harmonization, API translation, numerical normalization

The TrainClient runs the simulator-agnostic RL loop, while each SimServer hosts one heterogeneous engine. The Simulator Router then virtualizes all servers as one vectorized environment. Its “physics harmonization” maps unified physical specifications to per-engine configurations, including friction and contact parameters, actuator model and gains, gravity, dt\mathrm{d}t, and rigid-body inertias or centers of mass. Its “API translation” enforces identical observation, action, and reward tensor shapes and semantics across engines. Its “numerical normalization” standardizes units such as radians, rescales actions from a normalized range to per-engine joint limits or target ranges, and clips commands within actuator or controller constraints.

The data path is kept device-resident by GPU pass-through implemented as RPC over NCCL, using NVLink or PCIe to avoid host copies and serialization overhead. The paper evaluates policies trained in IsaacGym, IsaacSim, and Genesis and tested zero-shot in MuJoCo; versions and per-engine low-level solver settings are not disclosed. Control frequency, policy architecture, exact feature lists, and exact actuator modes are also not disclosed, although the router is stated to expose consistent physical state features and synchronized clocks.

At the training-loop level, the procedure is straightforward. For each iteration, environments sample simulator assignments and optionally within-engine parameters; observations are collected as GPU tensors; actions are computed and dispatched through the router; physics is stepped; trajectories are aggregated across engines; and the policy is updated with a simulator-agnostic on-policy actor–critic algorithm, described as following standard humanoid WBC practice and exemplified by PPO.

3. Kernel lifting and the theoretical sim-to-real bound

The theoretical contribution is expressed on a common real-world state space. The real system is written as M∗=(S0,A,T0,R0,γ)M^* = (S_0, A, T_0, R_0, \gamma), while simulator ii is p(s)p(s)0, with a projection p(s)p(s)1 and a measurable section p(s)p(s)2 satisfying p(s)p(s)3 (Lei et al., 2 Oct 2025). Each simulator kernel is lifted back to the real state space as

p(s)p(s)4

For any policy p(s)p(s)5, discounted return is

p(s)p(s)6

Using the p(s)p(s)7-Wasserstein distance p(s)p(s)8 induced by a metric p(s)p(s)9 on ww0, and assuming ww1 is ww2-Lipschitz on ww3, the single-simulator sim-to-real gap is defined as

ww4

The paper gives the upper bound

ww5

where

ww6

PolySim then defines a convex mixture of lifted kernels,

ww7

and the corresponding gap

ww8

The main theorem states that the PolySim gap is lower than any single simulator:

ww9

The proof sketch relies on Bellman equations under θ∼p(θ∣s)\theta \sim p(\theta \mid s)0 and θ∼p(θ∣s)\theta \sim p(\theta \mid s)1, the Lipschitz property of θ∼p(θ∣s)\theta \sim p(\theta \mid s)2, and the Kantorovich–Rubinstein dual for θ∼p(θ∣s)\theta \sim p(\theta \mid s)3. Under the paper’s convex-hull assumption, mixtures can place θ∼p(θ∣s)\theta \sim p(\theta \mid s)4 closer to the real dynamics than any single vertex. A plausible implication is that PolySim’s empirical gains should be strongest when individual engines exhibit complementary rather than redundant biases.

4. Training task, reward structure, and evaluation metrics

The task studied is closed-loop motion imitation for humanoid WBC. The policy tracks reference whole-body motions drawn from the ASAP dataset, with 14 motions spanning easy, medium, and hard difficulty tiers (Lei et al., 2 Oct 2025). Training is reported for 10k, 15k, and 20k iterations respectively for these tiers. In sim-to-sim evaluation on MuJoCo, parameter-level DR is withheld in order to isolate the effect of multi-simulator training; in sim-to-real deployment, PolySim is “augmented with parameter-based DR” using ASAP defaults.

The reward adopts ASAP settings, but the paper does not disclose a full analytic reward decomposition. It states that typical terms in physics-based imitation include joint-space tracking, velocity tracking, center-of-mass or root pose tracking, end-effector positional tracking, foot-contact consistency, orientation penalties, and smoothness or control regularizers. Exact weights are not provided.

Evaluation is based on a formally defined success criterion and four error metrics. An episode is unsuccessful if at any time the mean body position error exceeds θ∼p(θ∣s)\theta \sim p(\theta \mid s)5. The global MPJPE is

θ∼p(θ∣s)\theta \sim p(\theta \mid s)6

and the root-relative MPJPE is

θ∼p(θ∣s)\theta \sim p(\theta \mid s)7

Acceleration and root-velocity errors are

θ∼p(θ∣s)\theta \sim p(\theta \mid s)8

The protocol reports mean values across all sequences. Hyperparameters such as learning rate, batch size, θ∼p(θ∣s)\theta \sim p(\theta \mid s)9, GAE, entropy coefficient, number of parallel environments per simulator, horizon length, and curriculum are not disclosed, although the RL algorithm and settings are held identical across simulators for fair comparison.

5. Empirical results in sim-to-sim and sim-to-real transfer

The principal quantitative result is that training on a heterogeneous simulator mixture reduces cross-simulator failure and improves motion-tracking quality on unseen engines (Lei et al., 2 Oct 2025). On MuJoCo, which is unseen during training in the reported setup, the three-simulator policy trained with IsaacSim, IsaacGym, and Genesis achieves the strongest success score among the listed methods.

Training recipe Success on MuJoCo dt\mathrm{d}t0
Genesis-only 0.121 175.671 mm
IsaacGym-only 0.500 214.159 mm
IsaacSim-only 0.036 174.838 mm
IsaacSim+IsaacGym+Genesis 0.564 151.924 mm

The same MuJoCo evaluation reports dt\mathrm{d}t1, dt\mathrm{d}t2, and dt\mathrm{d}t3 for the three-simulator policy. The improvement in success is stated explicitly as dt\mathrm{d}t4 percentage points over IsaacSim-only, dt\mathrm{d}t5 over IsaacGym-only, and dt\mathrm{d}t6 over Genesis-only. The paper’s abstract summarizes this as an improvement of 52.8 over an IsaacSim baseline on MuJoCo.

Performance gains are not confined to unseen engines. On IsaacGym, IsaacGym-only training attains Success dt\mathrm{d}t7 with dt\mathrm{d}t8, whereas IsaacGym+Genesis also attains Success dt\mathrm{d}t9 but improves several errors, including J(π)=Es∼p(s)Eθ∼p(θ∣s)Eτ∼Ps,θ(τ∣π)[∑t=0Tγtr(st,at)].J(\pi) = \mathbb{E}_{s \sim p(s)} \mathbb{E}_{\theta \sim p(\theta \mid s)} \mathbb{E}_{\tau \sim P_{s,\theta}(\tau \mid \pi)} \left[\sum_{t=0}^{T} \gamma^{t} r(s_t, a_t)\right].0, J(π)=Es∼p(s)Eθ∼p(θ∣s)Eτ∼Ps,θ(τ∣π)[∑t=0Tγtr(st,at)].J(\pi) = \mathbb{E}_{s \sim p(s)} \mathbb{E}_{\theta \sim p(\theta \mid s)} \mathbb{E}_{\tau \sim P_{s,\theta}(\tau \mid \pi)} \left[\sum_{t=0}^{T} \gamma^{t} r(s_t, a_t)\right].1, and J(π)=Es∼p(s)Eθ∼p(θ∣s)Eτ∼Ps,θ(τ∣π)[∑t=0Tγtr(st,at)].J(\pi) = \mathbb{E}_{s \sim p(s)} \mathbb{E}_{\theta \sim p(\theta \mid s)} \mathbb{E}_{\tau \sim P_{s,\theta}(\tau \mid \pi)} \left[\sum_{t=0}^{T} \gamma^{t} r(s_t, a_t)\right].2. On IsaacSim, IsaacSim-only training achieves Success J(π)=Es∼p(s)Eθ∼p(θ∣s)Eτ∼Ps,θ(τ∣π)[∑t=0Tγtr(st,at)].J(\pi) = \mathbb{E}_{s \sim p(s)} \mathbb{E}_{\theta \sim p(\theta \mid s)} \mathbb{E}_{\tau \sim P_{s,\theta}(\tau \mid \pi)} \left[\sum_{t=0}^{T} \gamma^{t} r(s_t, a_t)\right].3 and J(π)=Es∼p(s)Eθ∼p(θ∣s)Eτ∼Ps,θ(τ∣π)[∑t=0Tγtr(st,at)].J(\pi) = \mathbb{E}_{s \sim p(s)} \mathbb{E}_{\theta \sim p(\theta \mid s)} \mathbb{E}_{\tau \sim P_{s,\theta}(\tau \mid \pi)} \left[\sum_{t=0}^{T} \gamma^{t} r(s_t, a_t)\right].4, while IsaacSim+IsaacGym and IsaacSim+Genesis each reach Success J(π)=Es∼p(s)Eθ∼p(θ∣s)Eτ∼Ps,θ(τ∣π)[∑t=0Tγtr(st,at)].J(\pi) = \mathbb{E}_{s \sim p(s)} \mathbb{E}_{\theta \sim p(\theta \mid s)} \mathbb{E}_{\tau \sim P_{s,\theta}(\tau \mid \pi)} \left[\sum_{t=0}^{T} \gamma^{t} r(s_t, a_t)\right].5 with lower or competitive errors. On Genesis, IsaacGym-only training on an unseen engine yields Success J(π)=Es∼p(s)Eθ∼p(θ∣s)Eτ∼Ps,θ(τ∣π)[∑t=0Tγtr(st,at)].J(\pi) = \mathbb{E}_{s \sim p(s)} \mathbb{E}_{\theta \sim p(\theta \mid s)} \mathbb{E}_{\tau \sim P_{s,\theta}(\tau \mid \pi)} \left[\sum_{t=0}^{T} \gamma^{t} r(s_t, a_t)\right].6, whereas multi-simulator recipes reach Success J(π)=Es∼p(s)Eθ∼p(θ∣s)Eτ∼Ps,θ(τ∣π)[∑t=0Tγtr(st,at)].J(\pi) = \mathbb{E}_{s \sim p(s)} \mathbb{E}_{\theta \sim p(\theta \mid s)} \mathbb{E}_{\tau \sim P_{s,\theta}(\tau \mid \pi)} \left[\sum_{t=0}^{T} \gamma^{t} r(s_t, a_t)\right].7 with competitive errors; IsaacGym+Genesis is reported with J(π)=Es∼p(s)Eθ∼p(θ∣s)Eτ∼Ps,θ(τ∣π)[∑t=0Tγtr(st,at)].J(\pi) = \mathbb{E}_{s \sim p(s)} \mathbb{E}_{\theta \sim p(\theta \mid s)} \mathbb{E}_{\tau \sim P_{s,\theta}(\tau \mid \pi)} \left[\sum_{t=0}^{T} \gamma^{t} r(s_t, a_t)\right].8, J(π)=Es∼p(s)Eθ∼p(θ∣s)Eτ∼Ps,θ(τ∣π)[∑t=0Tγtr(st,at)].J(\pi) = \mathbb{E}_{s \sim p(s)} \mathbb{E}_{\theta \sim p(\theta \mid s)} \mathbb{E}_{\tau \sim P_{s,\theta}(\tau \mid \pi)} \left[\sum_{t=0}^{T} \gamma^{t} r(s_t, a_t)\right].9, (o,r,d,info)(o, r, d, \text{info})0, and (o,r,d,info)(o, r, d, \text{info})1.

A notable comparison is between parallel and sequential multi-simulator training. Sequential training produces only marginal improvements on some unseen engines, such as (o,r,d,info)(o, r, d, \text{info})2 on IsaacSim and (o,r,d,info)(o, r, d, \text{info})3 on Genesis, and can degrade performance on seen engines through catastrophic forgetting. PolySim’s fully parallel joint training is presented as avoiding forgetting and consistently improving both success and precision. The paper’s visualization claim is specific: only the PolySim policy executes a forward jump successfully in MuJoCo.

Runtime overhead is reported as small relative to the slowest engine in the mixture. Iteration times are (o,r,d,info)(o, r, d, \text{info})4 for IsaacGym+IsaacSim, (o,r,d,info)(o, r, d, \text{info})5 for IsaacGym+Genesis, (o,r,d,info)(o, r, d, \text{info})6 for IsaacSim+Genesis, and (o,r,d,info)(o, r, d, \text{info})7 for IsaacGym+IsaacSim+Genesis, corresponding to overheads of (o,r,d,info)(o, r, d, \text{info})8, (o,r,d,info)(o, r, d, \text{info})9, dt\mathrm{d}t0, and dt\mathrm{d}t1 respectively.

The paper also compares PolySim to parameter-level DR on unseen engines. For “Kobe” tracking on Genesis, IsaacGym_DR and IsaacSim_DR both achieve Success dt\mathrm{d}t2 but with dt\mathrm{d}t3 of dt\mathrm{d}t4 and dt\mathrm{d}t5 respectively, while PolySim with IsaacSim+IsaacGym improves to dt\mathrm{d}t6. On MuJoCo, PolySim with three simulators achieves Success dt\mathrm{d}t7 with dt\mathrm{d}t8, while the single-simulator DR baselines show Success dt\mathrm{d}t9 and much larger errors. The paper therefore argues that dynamics-level randomization through multiple engines increases robustness beyond what parameter-only DR provides.

For hardware transfer, the reported result is zero-shot deployment on a Unitree G1 humanoid without additional fine-tuning. Quantitative real-world metrics are not disclosed; the evidence is qualitative success and visual demonstrations. Control mapping, latencies, and exact state-estimation pipelines are likewise not disclosed, although action clipping consistent with actuator limits is stated as part of the safety mechanism.

6. Limitations, positioning, and terminological scope

PolySim does not claim to eliminate the sim-to-real gap. The router harmonizes physics and aligns clocks, but residual model-class differences in contact solver, integrator, and actuator nonidealities remain; the framework treats this diversity as a useful regularization source rather than as a nuisance to be fully removed (Lei et al., 2 Oct 2025). The theory also depends on the convex-hull assumption: if all included engines share similar biases along an important real-world axis, adding them may not reduce the gap. Concurrent training further requires careful resource scheduling, and API brittleness is an explicit engineering concern because the router must track per-engine interface changes while preserving invariants and numerical normalization.

In the paper’s positioning, PolySim differs from ordinary domain randomization because DR perturbs parameters around one engine and therefore preserves that engine’s transition structure. It differs from system identification because SysID, even when effective, remains tied to a single engine’s physics and is often device-specific and data-intensive. It also differs from cross-simulator interface frameworks such as HumanoidVerse and RoboVerse/MetaSim, which unify interfaces and benchmarks but do not jointly update one policy across multiple engines in parallel. Future directions listed in the paper include adding more engines such as RaiSim and Bullet, adaptive weighting M∗=(S0,A,T0,R0,γ)M^* = (S_0, A, T_0, R_0, \gamma)0 across simulators, combining the approach with model-based components or online SysID, and exploring automated curricula over engine meta-parameters or worst-case-kernel optimization through distributionally robust MDP formulations.

The name “PolySim” should also be distinguished from unrelated simulation software in other domains. "Polyply: a python suite for facilitating simulations of (bio-)macromolecules and nanomaterials" is an open-source Python suite for topology and coordinate generation in molecular dynamics of polymers, biomacromolecules, and nanomaterials (Grünewald et al., 2021). "FFPopSim: An efficient forward simulation package for the evolution of large populations" is a forward-time population-genetics package based on Fourier or Walsh-Hadamard methods for multi-locus evolution (Zanini et al., 2012). "SimPoly: Simulation of Polymers with Machine Learning Force Fields Derived from First Principles" is a first-principles MLFF framework for predicting polymer densities and glass transition temperatures (Simm et al., 15 Oct 2025). These works share neither PolySim’s humanoid-control problem setting nor its multi-simulator WBC training architecture.

Taken together, PolySim defines a specific line of research in which simulator heterogeneity is elevated from a source of nuisance variation to a training signal. Within the limits stated by the paper, its significance lies in unifying concurrent multi-engine training, a Wasserstein-style kernel-mixture analysis of simulator inductive bias, and empirical zero-shot transfer from simulation to a real humanoid platform.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to PolySim.