Papers
Topics
Authors
Recent
Search
2000 character limit reached

What Matters for Simulation to Online Reinforcement Learning on Real Robots

Published 23 Feb 2026 in cs.RO and cs.AI | (2602.20220v1)

Abstract: We investigate what specific design choices enable successful online reinforcement learning (RL) on physical robots. Across 100 real-world training runs on three distinct robotic platforms, we systematically ablate algorithmic, systems, and experimental decisions that are typically left implicit in prior work. We find that some widely used defaults can be harmful, while a set of robust, readily adopted design choices within standard RL practice yield stable learning across tasks and hardware. These results provide the first large-sample empirical study of such design choices, enabling practitioners to deploy online RL with lower engineering effort.

Summary

  • The paper shows that standard off-policy RL, including SAC and vision-based DrQ, can adapt simulation-trained policies to real robots when practitioners retain prior data, warm-start replay buffers, and delay actor updates.
  • Across more than 100 runs on a Panda, Go1, and race car, annealed replay-data mixing and asymmetric actor-critic updates consistently improve stability, while symmetric updates often fail to outperform the starting policy.
  • The study finds that scaling simulation update rates with parallel environments and using roughly 1,000 or more randomized environments can improve hardware transfer, even when simulated returns appear unchanged.

Overview

This paper presents a large-sample empirical study of what enables stable online reinforcement learning (RL) on real robots when starting from simulation-trained policies — the "sim-to-online" setting. Across more than 100 real-world training runs on three platforms spanning manipulation (Franka Emika Panda), locomotion (Unitree Go1), and agile navigation (a remote-controlled race car at 60 Hz), the authors systematically ablate algorithmic, systems, and experimental design choices that prior work typically leaves implicit. The central finding is that standard off-policy RL — specifically Soft Actor-Critic (SAC) with a BRO critic architecture, extended to vision via DrQ — requires no fundamental algorithmic modification, but does depend critically on a small set of practical choices: retaining prior data, warm-starting the replay buffer, and asymmetric actor-critic updates. The full pipeline and the Panda hardware/software stack are open-sourced to lower the entry barrier for real-world RL research (2602.20220).

Motivation and the sim-to-online problem

The paper is motivated by the observation that most successful robotic RL systems learn entirely offline, in simulation, or from fixed demonstration datasets, leaving performance bounded by the quality of available priors. Since simulators remain inaccurate for contact-rich and vision-based tasks, and real-world robotics data is orders of magnitude scarcer than Internet-scale text, online adaptation on hardware is argued to be necessary for continued improvement. Prior demonstrations of real-robot RL either start from scratch (unsafe and wasteful), rely on custom hardware, or focus on algorithmic novelty without systematic study of deployment challenges.

The core technical difficulty is formalized through approximate policy iteration: after NN greedy updates, cumulative improvement is lower-bounded by the sum of greedy policy-improvement terms minus a penalty from approximation and modeling errors ϵ(s,a)\epsilon(s,a) in the learned action-value function QϕQ_\phi. In the sim-to-online setting, the deployed prior policy π0\pi_0 follows actions that are optimal under the simulator's dynamics p0p_0 but may visit state-action pairs where QϕQ_\phi has large errors on the real system. These errors compound across episodes and can dominate the improvement term, producing the "downward spiral" in which the policy underperforms π0\pi_0 — effectively unlearning the pretraining. The authors demonstrate this empirically: on a simulated Race Car under mild dynamics mismatch, Monte Carlo estimates show that the learned QϕQ_\phi substantially overestimates true action values over a large fraction of states added to the buffer during vanilla SAC training, while a stabilized variant keeps error mass concentrated near zero. This diagnostic directly connects the observed instability to value overestimation under distribution shift, rather than to generic optimization noise.

Stabilizing techniques

Three interventions are studied, all compatible with standard off-policy practice.

Data retention. Because QϕQ_\phi's approximation errors on the pretraining data D0D_0 are smaller on average than on newly collected data, retaining ϵ(s,a)\epsilon(s,a)0 and sampling critic minibatches from a mixture ϵ(s,a)\epsilon(s,a)1 regularizes the Bellman updates. Extending prior offline-to-online work, the authors anneal ϵ(s,a)\epsilon(s,a)2 during training, which they identify as essential when ϵ(s,a)\epsilon(s,a)3 comes from mismatched simulator dynamics — otherwise the mismatched transitions would permanently bias the critic.

Warm starts. When ϵ(s,a)\epsilon(s,a)4 cannot be retained, collecting a small number of episodes (ϵ(s,a)\epsilon(s,a)5 trials, roughly 1250–5000 transitions depending on the platform) with the frozen prior policy before any gradient updates serves as a partial proxy. This is shown to be necessary for stable learning on the Go1 and Race Car, though not on the Panda.

Asymmetric updates. Updating the actor only every ϵ(s,a)\epsilon(s,a)6 critic steps (with ϵ(s,a)\epsilon(s,a)7 and a reduced actor learning rate in the main experiments) drastically improves stability, consistent with two-timescale stochastic approximation theory. The sim-to-sim ablations sweep ϵ(s,a)\epsilon(s,a)8 and show monotone stability gains with larger ϵ(s,a)\epsilon(s,a)9 across all platforms; TD3 exhibits the same sensitivity to policy delay, indicating the effect is not SAC-specific.

Scaling off-policy RL in massively parallel simulation

A prerequisite for the pipeline is pretraining off-policy agents at scale, which is often considered impractical. The authors identify a concrete failure mode: popular SAC implementations perform a fixed number of updates per environment step, so the effective update-to-data (UTD) ratio QϕQ_\phi0 collapses as the number of parallel environments QϕQ_\phi1 grows, causing undertraining. Scaling QϕQ_\phi2 proportionally to QϕQ_\phi3 resolves this, with performance improving up to a task-dependent saturation point beyond which wall-clock cost grows without benefit. Notably, the number of domain-randomized environments matters for transfer even when simulated performance looks identical: a Go1 policy trained with QϕQ_\phi4 matches the QϕQ_\phi5 policy in simulation but transfers markedly worse to hardware. This is a strong and practically important claim — simulation return alone is not a reliable predictor of sim-to-real transfer quality, and roughly QϕQ_\phi6 parallel randomized environments appear necessary.

Real-world findings

Three headline results emerge from the hardware experiments, each with three random seeds per configuration.

First, data recycling across experiments accelerates online learning on every platform. Experiments are structured as sequences of trials sharing a random seed, where the online buffer of one trial seeds the prior buffer of the next. Performance and robustness increase monotonically with accumulated retained data, and on the Panda, roughly ten minutes of wall-clock training (including reset overhead and network latency) suffices to recover near-perfect pick-and-lift success from a prior policy that frequently failed to grasp. An ablation over the initial mixing weight QϕQ_\phi7 shows the recipe is robust: any scheme that uses prior data early and anneals toward purely online data works, with earlier online-data usage trading stability for speed.

Second, retaining simulation data itself is sufficient to stabilize learning, even without a warm start. Loading the simulator replay buffer used to train QϕQ_\phi8 and annealing QϕQ_\phi9 over five episodes substantially improves both efficiency and stability, which the authors interpret as the simulation data acting as a regularizer that biases critic updates toward low-error state-action pairs. This partially contradicts the position of Zhou et al. that offline data need not be retained for efficient online fine-tuning: in the sim-to-online regime with genuinely mismatched dynamics, the authors find retention valuable, though they concede that warm-started buffers are a viable fallback.

Third, asymmetric actor-critic updates are critical on hardware, not merely helpful. In the real-robot comparisons, the symmetric-update baseline fails to improve over π0\pi_00 on all three robots due to instability — and this remains true even when both runs are warm-started identically. The implication is that practitioners who omit this change should expect transfer failures that are attributable to update scheduling rather than to the choice of off-policy algorithm.

The paper also documents a set of engineering pitfalls with outsized effects: failing to restore optimizer state, target networks, or the SAC temperature π0\pi_01 (and its optimizer) upon resuming from pretraining each silently destabilizes fine-tuning. These are easy to overlook and can produce wrong conclusions, and the authors' explicit enumeration is a useful contribution for reproducibility.

Limitations and open questions

Several limitations are acknowledged directly. All experiments use episodic learning with manual resets, so the setting remains semi-automated; fully autonomous, reset-free learning is left open. The UTD ratios used on hardware are conservative (π0\pi_02 for the Panda and Race Car, π0\pi_03 for the Go1), so the sample-efficiency ceiling of the recipe on hardware is not established. The claim that π0\pi_04 is necessary for transfer rests on a comparison at two values of π0\pi_05 on a single platform, and the optimal π0\pi_06 saturation point is task-dependent and not characterized in closed form. The paper also leaves open how to select samples from π0\pi_07 optimally, how to reuse data across different tasks, and whether better regularization strategies could permit faster online learning.

Conclusion

This work provides the first large-sample, multi-platform empirical characterization of the design choices that determine whether simulation-pretrained off-policy policies can be fine-tuned stably on physical robots. Its actionable conclusions — annealed data retention, warm starts when retention is impossible, asymmetric actor-critic updates, and sufficiently broad domain randomization during pretraining — are simple, require no algorithmic novelty, and are validated on manipulation, locomotion, and agile navigation tasks, including a vision-based sparse-reward manipulation task solved within minutes of online training. The accompanying open-source stack makes these results directly reproducible on commodity hardware.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 4 tweets with 162 likes about this paper.