Sim-to-real reliability of learning-based methods

Determine whether learning-based methods developed and evaluated in simulation reliably transfer to real-world robotic systems, and characterize the conditions that affect successful sim-to-real transfer.

Background

Simulation enables scalable evaluation and training, but domain gaps—especially in sensing and contact dynamics—can undermine real-world performance.

The authors note persistent uncertainty about whether methods validated in simulation will work in the real world, underscoring the need to understand and mitigate sim-to-real gaps.

References

While many learning-based methods are developed and evaluated in simulation, it is unclear whether they would work in the real world.

— From Machine Learning to Robotics: Challenges and Opportunities for Embodied Intelligence  (2110.15245 - Roy et al., 2021) in Section 6.2 (Assessing Robot Learning: Performance Evaluation)

However, future work should establish whether the behavioural effects on successful trajectories observed here extend to physical robot deployment.

— Is Success All You Need? Investigating the Impact of Input Perturbations on VLA Behaviour in Tabletop Manipulation Tasks  (2610.01351 - Higham et al., 1 Oct 2026) in Section LIMITATIONS AND FUTURE WORK

Transfer across embodiments remains an open problem due to significant differences in hydraulic actuation, dynamics, and geometry across commonly used machinery.

— Size Doesn't Matter: Material-State Reinforcement Learning for Excavator Transferable Soil Manipulation  (2609.12677 - Werner et al., 11 Sep 2026) in Section 2.4, “Transfer, Embodiment Dependence, and Field Deployment”

Robustness to real perception and contact mismatch remains unverified without physical arm--hand experiments.

— Outcome-Sensitive Motion Search for Impact-Aware Dexterous Catching  (2609.29020 - Pei et al., 24 Sep 2026) in Section 4, subsection "Ablation Studies," subsection "Limitations"

Real roofs contain shingles, seams, ridges, debris, damaged regions, compliance, and spatially varying friction. Although friction is randomized in simulation, robustness to scanned-mesh noise, surface uncertainty, roof edges, and weather conditions has not been established. Future work should incorporate local surface estimation, uncertainty-aware safety margins, and testing on more diverse roof materials and geometries.

— Learning Slope-Adaptive Whole-Body Locomotion for Humanoid Robots in Roofing Construction  (2609.20558 - Liu et al., 17 Sep 2026) in Section Limitations and Future Work

Simulations are, by nature, abstractions of reality, and the domain gap between simulated and real environments remains challenging and constitutes an open research question.

Validating GS-VLA on a real robot requires both an external metric-depth source (per the previous point) and the engineering effort to instrument a physical rig, neither of which we were able to put in place within the resource and personnel constraints of this project. We therefore report the simulator results as a strong but ultimately preliminary signal, and treat a real-robot replication as the natural follow-up.

— GS-VLA: Plug-and-Play Viewpoint Canonicalization for Frozen VLA Policies via Gaussian Splatting  (2608.19066 - Park et al., 19 Aug 2026) in Section 5, “Limitation,” subsection “Real-world evaluation”

Finally, evaluation on larger datasets and real field deployments is necessary to determine whether the learned reliability adaptation transfers beyond the controlled UMOD setting. Of particular interest is whether the increase in acoustic reliance observed under synthetic turbidity and blur also appears naturally as visibility changes during an underwater mission.

— Robust Cross-Modal Foundation Model Perception for Underwater Robots under Degraded Visual Conditions  (2608.19710 - Alam, 20 Aug 2026) in Section 7.6, Future Directions

The benchmark's conclusions are accordingly stated for simulated visuo-tactile policy learning; characterizing the simulation-to-real tactile gap on these assets is left to future work.

— SoftVTBench: A Deformation-Aware Visuo-Tactile Dataset and Benchmark for Deformable-Object Manipulation  (2608.18701 - Jing et al., 19 Aug 2026) in Supplementary Section 7, “Tactile Simulation Pipeline and Scope,” paragraph “Scope”

Whether the split holds under a different simulator, a learned policy, or a physical robot is open; repeating the audit on real hardware is the test we would run first.

— When Should a Failing Robot Ask? Initiating Corrective Human-Robot Dialogue from Audited Sensor Evidence  (2609.21942 - Pathak et al., 18 Sep 2026) in Section 6, Limitations and Future Work

However, extending ASGARD to other types of robotic systems and sim-to-real transfer remains an open direction for future work.

— ASGARD: Action-Space Guard for UAV Resilience via Reinforcement Learning  (2609.20982 - Salehi et al., 17 Sep 2026) in Discussion and Conclusion, Section titled “Discussion {content} Conclusion”

Though we did a lot to mitigate this issue in terms of the design of the eval the question still exists of whether models would behave this way in real-life.

— HarvestBench: Measuring Whether LLM Agents Will Pay to Avoid Killing Animals  (2609.04444 - Brazilek et al., 3 Sep 2026) in Section 6, Limitations, subsection “Game as a benchmark”

Identifying the factors behind the advantage over the dual-expert control still requires further controls in the matched setting and independent training seeds, and all results remain to be tested on physical robots.

— An Action Is Worth One Patch: Unified World-Action Modeling with PatchWAM  (2609.25961 - Wang et al., 22 Sep 2026) in Conclusion; Section 5, “Inference and deployment”

Additionally, most studies perform experiments and evaluations primarily in simulation, making it unclear whether the results translate to the real world.

The offline training data (Section~\ref{subsec_dataset}) also cover a limited set of operating conditions and disturbance patterns, so generalization to unseen scenarios remains open, and the cubic scaling of the NLP sensitivity computation with the number of decision variables will require decomposition or approximate sensitivities for larger networks.

— Two-Timescale Reinforcement Learning for Real-Time Optimization and Economic NMPC: Experimental Validation  (2609.35141 - Adhau et al., 28 Sep 2026) in Section 5, “Limitations and future work”

How to use these diagnostics to improve actual task success remains open.

— Natural State-Prediction Accuracy can Hide Weak Controlled Responsiveness in VLA Readouts  (2609.34684 - Kim et al., 28 Sep 2026) in Section 6.2, “Limitations and scope”

Future work will investigate whether the motion-grounded representations learned in simulation can generalize to real-robot settings with different viewpoints, object appearances, contact dynamics, and execution noise.

— MotionWeave: Learning Motion-Centered Future Dynamics for Vision-Language-Action Policies  (2609.39324 - Wang et al., 30 Sep 2026) in Conclusion

As with any advance in autonomous control, the methods could eventually contribute to robotic systems whose deployment raises safety considerations; we note that F-CIP encourages controllable rather than uncontrolled behavior, and that sim-to-real transfer remains an open problem for our method (Section 6).

— Bellman Meets Lyapunov: Unsupervised Reinforcement Learning via Mastering Chaos  (2610.02012 - Shah et al., 1 Oct 2026) in Ethics Statement; related limitation discussed in Section 6

We do not have ground-truth measurements of these properties for the reconstructed objects. Simulation export and consistency checks therefore do not establish that the scenes reproduce real-world dynamics. Collecting measured physical properties and validating simulated behaviour against real interactions are important future work.

— LiteReality-Agent: An Agentic System for Interactable 3D Indoor Scene Reconstruction  (2610.01863 - Huang et al., 1 Oct 2026) in Section "Limitations and Future Work," paragraph "Physical properties lack ground-truth validation"

Although P2 is unseen with respect to real contact coverage, it lies spatially between the seen placements P1 and P3; we therefore conjecture that its relatively strong performance may reflect spatial interpolation rather than pure extrapolation.

— SARI: Phase-Split Sim-Real Co-Training for Contact-Rich Manipulation  (2610.02804 - He et al., 2 Oct 2026) in Section 5.1, “Spatial Generalization and Real-Data Efficiency” (subsection label: Section 5.1)

Useful downstream transfer thus does not require perfect direct simulator fidelity; fidelity metrics alone are an incomplete measure of a simulator's value as a training-data generator, and characterizing which fidelity properties govern transfer remains open.

— Forecasting from Counterfactual Simulator Rollouts: A Sim2Real Evaluation  (2610.03662 - Wang et al., 2 Oct 2026) in Section 5, Discussion and Limitations

Our evaluation is primarily synthetic, leaving open how well these findings transfer to real-world systems.

— Time-series Foundation Models for Predictive Control: The Role of Excitation  (2610.06447 - Amria et al., 5 Oct 2026) in Section 5, “Conclusion”