Overcoming the Reality Gap for Simulation-Based Policy Evaluation
Develop calibration and alignment techniques that overcome the reality gap specifically for simulation-based evaluation of robotic policies, ensuring that performance measured in simulation reliably matches performance in the real environment.
References
While methods to reduce this gap are very similar to the ones used for sim-to-real transfer, an open question is how to overcome the reality gap for the specific downstream purpose of model evaluations such that the performance of a policy in simulation matches the performance in the real environment.
Sim-to-real evaluation reads the sim-to-real gap directly, as the agreement between a ranking obtained in simulation and a ranking obtained on the robot, a metric that \citet{yang2025simtorealeval} propose but have not yet measured.
It also exceeds matched Domain Adaptation with Rewards from Classifiers (DARC) return in every environment, but its return differences from direct target Model-Based Policy Optimization (MBPO) remain unresolved.