Sim-to-real reliability of learning-based methods
Determine whether learning-based methods developed and evaluated in simulation reliably transfer to real-world robotic systems, and characterize the conditions that affect successful sim-to-real transfer.
References
While many learning-based methods are developed and evaluated in simulation, it is unclear whether they would work in the real world.
However, future work should establish whether the behavioural effects on successful trajectories observed here extend to physical robot deployment.
Transfer across embodiments remains an open problem due to significant differences in hydraulic actuation, dynamics, and geometry across commonly used machinery.
Robustness to real perception and contact mismatch remains unverified without physical arm--hand experiments.
Real roofs contain shingles, seams, ridges, debris, damaged regions, compliance, and spatially varying friction. Although friction is randomized in simulation, robustness to scanned-mesh noise, surface uncertainty, roof edges, and weather conditions has not been established. Future work should incorporate local surface estimation, uncertainty-aware safety margins, and testing on more diverse roof materials and geometries.
Simulations are, by nature, abstractions of reality, and the domain gap between simulated and real environments remains challenging and constitutes an open research question.
Validating GS-VLA on a real robot requires both an external metric-depth source (per the previous point) and the engineering effort to instrument a physical rig, neither of which we were able to put in place within the resource and personnel constraints of this project. We therefore report the simulator results as a strong but ultimately preliminary signal, and treat a real-robot replication as the natural follow-up.
Finally, evaluation on larger datasets and real field deployments is necessary to determine whether the learned reliability adaptation transfers beyond the controlled UMOD setting. Of particular interest is whether the increase in acoustic reliance observed under synthetic turbidity and blur also appears naturally as visibility changes during an underwater mission.
The benchmark's conclusions are accordingly stated for simulated visuo-tactile policy learning; characterizing the simulation-to-real tactile gap on these assets is left to future work.
Whether the split holds under a different simulator, a learned policy, or a physical robot is open; repeating the audit on real hardware is the test we would run first.
However, extending ASGARD to other types of robotic systems and sim-to-real transfer remains an open direction for future work.
Though we did a lot to mitigate this issue in terms of the design of the eval the question still exists of whether models would behave this way in real-life.
Identifying the factors behind the advantage over the dual-expert control still requires further controls in the matched setting and independent training seeds, and all results remain to be tested on physical robots.
Additionally, most studies perform experiments and evaluations primarily in simulation, making it unclear whether the results translate to the real world.
The offline training data (Section~\ref{subsec_dataset}) also cover a limited set of operating conditions and disturbance patterns, so generalization to unseen scenarios remains open, and the cubic scaling of the NLP sensitivity computation with the number of decision variables will require decomposition or approximate sensitivities for larger networks.
How to use these diagnostics to improve actual task success remains open.
Future work will investigate whether the motion-grounded representations learned in simulation can generalize to real-robot settings with different viewpoints, object appearances, contact dynamics, and execution noise.
As with any advance in autonomous control, the methods could eventually contribute to robotic systems whose deployment raises safety considerations; we note that F-CIP encourages controllable rather than uncontrolled behavior, and that sim-to-real transfer remains an open problem for our method (Section 6).
We do not have ground-truth measurements of these properties for the reconstructed objects. Simulation export and consistency checks therefore do not establish that the scenes reproduce real-world dynamics. Collecting measured physical properties and validating simulated behaviour against real interactions are important future work.
Although P2 is unseen with respect to real contact coverage, it lies spatially between the seen placements P1 and P3; we therefore conjecture that its relatively strong performance may reflect spatial interpolation rather than pure extrapolation.
Useful downstream transfer thus does not require perfect direct simulator fidelity; fidelity metrics alone are an incomplete measure of a simulator's value as a training-data generator, and characterizing which fidelity properties govern transfer remains open.
Our evaluation is primarily synthetic, leaving open how well these findings transfer to real-world systems.