Broaden repeated-run coverage across benchmark tasks
Broaden the repeated-training-run coverage of PhysicsBench so that all tasks and data-scale findings are placed on the same statistical footing.
References
Where the reruns exist, the effect survives averaging and is therefore a property of the model rather than of one draw. Where they do not, part of the swing may still be run-to-run. We deliberately retain these anomalies in the published results rather than smoothing them over, and broadening the repeated-run coverage remains open work.
— PhysicsBench: A Unified Leaderboard for Generative and Predictive Models in Engineering Design and Simulation
(2608.24056 - Lee et al., 25 Aug 2026) in Section 'Limitations', subsection 'Statistical' (Section 5.1)