Broaden repeated-run coverage across benchmark tasks

Broaden the repeated-training-run coverage of PhysicsBench so that all tasks and data-scale findings are placed on the same statistical footing.

Background

PhysicsBench reports the mean over repeated training runs for five of its seven tasks, but 3D generation and 3D field prediction use single baseline runs in some or all cells because their rerun sets are incomplete. The authors note that non-monotonic scaling effects may therefore include residual run-to-run variation where repeated runs are unavailable.

Expanding the repeated-run coverage would provide more consistent estimates of variability and increase confidence in scale-dependent ranking and crossover claims across the entire benchmark.

References

Where the reruns exist, the effect survives averaging and is therefore a property of the model rather than of one draw. Where they do not, part of the swing may still be run-to-run. We deliberately retain these anomalies in the published results rather than smoothing them over, and broadening the repeated-run coverage remains open work.

PhysicsBench: A Unified Leaderboard for Generative and Predictive Models in Engineering Design and Simulation  (2608.24056 - Lee et al., 25 Aug 2026) in Section 'Limitations', subsection 'Statistical' (Section 5.1)