Compute-optimal solver fidelity and dataset design

Characterize the compute-optimal trade-off among solver fidelity, number of training trajectories, trajectory length, and neural model capacity for training neural emulators of partial differential equations.

Background

High-fidelity PDE data generation can be computationally expensive, while the thesis suggests that current neural emulators may not exploit solver accuracy beyond roughly 10-2 and may sometimes correct structured numerical errors. This motivates mixing trajectories generated at different solver fidelities rather than producing all data with a fully converged solver.

The unresolved problem is to determine how solver-side computation should be allocated relative to dataset size, trajectory duration, and model capacity. The question extends scaling-law analyses for neural operators by incorporating the computational cost and fidelity of the numerical solver.

References

This raises a question that, to our knowledge, has no answer yet: what is the compute-optimal trade-off between solver fidelity, number of trajectories, trajectory length, and model capacity?

From Numerical Simulators of PDEs to Neural Emulators and Back  (2608.24547 - Koehler, 25 Aug 2026) in Section 11.2, “Compute-Optimal Data Generation”