Physics-Informed Multi-Machine Scaling Laws
- Physics-informed multi-machine scaling laws are empirical relations that use physical variables, such as divertor geometry or compute resources, to predict outcomes like plasma detachment and ACOPF feasibility.
- The approach conducts regressions constrained by dimensionless learning, ensuring that model features mirror key physical invariances such as plasma opacity and density limits.
- Applications in magnetic-confinement fusion and optimal power flow demonstrate that incorporating physics-based variables improves predictive accuracy and operational relevance.
Searching arXiv for papers on physics-informed multi-machine scaling laws and closely related work. I’ll look up the relevant arXiv records to ground the article in the latest preprints. Physics-informed multi-machine scaling laws are empirically calibrated relations in which the independent variables are selected or constrained by physical considerations and the fitting data are aggregated across multiple machines, devices, or operating configurations. In recent arXiv literature, the term appears most directly in magnetic-confinement fusion, where multi-machine laws are used to predict the fuel and impurity puffing rates sufficient to access plasma detachment across 32 devices, and more broadly in machine-learning studies of multi-machine AC optimal power flow, where power-law relations quantify how accuracy, constraint violations, and speed vary with data and compute resources on a 118-bus system with 54 generators (Moscheni et al., 28 Jul 2025, Liu et al., 6 Jan 2026). Across these settings, the central premise is that scaling should respect the relevant governing structure—geometry, opacity, density limits, feasibility constraints, or dimensional invariance—rather than rely on purely ad hoc regressions.
1. Conceptual scope and defining characteristics
Within the fusion literature, a multi-machine scaling law is a relation fitted on data pooled from many devices and then used to predict a control-relevant actuator, such as the deuterium or impurity puffing rate required for detachment onset. The underlying variables are not arbitrary statistical covariates: they are quantities such as divertor volume, divertor wetted surface, separatrix density, plasma opaqueness, and power fall-off length. This makes the law “physics-informed” in the sense that the regression is organized around variables motivated by edge-plasma transport and divertor geometry rather than around machine labels alone (Moscheni et al., 28 Jul 2025).
In the optimal-power-flow literature, the same phrase does not denote cross-device actuator laws. Instead, it refers to scaling relations for machine-learning surrogates operating on a multi-machine power-system benchmark. There, the resource dimensions are the number of training samples, denoted , and total training compute in FLOPs, denoted . The reported empirical ansatz is
with and taken from prediction error, constraint violations, or inference-time related metrics (Liu et al., 6 Jan 2026).
A common feature of both usages is that the scaling relation is intended to be operational. In fusion, the target is detachment access. In ACOPF, the target is predictable selection of data volume, model complexity, and compute budget for a prescribed accuracy-feasibility-speed operating point.
2. Methodological basis: physical invariance, constrained variables, and nested fitting
A general methodological foundation for physics-informed scaling is provided by dimensionless learning, which embeds dimensional invariance into a two-level machine-learning scheme. Given measured physical inputs whose units are encoded in a dimension matrix , any admissible dimensionless monomial
must satisfy
Because , the null space has dimension 0, and any solution can be parameterized as
1
where 2 is a null-space basis and 3 contains the free coefficients. Regression is then performed on 4 or on features constructed from it, while 5 is selected by maximizing predictive performance under the dimensional constraint (Xie et al., 2021).
This formalism is not itself a multi-machine law, but it clarifies what “physics-informed” means in a rigorous sense. The admissible forms are restricted before fitting. In the reported case studies, the approach recovers the Rayleigh number in turbulent Rayleigh-Bénard convection, the keyhole number in laser melting, and interpretable porosity-governing numbers in metal additive manufacturing, with reported 6 values ranging from 7 to 8 depending on the problem (Xie et al., 2021).
A plausible implication is that multi-machine scaling laws become more transferable when their variables are chosen to preserve physical invariance or known operating limits. This perspective is explicit in the fusion laws, where progressively richer models move from pure geometry to geometry-plus-opacity and then to density-limit-based simplifications.
3. Fuel-puffing laws for detachment access in magnetic-confinement devices
The clearest example of a physics-informed multi-machine law is the systematic review of fuel and impurity puffing rates sufficient for detachment access in magnetic-confinement fusion devices. The database contains 457 experimental and numerical data from 32 machines among solid-walled tokamaks, stellarators and linear plasma devices. The fitting variables include divertor wetted surface 9, divertor volume 0, separatrix electron density 1, plasma opaqueness 2, and the power fall-off length 3. The working assumption is that, to reach detachment onset with degree-of-detachment 4, plasma neutral opacity and divertor geometry govern the required deuterium puffing rate 5 (Moscheni et al., 28 Jul 2025).
The first step is purely geometric: 6 with 7 over 32 machines and 457 points. Here 8 is expressed in electron-equivalent particles per second.
Adding plasma opaqueness yields the more informative relation
9
for which the reported fit quality increases to 0. The exponent 1 is interpreted in the source as nearly linear scaling with the neutral-opacity weighting of plasma density and machine scale, while the 2 exponent on 3 indicates that more compact divertors require more fueling to achieve detachment.
To eliminate explicit dependence on 4, the study replaces it by the H-mode density limit and uses the Eich-Scarabosio fit for 5, obtaining
6
A stricter Greenwald-limited form is also reported: 7
These laws occupy different roles. The volume-only law provides a geometric baseline. The opacity-weighted law is the principal empirical predictor. The HDL and GES forms are simplifications or upper-bound-type constructions designed to remove unavailable inputs such as 8.
4. Impurity seeding, validation protocol, and stellarator extensions
The impurity-seeding problem is treated separately because impurity dynamics are described as inherently nonlinear and dependent on both the fuel puff and the pre-existing impurity level. The study introduces the composite variable
9
and reports a physics-guided regression with 0 for 1, based on characteristic divertor length, height, and radial location. Solving back for 2 yields the generalized detachment-onset impurity rate as a function of 3. Under the Greenwald-Eich-Scarabosio simplification, the impurity law becomes
4
where the weak negative exponent on 5 is interpreted as reflecting that broader scrape-off layers require less impurity to radiate away power (Moscheni et al., 28 Jul 2025).
The validation protocol separates fitting and onset validation by using 6 cases for fitting and 7 cases for validation. The full database contains 457 cases, comprising 193 from edge plasma codes and 264 from experiments, spanning conventional, spherical, and high-field tokamaks, stellarators, and linear plasma devices, as well as L-mode, H-mode, and advanced configurations. On the 8 validation set of 40 plasmas, the fuel law “OPQ” places 43% of predictions within 9 and 63% within 0, with geometric mean error about 1. The HDL bound has average error about 2 for H-modes. The GES upper bound is reported as never violated. For impurity seeding, the GES law places 50% of cases within 3 and 79% within 4, with geometric mean about 5 (Moscheni et al., 28 Jul 2025).
Stellarator-specific analogues are also proposed by replacing the H-mode density limit with the Sudo limit and 6 with a pedestal-width proxy or 7 scaling: 8
9
No 0 stellarator data were available for validation, so these relations remain untested.
5. Multi-machine ACOPF: scaling of data, compute, accuracy, and feasibility
A distinct but related line of work studies scaling laws for machine learning in optimal power flow. The focus is not on actuator scaling across machine classes, but on how surrogate performance scales in a multi-machine ACOPF setting. The principal benchmark is the IEEE 118-bus network, which has 54 generators and 186 branches. Training data consist of 50 K random load vectors 1 per case, generated by Gaussian perturbation around nominal values and retained only when OPF solves converge in Ipopt through PowerModels.jl. Fully connected ReLU DNNs are used as baselines, while the ACOPF PINN uses an identical 2 backbone that predicts 3 and adds a physics penalty 4 computed by re-running a full AC power flow in PyPower and accumulating violations of 5, 6, 7, and branch-flow limits (Liu et al., 6 Jan 2026).
The training loss is
8
The study reports consistent power-law relationships in both data scale and compute scale. On the 118-bus system with 1 K training epochs, the data-scaling fits are
9
0
with reported 1 of 2, 3, and 4, respectively. ACOPF constraint-violation fits are also reported: 5 with 6, 7, and 8 (Liu et al., 6 Jan 2026).
For compute scaling on the 118-bus case with 10 K samples and 9 measured in TFLOPs, the fitted laws include
0
1
Inference time per sample is reported to scale nearly linearly with model FLOPs; a 2 PINN on 118 buses achieves sub-millisecond per-sample inference on a single CPU core, and models below 3 FLOPs are stated to meet real-time 4 requirements.
An important finding is the divergence between prediction accuracy and constraint feasibility. PINNs are described as showing flatter accuracy-vs-data scaling but an almost step-function drop in violations. Concretely, the DNN ACOPF 5 scales as 6, whereas the PINN counterpart is observed empirically to scale at about 7, while 8 and 9 violations fall by over two orders of magnitude even at modest data sizes. On the 118-bus, 60-epoch comparison, ACOPF DNN at 40 K samples achieves 0 and 1 violation 2, while the PINN achieves 3 and 4 violation 5, summarized in the source as a 6 MAE penalty and 7 reduction in violations.
6. Interpretation, limitations, and related developments in complex multi-component systems
The primary misconception addressed by this body of work is that scaling laws can be treated as purely statistical trend lines independent of omitted physical interactions. The fusion study explicitly argues the opposite: divertor geometry alone correlates to fuelling, but the addition of plasma opaqueness improves predictive structure, and density-limit simplifications alter the role of directly measured edge density (Moscheni et al., 28 Jul 2025). The ACOPF study similarly shows that low prediction error does not guarantee constraint feasibility; the feasible manifold must be encoded through the PINN loss if the target is operational deployment rather than unconstrained regression (Liu et al., 6 Jan 2026).
Related arXiv work in other multi-component physical systems points in the same direction. In multi-RIS aided MIMO systems, a physics-compliant model derived from multiport network theory differs from the widely used multiplicative model because it retains the RIS structural-scattering term 8 rather than replacing it by 9. The reported discrepancy increases with the number of RISs and with multipath richness; in a four-RIS, 128-element-per-RIS multipath example, optimizing under the conventional model and evaluating under the physics-compliant model yields only 7% of the maximum channel gain (Nerini et al., 2024). In synchronously pitching foil schools, pure-pitching terms alone are reported to give inadequate estimates, whereas adding induced-velocity terms yields near-perfect linear collapse for both thrust and power in two-foil systems and extends without refitting to a three-foil configuration (Gungor et al., 2023).
This suggests a broader interpretation of physics-informed scaling laws. Their value is not merely in achieving compact regressions, but in preserving the interaction terms that dominate when multiple units, machines, or bodies are coupled. In that sense, the most durable scaling relations are those that remain aligned with first-principles structure even after aggregation across heterogeneous operating conditions.