- The paper introduces surrogate deep learning models leveraging a synthetic equilibrium dataset to accurately predict MHD parameters in the ADITYA-U tokamak.
- It integrates PCA, CNN architectures, and PINN constraints to achieve median errors below key thresholds and enable swift, real-time control.
- Physical insights are validated by comparing model predictions with experimental distributions, highlighting potential for broader fusion device applications.
Deep Learning Models for ADITYA-U MHD Equilibrium: Technical Analysis
Synthetic Dataset Construction and Physical Operational Domain
Accurate modeling of MHD equilibrium in ADITYA-U tokamak is addressed through the development of surrogate models trained on an expansive synthetic equilibrium dataset. The dataset comprises 100,760 instances, generated using the pyIPREQ Grad-Shafranov free-boundary solver with inputs from 766 ADITYA-U discharges. Physical realism is ensured via physics-informed filtering criteria, targeting flat-top phase, circular limiter configurations, and exclusion of nonphysical or numerically unstable cases. Key current profile parameters (α, β, γ) are determined through data-driven procedures, including Bayesian optimization and empirically constructed probabilistic models for poloidal beta (βp​). The operational domain spans Ip​∈[1.01×105,1.95×105] A, q0​∈[0.9,1.5], q1​>2.0, and βp​∈[0.05,0.4], accurately capturing ADITYA-U's experimental regime.
The dataset's statistical distributions of equilibrium quantities are shown to be consistent with experimental measurements; e.g., the sampled plasma current centroid matches the observed distribution (Figure 1).
Figure 1: Distribution of plasma current centroid positions sampled during equilibrium generation, showing consistency with experimentally observed flat-top data.
Comprehensive visualization of equilibrium distributions in (βp​,ℓi​) space, q-profile ensembles, and LCFS positions confirm broad but physically plausible coverage within the prescribed operational ranges.
Forward Models for Equilibrium Scalars and Inverse Coil Currents
Dense neural network architectures are constructed for scalar equilibrium parameters (β0) and coil currents (β1, β2, β3), optimized via Hyperband search. Scalar prediction models achieve high accuracy, with median absolute errors well within physically relevant thresholds (e.g., β4 for β5 and β6 for β7), despite operational variability. The inverse model for coil currents is shown to infer actuator settings with median errors of β8 A/turn for β9 and γ0 A/turn for γ1, supporting real-time experiment planning and actuator trend analysis.
Reduced-Order and Convolutional Surrogate Models for Profile Prediction
Safety Factor (γ2) Profile Models
Principal Component Analysis is leveraged for dimensionality reduction, revealing that γ3-profiles are nearly four-dimensional (99.99% explained variance). Dense regressors for PCA-coefficients yield median absolute errors γ4 for γ5 and γ6 for γ7 (Figure 2).
Figure 2: Mean absolute errors for reconstructed γ8-profiles demonstrate highly efficient low-rank representations and accurate predictions for both core and edge safety factors.
Alternatively, 1d-CNNs predicting γ9 enable physics-constrained monotonicity and outperform PCA models on local edge accuracy, with core errors βp​0 and edge errors βp​1 (Figure 3).
Figure 3: Distribution of CNN-predicted βp​2-profiles visualized for median, 99th percentile, and worst-case errors, highlighting robustness and physical consistency.
For βp​4 profiles, PCA models achieve 99.99% explained variance in five modes, but optimal accuracy requires up to 13 modes, as shown by improved mean absolute errors and reduced Grad-Shafranov residuals. Both PCA and CNN models are formulated as PINNs, directly incorporating Grad-Shafranov operator residual as a penalizing loss term. CNN models, via upsampling, achieve finer local accuracy but introduce spatially confined distortions; PCA models distribute errors globally and preserve dominant equilibrium structures (Figure 4).
Figure 4: Median and 99th percentile CNN reconstructions of βp​5 profiles display agreement in major features with some localized error, confirming suitability for real-time equilibrium inference.
Strong numerical claims are substantiated: PCA and CNN PINNs reconstruct full 2D profiles with median errors βp​6 across the operational domain, residuals approaching solver accuracy, and inference times βp​7 ms per equilibrium on CPU, suggesting applicability for real-time control systems.
Comparative Analysis and Physical Insights
The study establishes the operational domain as admitting efficient low-rank representations, with most variability encapsulated in a small number of latent dimensions, consistent with global equilibrium physics. PCA-based models excel at smooth global reconstructions, while CNNs capture local spatial variability more effectively. PINN formulations regularize out-of-domain predictions, ensuring physical compatibility.
Difficulties in predicting certain parameters (e.g., βp​8, βp​9, Ip​∈[1.01×105,1.95×105]0) indicate information bottlenecks in available diagnostics, motivating future inclusion of richer profile measurements and advanced regression strategies. Model limitations arise from dataset restriction to circular limiters; expansion to shaped plasmas and broader current-profile parameterization is feasible given the framework.
Practical and Theoretical Implications
Practically, these models allow rapid equilibrium prediction and actuator planning via direct mapping from measured signals to key equilibrium descriptors. Computational efficiency facilitates integration with plasma control systems, and PINN-based regularization enhances trustworthiness of real-time surrogates. Theoretically, the methodology demonstrates that well-designed synthetic datasets and low-rank surrogates can replace slow iterative solvers for operational parameter regimes. Extrapolation to other devices (e.g., KSTAR, EAST, JET) is supported by analogous studies using similar approaches [2020Joung, 2026Rutigliano, 2025Zheng].
Anticipated future developments include incorporation of advanced diagnostic signals, uncertainty-aware ensembles, transformer-based architectures, and extension to shaped plasma regimes. Inclusion of formal stability analysis and broader data-driven modeling will address current limitations and increase applicability for fusion device control.
Conclusion
This technical evaluation demonstrates that deep learning models, particularly PINNs and reduced-rank representations, are highly effective for fast MHD equilibrium prediction in ADITYA-U tokamak within experimentally relevant operational domains. The study provides authoritative evidence for model accuracy (Ip​∈[1.01×105,1.95×105]1 ms inference, absolute errors Ip​∈[1.01×105,1.95×105]2 in key quantities), scalability, and robustness. Physical regularization via PDE constraints is essential for maintaining predictive fidelity. The framework is immediately applicable to real-time control, experiment planning, and rapid data interpretation; future work will broaden operational scope and diagnostic integration. The systematic construction and evaluation of surrogate models on comprehensive synthetic datasets set a paradigm for data-driven equilibrium modeling across tokamak devices (2607.04865).