- The paper introduces mCGCNN, which combines structural and magnetic graph streams, angle-aware exchange-pathway features, ligand identities, and magnetic-site pooling to model local magnetic interactions.
- The model reduces magnetic-moment prediction MAE to 2.0163 μB and raises test R² to 0.7756, outperforming CGCNN with a 20.7% MAE improvement and reducing dilution-related errors.
- The paper shows that moment-regression pretraining enables mCGCNN to achieve 75.68% FM/AFM classification accuracy, while limited data, high-moment scarcity, DFT label noise, and FM-biased predictions remain challenges.
Motivation and problem statement
The paper addresses a well-documented failure mode of crystal graph neural networks (GNNs): the prediction of magnetic properties from crystal structure alone. While architectures such as CGCNN (Burghardt, 2018), iCGCNN, MEGNet, ALIGNN, DimeNet, SchNet, PaiNN, and CrysXPP achieve strong performance on smooth scalar targets like formation energy and band gap, they perform poorly on magnetic quantities. The authors attribute this to four specific representational deficiencies:
- Homogeneous node treatment: standard CGCNN treats all atoms equivalently, whereas magnetism is concentrated on partially filled $3d$/$4f$ centers (Fe, Co, Ni, Mn, Cr, Gd, Nd).
- Mean-pool dilution: the crystal-level vector c=N1∑ihi(L) weights magnetic sites by Nm/N, systematically suppressing the signal in dilute systems.
- Absence of angular information: distance-only Gaussian RBF bond features cannot distinguish M–X–M pathways at 90∘ versus 180∘, which the Goodenough–Kanamori–Anderson (GKA) rules associate with ferromagnetic versus antiferromagnetic superexchange.
- Multi-valued DFT targets: spin-polarized DFT admits multiple self-consistent solutions (FM, AFM, ferrimagnetic) for the same structure depending on spin initialization, introducing structured label noise.
The paper's central claim is that these are architectural deficiencies that cannot be remedied by additional depth or width, motivating an explicitly magnetism-aware graph design.
Architecture
mCGCNN augments the full structural graph with a dedicated magnetic subgraph containing only moment-carrying nodes, defined either by a configurable element list or, where per-site moments are available, by a data-driven threshold ∣miDFT∣>0.1μB. The architecture has four distinguishing components.
Angle-aware magnetic bond features. For each directed edge (i,j) in the magnetic subgraph, the shortest bridging ligand k∗ is identified by minimizing rik+rkj, and the bond feature concatenates RBF expansions of all three leg distances, a Fourier angular basis $4f$0 with $4f$1 terms capable of representing the non-monotonic GKA dependence, and a one-hot ligand encoding. When no ligand exists (elemental metals, TM alloys), M–M–M angles are used instead.
Dual-stream message passing. The structural stream uses the original gated CGCNN convolution over all atoms; the magnetic stream runs an identical gated aggregation over magnetic centers using the angle-aware features. Magnetic nodes additionally receive Hund's-rule-derived spin and orbital momentum features.
Layer-wise cross-coupling. At each layer, structural representations are injected into the magnetic stream via a learned projection $4f$2. The coupling is deliberately unidirectional, which permits pretraining the structural stream on large formation-energy datasets and freezing it — though this option is not exercised in the reported experiments.
Magnetic sublattice pooling. A second pool restricted to the $4f$3 magnetic centers gives the magnetic contribution full weight independent of concentration; for $4f$4 the model reduces gracefully to standard CGCNN. The concatenated dual-pool representation feeds a fusion layer with layer normalization, dropout ($4f$5), and residual MLP blocks.
The computational overhead is bounded by the magnetic subgraph cost $4f$6 per layer, which never exceeds the structural stream cost since $4f$7; for dilute systems the overhead is negligible.
Regression results
The regression task predicts the DFT total magnetic moment over 24,916 Materials Project inorganic compounds (GGA+$4f$8, VASP), split 80/10/10 into 19,932/2,491/2,493 samples. The target distribution is strongly skewed toward low moments, with a sparse high-moment tail. Three-way benchmarking cleanly decouples generic architectural improvements from the magnetic physics: CGCNN+ retains single-graph message passing but adopts mCGCNN's enriched readout head, so the CGCNN+ → mCGCNN delta isolates the magnetic subgraph contribution.
| Model |
Test MAE ($4f$9) |
Test RMSE (c=N1∑ihi(L)0) |
Test c=N1∑ihi(L)1 |
| CGCNN |
2.5433 |
4.4230 |
0.6441 |
| CGCNN+ |
2.4129 |
3.9287 |
0.7192 |
| mCGCNN |
2.0163 |
3.5123 |
0.7756 |
mCGCNN reduces test MAE by 16% relative to CGCNN+ (20.7% relative to CGCNN) and improves c=N1∑ihi(L)2 from 0.644 to 0.776. Notably, roughly half of the total improvement over vanilla CGCNN comes from the modernized readout head alone, underscoring the value of the ablated baseline design. Parity analysis shows close agreement for moments below ~30 c=N1∑ihi(L)3 but systematic underestimation above ~40 c=N1∑ihi(L)4, with some high-moment compounds predicted at less than half their true value. The authors attribute this residual error to data scarcity in the high-moment tail rather than to architectural limits — a plausible reading given the distribution skew, though not directly demonstrated.
Classification results
The FM/AFM task uses 5,872 compounds balanced exactly 1:1 between classes (4,697/587/588 splits), so accuracy equals macro c=N1∑ihi(L)5 and the random baseline is exactly 50%. Four configurations are compared: direct training and moment-regression-pretrained transfer learning for both CGCNN and mCGCNN, with transfer following a two-phase protocol (frozen-backbone linear probing for 15 epochs, then joint fine-tuning with a lower backbone learning rate).
The results reveal a pronounced capacity–data tension. Direct mCGCNN dominates on training data (84.48% accuracy, over 8 points above CGCNN) but generalizes worse than CGCNN on test data (72.79% vs. 73.64%), exhibiting an 11.7-point train-to-test gap — the largest of any configuration. The authors conclude that the dual-stream capacity cannot be reliably regularized on fewer than 5,000 samples.
Transfer learning produces divergent outcomes that constitute the paper's most interesting finding:
- Negative transfer for CGCNN: pretraining degrades test accuracy from 73.64% to 71.94%, driven primarily by reduced FM recall. The authors hypothesize that a backbone without explicit magnetic-sublattice structure learns bulk structural correlates of moment magnitude that misalign with local exchange physics, while explicitly conceding this remains a hypothesis.
- Decisive positive transfer for mCGCNN: mCGCNN-Tr achieves the best test accuracy (75.68%) and macro c=N1∑ihi(L)6 (0.7563), compressing the generalization gap from 11.7 to 6.6 points. The gain is concentrated almost entirely in FM identification (FM errors drop from 76 to 58 samples relative to direct mCGCNN).
A systematic precision–recall asymmetry persists across all partitions despite exact class balance: FM recall exceeds AFM recall (0.8027 vs. 0.7109), indicating the model defaults to FM predictions under uncertainty. The authors attribute this to physical correlation between large net moment and FM order in the pretraining corpus — a prior that fine-tuning does not fully overcome. This is a candid acknowledgment that the pretraining signal is not neutral with respect to the downstream task.
The overall conclusion of this section is that the magnetic subgraph and regression-based pretraining are complementary: neither alone outperforms the simpler baseline on held-out data, but together they yield the best configuration.
Limitations and open questions
Several limitations are stated or implicit in the work. The classification gain, while consistent, is modest (+2.04 points over direct CGCNN), and the negative-transfer result for CGCNN-Tr lacks a mechanistic explanation beyond hypothesis. The multi-valued DFT label-noise problem identified in the introduction is acknowledged as unresolved; the authors note that larger curated datasets with explicit uncertainty quantification over multiple magnetic solutions are needed. Evaluation is limited to collinear order, total (not site-resolved) moments, and GGA+c=N1∑ihi(L)7 reference data whose own accuracy for exchange couplings is contested. The element-list definition of magnetic nodes introduces a hard-coded prior, although the threshold-based alternative partially mitigates this. Open questions left by the paper include whether the frozen-structural-stream pretraining strategy the architecture was designed to enable delivers additional gains, and whether the FM-biased prior induced by moment pretraining can be corrected.
Conclusion
mCGCNN demonstrates that encoding exchange geometry — magnetic sublattices, M–X–M angles via Fourier bases, ligand identity, and sublattice-restricted pooling — directly into a CGCNN-style architecture yields measurable gains on magnetic property prediction: test c=N1∑ihi(L)8 improvement from 0.644 to 0.776 and MAE reduction from 2.54 to 2.02 c=N1∑ihi(L)9 for moment regression, and best-in-class FM/AFM accuracy of 75.68% when combined with moment-regression pretraining. Equally instructive are the negative results: raw architectural capacity hurts without sufficient data or physically informed initialization, and pretraining can actively harm a representationally weaker backbone. The work establishes exchange-pathway-aware graph construction as a principled design axis for magnetic materials ML, while leaving site-resolved moments, ordering temperatures, non-collinear structures, and uncertainty-aware handling of multiple DFT magnetic solutions as open problems.