---
title: Interaction-based Material Network (IMN)
url: https://www.emergentmind.com/topics/interaction-based-material-network-imn
type: topic
---

# Interaction-based Material Network (IMN)

Searching arXiv for IMN/DMN/ODMN papers and related material-network work.
arxiv_search(query="interaction-based material network deep material network ODMN polycrystalline materials modeling", max_results=10, sort_by="relevance")
Interaction-based Material Network (IMN) denotes a family of structure-preserving, mechanistic machine-learning surrogates for multiscale material modeling that retain an explicit homogenization architecture rather than replacing it with a purely black-box predictor. Across the literature, the term is used for closely related formulations derived from Deep Material Networks (DMNs): an interface-enriched network with cohesive layers for heterogeneous media with interfacial failure [1908.04858], a Hill–Mandel-based interaction formulation for polycrystalline homogenization and nonlinear upscaling [2502.02457], and a rotation-free alternative to the original DMN in which interaction planes are parameterized directly by unit normals [2602.07192]. In all of these uses, the central idea is that microstructural subregions are organized in a binary tree, combined through analytical two-phase building blocks, trained primarily on linear-elastic data, and then deployed for nonlinear, path-dependent online prediction.

## 1. Terminology, lineage, and scope

The IMN emerged from the DMN program as a way to make phase interaction more explicit while preserving the interpretability of analytical homogenization blocks. In the work by Liu et al., "Deep material network with cohesive layers: Multi-stage training and interfacial failure analysis" [1908.04858], IMN refers to a DMN enriched by a hierarchy of zero-thickness cohesive building blocks. Each selected bottom-layer node is attached to a small subnetwork of serial cohesive layers, so that interfacial traction–separation behavior can be represented within the same reduced-order architecture that handles bulk homogenization.

In later polycrystal modeling, the term is used for a Hill–Mandel-based interaction mechanism that directly encodes stress-equilibrium directions among RVE subregions. The paper "Orientation-aware interaction-based deep material network in polycrystalline materials modeling" [2502.02457] describes this interaction mechanism as the core of the ODMN. There, the IMN component captures equilibrium among subregions, while an orientation-aware mechanism learns crystallographic textures and enables texture evolution prediction.

A systematic comparison later characterized IMN as a rotation-free formulation relative to the original DMN. Rather than orienting each interaction by three Euler angles, IMN directly parameterizes the unit normal of the interaction plane by two angles, yielding a more compact offline training problem while maintaining comparable online prediction accuracy for sufficiently deep networks [2602.07192].

This usage suggests that IMN is best understood not as a single fixed architecture but as a class of analytical material networks in which interaction physics is made explicit at the building-block level. The invariant features are the binary-tree topology, physically interpretable parameters, and the reuse of linear-elastic offline data to support nonlinear online extrapolation.

## 2. Governing mechanics and interaction laws

A central formulation of IMN is based on the Hill–Mandel macro–micro energetic consistency condition. In the ODMN presentation, the virtual-work statement is written as
\[
\delta W_{\mathrm{macro}}=\bar{\sigma}:\delta\varepsilon,\qquad
\delta W_{\mathrm{micro}}=\sum_{i=0}^{2^N-1} W^i\,\sigma^i:\delta\varepsilon^i,
\]
which implies
\[
\sum_i W^i\,\sigma^i:(\delta\varepsilon^i-\delta\varepsilon)=0.
\]
The interaction assumption is a rank-one decomposition of each subregion strain increment,
\[
\delta\varepsilon^i=\delta\varepsilon+\sum_{j=0}^{2^N-2}\alpha^{i,j}\,a^j\otimes N^j,
\]
and substitution into the Hill–Mandel condition yields stress-equilibrium constraints
\[
\sum_{i=0}^{2^N-1} W^i\,\alpha^{i,j}\,\sigma^i\cdot N^j=0
\qquad
\text{for } j=0,\dots,2^N-2.
\]
In the linear regime, these relations define a small system for the interaction variables \(a^j\), while the equilibrium directions \(N^j\) themselves are trainable in the homogenization stage [2502.02457].

The same interaction physics can be expressed at the level of each binary homogenization block. The rotation-free IMN formulation defines a unit normal
\[
\vec n=\bigl[\cos(2\pi\theta)\sin(\pi\phi),\;\sin(2\pi\theta)\sin(\pi\phi),\;\cos(\pi\phi)\bigr],
\]
and imposes rule-of-mixtures relations for strain and stress together with traction continuity,
\[
\boldsymbol\varepsilon^h=f^1\boldsymbol\varepsilon^1+f^2\boldsymbol\varepsilon^2,\qquad
\boldsymbol\sigma^h=f^1\boldsymbol\sigma^1+f^2\boldsymbol\sigma^2,
\]
\[
\mathbf H(\vec n)^T(\boldsymbol\sigma^2-\boldsymbol\sigma^1)=0.
\]
Solving these relations yields a closed-form homogenized stiffness,
\[
\mathbf C^h
= f^1 \mathbf C^1 + f^2 \mathbf C^2
-\left(\mathbf C^2-\mathbf C^1\right)\mathbf H\mathbf b,
\]
with \(\mathbf b\) defined analytically in terms of \(\mathbf H\), \(f^1\), \(f^2\), and the child stiffness tensors [2602.07192].

The ODMN paper states the same binary merge through the operator
\[
\bar C = H_2(C^0,C^1,f^0,f^1,N)
= f^0 C^0 + f^1 C^1 - f^0 f^1 (C^0-C^1)Q(C^0-C^1),
\]
where
\[
Q = H\,S^{-1}\,H^T,\qquad
S=H^T(f^1C^0+f^0C^1)H.
\]
This formulation makes the interaction direction \(N\) part of the learned physics [2502.02457].

In the cohesive IMN, interaction is enriched further by explicit interface kinematics. The compliance of a cohesively bonded block is propagated analytically as
\[
D_{\mathrm{block}}
=
D^0 + \hat z\,R(\tilde\alpha,\tilde\beta,\tilde\gamma)\,\hat G\,R(\tilde\alpha,\tilde\beta,\tilde\gamma)^{-1},
\]
where \(\hat z=a(\tilde z)/L\) is a reciprocal-length activation and \(T_3=1/\hat z\) is the effective layer thickness. The online cohesive law follows the Camacho–Ortiz bilinear form with
\[
t_m = K d_m \qquad (0\le d_m\le d_0),
\]
\[
t_m = \sigma_c \frac{d_m-d_f}{d_0-d_f}\qquad (d_0\le d_m\le d_f),
\]
and
\[
G_c=\tfrac12 \sigma_c d_f,\qquad d_f=\frac{2G_c}{\sigma_c},
\]
with irreversible softening enforced by updating \(d_0\) when \(d_m\) exceeds it during loading [1908.04858].

## 3. Binary-tree architecture and parameterization

IMN retains the binary-tree topology of DMN. A depth-\(N\) network has \(2^N\) leaf nodes and \(2^N-1\) internal interaction nodes in the polycrystal formulations, while the systematic assessment states that an \(N\)-layer tree has \(2^N\) base leaves and \(2^{N+1}-1\) total nodes when the entire hierarchy is counted [2502.02457] [2602.07192]. Each parent node combines two child subtrees through an analytical two-phase homogenization block.

In the ODMN, each material leaf \(i\) carries three trainable Tait–Bryan angles \((\alpha^i,\beta^i,\gamma^i)\) and a scalar \(z^i\), with weight
\[
W^i=\mathrm{softplus}(z^i)=\ln(1+e^{z^i}).
\]
In the crystal frame, the node obeys a linear law
\[
\sigma_c^i = C^i \varepsilon_c^i,
\]
and the rotated stiffness in the specimen frame is
\[
C_R^i = Z^{R1}Y^{R1}X^{R1}\;C^i\;(X^{R2})^{-1}(Y^{R2})^{-1}(Z^{R2})^{-1}.
\]
During plastic loading, the elastic deformation gradient is polar-decomposed as
\[
F_e^i(t)=R_t^i U_t^i,
\]
and the updated orientation \(R_t^i\) is used to reconstruct the orientation distribution function from \(\{R_t^i,W^i\}\) [2502.02457].

The later texture-generalizable framework states the full ODMN parameter set as
\[
\mathcal F
=
\{z^i,\alpha^i,\beta^i,\gamma^i\}
\cup
\{\theta_p^l,\phi_p^l\},
\]
where \((\theta_p^l,\phi_p^l)\) parameterize each internal interaction node and the leaf angles represent crystallographic orientation [2512.06779].

In the original DMN and cohesive IMN, bottom-layer activations use ReLU,
\[
w_N^n=\mathrm{ReLU}(z^n)=\max(z^n,0),
\]
and parent weights are accumulated recursively. The 2019 cohesive IMN emphasizes that all fitting parameters \(\{z^j,\alpha^{ik},\beta^{ik},\gamma^{ik}\}\) have clear physical interpretations in terms of topology and orientation of phase patterns, while each added cohesive layer contributes its own reciprocal length, orientation, and softening law [1908.04858]. When all cohesive blocks are deactivated, the original DMN is exactly recovered.

This architecture yields a constrained hypothesis space. The network does not learn an arbitrary constitutive map; it learns the parameters of a hierarchical assembly of physically meaningful homogenization and interaction blocks. A plausible implication is that this architectural bias underlies the repeatedly reported extrapolation from linear-elastic training data to nonlinear deployment.

## 4. Offline training and identification from linear-elastic data

A defining property of IMN formulations is that offline training uses linear-elastic DNS data only. In the ODMN, the training set consists of 500 linear-elastic RVE stiffnesses \(\bar C^{DNS}\) generated by DAMASK-FFT with random/stability-filtered cubic \(C_{11},C_{12},C_{44}\), and the loss is a relative MSE over a minibatch,
\[
\mathrm{Loss}(\mathcal A)
=
\frac{1}{N_{\mathrm{batch}}}
\sum_{k=1}^{N_{\mathrm{batch}}}
\frac{\|\bar C_k^{DNS}-\bar C_k^{ODMN}\|_F^2}
{\|\bar C_k^{DNS}\|_F^2}.
\]
The paper also gives an optional texture-error monitor,
\[
\hat T^d
=
\frac{\int [f_{ODMN}(g)-f_{DNS}(g)]^2\,dg}
{\int [f_{DNS}(g)]^2\,dg},
\]
and trains with AdamW, weight-decay \(=10^{-2}\), learning rate \(=10^{-3}\), batch-size \(=20\), epochs \(=200\), with a hold-out set of 100 samples [2502.02457].

The cohesive IMN adopts a two-stage training strategy on linear-elastic DNS data only. Stage I fits the elastic DMN parameters by minimizing
\[
J_1(\{z,\alpha,\beta,\gamma\})
=
\frac{1}{2N_s}\sum_{s=1}^{N_s}
\frac{\|\bar C_s^{dns}-\bar C_s^{rve}\|^2}{\|\bar C_s^{dns}\|^2}
+\lambda\sum_j z_j^2,
\]
using SGD with back-propagation of analytical gradients, mini-batch size \(\sim 20\), and a Bold-driver learning-rate schedule. Stage II fixes the DMN parameters and fits the cohesive activations and angles for each enriched node and layer, again using the overall stiffness MSE. The procedure includes model compression by deleting unused nodes, merging similar sub-trees, and, for the cohesive stage, merging layers whose normals coincide by summing activations [1908.04858].

The 2026 assessment reformulates the shared training objective for DMN and IMN as
\[
\mathcal L(\psi)
=
\frac1{2N_s}\sum_{j=1}^{N_s}
\frac{\Vert \mathbf C_j^{DNS}-\hat{\mathbf C}_j\Vert^2}
{\Vert \mathbf C_j^{DNS}\Vert^2}
+
\eta\Bigl(\sum_{n=1}^{2^N}\mathrm{ReLU}(z^n)-\xi\Bigr)^2,
\]
where the second term regularizes the total active base-node mass and controls network complexity. Best-practice settings reported there are Adam, 10,000 epochs, initial learning rate \(1\times10^{-2}\), learning-rate reduction by \(0.8\) if no validation-loss improvement for 50 epochs, and activation regularization with \(\eta=1,\;\xi=1\) as the best trade-off [2602.07192].

For texture-generalizable ODMN, the later TACS-GNN-ODMN framework extends the offline stage in two directions. Texture-Adaptive Clustering and Sampling initializes the texture-related parameters from an ODF histogram, and a two-layer GATv2Conv graph neural net with global mean pooling predicts the \(2(2^N-1)\) stress-equilibrium angles. Its dataset comprises \(4\) texture classes \(\times\) \(360\) RVEs \(=1440\) RVEs, with 500 crystal-stiffness triplets per RVE, giving \(720\,000\) pairs of \((C^{\mathrm{crystal}},C^{DNS})\), and uses AdamW with learning rate \(1e^{-3}\), \(\beta_1=0.9\), \(\beta_2=0.98\), weight-decay \(1e^{-4}\), batch size \(512\), and early stopping around epoch \(35\) for \(N=7\) [2512.06779].

## 5. Online nonlinear prediction and constitutive deployment

Once trained, IMN is used as an online multiscale solver rather than merely as a stiffness regressor. In the ODMN, the inputs are the macroscopic deformation gradient \(\bar F(t)\) and time increment \(dt\). Internal state variables per leaf include \(F_e^i\), \(F_p^i\), plastic variables \(Z^i\), interaction variables \(a^j\), weights \(W^i\), and orientations \(R^i\). The outputs are the homogenized stress \(\bar P(t)\), tangent \(\bar L(t)\), and updated textures \(\{R_t^i,W^i\}\) [2502.02457].

The online solve uses the same tree as the offline homogenization. The downscaling map is
\[
F^i=\bar F+\sum_j \alpha^{i,j}(a^j\otimes N^j).
\]
Local constitutive responses are then evaluated as
\[
P^i=\mathcal P^i(F^i,Z^i),
\]
internal variables are updated, and the interaction variables are obtained from the equilibrium conditions
\[
\sum_i W^i \alpha^{i,j} P^i\cdot N^j = 0
\]
by Newton–Raphson. Upscaling gives
\[
\bar P = \frac{\sum_i W^i P^i}{\sum_i W^i}.
\]
Algorithm 2 in the paper summarizes the forward pass: initialize \(F_e^i\leftarrow R^i\), \(F_p^i\leftarrow(R^i)^{-1}\), \(a^j\leftarrow 0\) at \(t=0\); iterate downscaling, local-law evaluation, residual assembly, and Newton update until convergence; then upscale \(\bar P\) and \(\bar L\); finally update texture through \(R_t^i\leftarrow \mathrm{polar}(F_e^i)\) [2502.02457].

The 2026 assessment describes two online IMN schemes for inelastic prediction: a fixed-point iteration that forward-homogenizes tangent stiffnesses and back-dehomogenizes strains until convergence, and a Newton iteration that solves directly for the “jumping” vector \(\mathbf a\). In the latter,
\[
\mathbf R=\mathbf A^T\mathbf W\mathbf K(\mathbf A\mathbf a+\boldsymbol\varepsilon^{\mathrm{macro}}),
\qquad
\mathbf J=\frac{d\mathbf R}{d\mathbf a}=\mathbf A^T\mathbf W\mathbf K\mathbf A.
\]
This formulation is used to compare online accuracy and uncertainty across multiple random initializations [2602.07192].

The cohesive IMN deploys a different nonlinear local law but follows the same offline/online division. After elastic-only training, the online stage attaches the Camacho–Ortiz irreversible mixed-mode cohesive law to each cohesive layer and solves the IMN implicitly for new loading histories, including loading–unloading, tension/compression, shear, and large local separations, over 300 time steps. The benchmark RVE is a UD composite with \(29.4\%\) fiber volume, elastic fiber \(E=500\) GPa, matrix either elastic \(E=100\) GPa or von Mises plasticity with \(\sigma_y=0.1\) GPa and piecewise hardening, and interface parameters \(\sigma_c=0.5\) GPa, \(G_c=2.5\times10^{-6}\) GPa·m, \(\beta=0.5\), \(\zeta=2\times10^{-5}\) GPa·s, \(L=2.5\) mm [1908.04858].

These online procedures define IMN as a reduced-order constitutive solver. The network is not only predicting homogenized quantities but also carrying microstructurally organized internal states, whether those states correspond to cohesive openings, crystal-plastic internal variables, or evolving orientations.

## 6. Accuracy, efficiency, extensions, and limitations

The reported performance of IMN depends on the physical setting and the specific variant. For polycrystals, the offline-trained ODMN, coupled with a phenomenological crystal-plasticity law per node with 12 FCC slip systems and generalized Hooke’s law, matches DNS under uniaxial, cyclic, and shear loading within \(<5\%\) RMSE even for \(N=5\), with errors dropping below \(2\%\) for \(N\ge 7\). At \(F_{11}=1.3\) under uniaxial loading, texture evolution pole figures reproduce the DNS-ODF with \(\hat T^d\approx 0.1\) for \(N=8\). The paper also states a speed-up over DNS of \(10^2\)–\(10^3\times\) even with pure-Python implementation, with an optimal trade-off at \(N=7\) and approximately 600 trainable parameters [2502.02457].

For interface-governed composites, the cohesive IMN reports a Stage II average test error on stiffness of \(\le 0.33\%\) for \(N=9,N_c=4\). In nonlinear extrapolation, the network captures debonding softening, unloading stiffness degradation, and re-hardening in compression within \(\pm5\%\) of DNS. The online cost for a 300-step RVE simulation drops from 5.6 h on 10 CPUs for DNS with cohesive elements to 33.3 s on a single CPU for IMN \(N=9,N_c=4\), corresponding to \(>6000\times\) speed-up; smaller networks give 5.8 s for \(N=7,N_c=4\) and 1.4 s for \(N=5,N_c=4\) [1908.04858].

The 2026 systematic assessment sharpens the comparison with the original DMN. Across depths \(N=4\dots 8\), IMN trains in only \(\tfrac1{3.4}\)–\(\tfrac1{4.7}\) of the time required by DMN. For \(N\ge 6\), both achieve comparable mean stress error \(e_\sigma\approx 2\%-3\%\) with nearly identical variance across initializations. In very shallow networks, \(N\le 5\), DMN is marginally more accurate by approximately \(0.5\%\). Online wall-clock times are essentially the same because DMN converges in fewer iterations, while IMN has a lower per-node, per-iteration cost of about \(1.9\times\)–\(2.4\times\) [2602.07192].

The most explicit limitation identified for ODMN is texture specificity. The later paper on texture-generalizable ODMN states that the original ODMN remains limited by the need to retrain for each distinct crystallographic texture. TACS-GNN-ODMN is proposed to remove this requirement by combining Texture-Adaptive Clustering and Sampling for initializing texture-related parameters with a GNN for predicting stress-equilibrium-related parameters. On validation, it reports an effective stiffness reconstruction error of approximately \(0.0205\), nonlinear cyclic-loading mean-relative error below \(2\%\), max-relative error below \(5\%\), texture-difference indices below \(0.12\), and CPU speed-ups of \(200\)–\(280\times\) relative to DAMASK-FFT for the stated loading cases [2512.06779].

A recurrent misconception is that the interaction directions in IMN are equivalent to crystallographic orientations. The ODMN results explicitly separate these roles: standard DMN cannot disentangle node orientation from stress-equilibrium directions and therefore cannot predict texture evolution, while interaction-based IMN without orientation captures equilibrium but not texture [2502.02457]. Another common misconception is that training on linear-elastic data implies linear-elastic use only. Across cohesive IMN, ODMN, and the systematic assessment, the stated workflow is the opposite: train on elastic data, then extrapolate to nonlinear, anisotropic, or path-dependent regimes during online prediction [1908.04858] [2602.07192].

The extension space is correspondingly broad but remains tied to the same architecture. Reported or proposed directions include other crystal systems such as HCP and BCC by supplying appropriate \(C^{\mathrm{crystal}}\) and slip systems, FE\(^2\) coupling with ODMN as the microscale solver at each Gauss point, enrichment of local laws to include damage or twinning, uncertainty quantification through Bayesian training of the weights \(z^i\), and inverse identification of interface laws from macroscopic experiments [2502.02457] [1908.04858]. Taken together, these developments position IMN as an interpretable reduced-order formalism for multiscale constitutive modeling in which the network architecture itself encodes the material-mechanics assumptions.

Source: https://www.emergentmind.com/topics/interaction-based-material-network-imn