---
title: Automated Material Model Discovery
url: https://www.emergentmind.com/topics/automated-material-model-discovery
type: topic
---

# Automated Material Model Discovery

Searching arXiv for recent papers on automated material model discovery and related constitutive discovery frameworks.
Automated material model discovery denotes the data-driven identification of constitutive laws, free-energy densities, dissipation potentials, or related predictive material relations directly from experiments or simulations while preserving interpretability and, in many formulations, embedding mechanical admissibility by design. In the recent literature, the topic spans sparse regression over curated feature libraries, constitutive neural networks that output thermodynamic potentials, grammar-based and genetic symbolic regression, unsupervised full-field inverse methods based on equilibrium, and differentiable finite-element model updating for history-dependent materials. Across these variants, the common objective is simultaneous model selection and parameter identification under explicit trade-offs between predictive accuracy, sparsity, physical consistency, and computational tractability [2310.06872], [2211.04453], [2505.07801].

## 1. Emergence of the field and its problem classes

Automated constitutive discovery has been framed as an inverse problem in which one seeks a constitutive law that maps deformation or strain histories to stresses in an interpretable and physically admissible form. For hyperelasticity, the target is typically a strain-energy density $W$ or $\psi$ from which stresses are obtained by differentiation; for inelasticity, the target extends to elastic and inelastic potentials together with evolution laws for internal variables [2210.02202], [2602.17750].

Several distinct problem classes now coexist. A first class is **local/direct discovery**, in which strain–stress pairs or stress–deformation curves are available, as in supervised hyperelastic characterization of brain tissue from uniaxial and torsional data [2305.16362] or local history-dependent fitting in ADiMU [2505.07801]. A second class is **global/indirect discovery**, in which stresses are not observed and the constitutive law is inferred from full-field displacement measurements and net reaction forces by enforcing equilibrium in weak form, as in EUCLID for plasticity and generalized standard materials [2202.04916], [2211.04453]. A third class is **symbolic or grammar-based discovery**, where the search variable is the mathematical expression itself rather than only coefficients in a fixed library [2402.04263], [2402.05238]. A fourth class is **physics-augmented neural discovery**, where trainable architectures encode kinematics, thermodynamics, convexity, or monotonicity and can subsequently be sparsified or symbolified [2210.02202], [2408.14615], [2602.17750].

This heterogeneity reflects the fact that the constitutive search space is strongly problem-dependent. Hyperelastic discovery often reduces to selecting functions of invariants or principal stretches, whereas elastoplastic and viscoelastic discovery must also identify internal variables, rate equations, and dissipation mechanisms. A plausible implication is that “automated material model discovery” is best understood not as a single algorithmic family but as a family of inverse constitutive identification strategies organized by data modality, physics assumptions, and representational bias.

## 2. Physical structure and admissibility constraints

A defining feature of the field is that successful methods do not treat constitutive modeling as unconstrained function approximation. They instead encode objectivity, symmetry, incompressibility or compressibility structure, and thermodynamic consistency either a priori in the representation or a posteriori through admissibility checks.

For hyperelasticity, the usual kinematic basis is the deformation gradient $F$, the right Cauchy–Green tensor $C = F^\top F$, the Jacobian $J = \det F$, and the invariants
$$
I_1 = \operatorname{tr}(C), \qquad
I_2 = \tfrac{1}{2}\big[(\operatorname{tr} C)^2 - \operatorname{tr}(C^2)\big], \qquad
I_3 = \det(C) = J^2.
$$
Energy-based formulations then define stresses by differentiation, for example
$$
P = \frac{\partial W}{\partial F}, \qquad
\sigma = \frac{1}{J} F \left(\frac{\partial W}{\partial F}\right)^\top,
$$
or, in incompressible settings,
$$
P = \frac{\partial W}{\partial F} - p F^{-T},
$$
with $p$ acting as a Lagrange multiplier enforcing $J=1$ [2310.06872], [2210.02202].

These constructions are not merely formal. They are used to guarantee that learned stresses remain derivable from a potential and therefore satisfy the Clausius–Duhem inequality for purely elastic processes when the potential is well posed. In constitutive neural networks, this is achieved by making the network output a scalar free energy rather than stress components directly, with inputs chosen as invariants or principal stretches so that frame indifference and isotropy are enforced by construction [2210.02202]. In grammar-based hyperelastic discovery, objectivity and isotropy are embedded by generating only expressions in the invariants of $C$ and by appending normalization corrections so that $W(I)=0$ and stress is near zero at the identity [2402.04263]. In symbolic regression for brain cortex, invariant-based models are constructed as sums of convex functions in invariants, while stretch- and strain-based models undergo posterior Hessian or ellipticity checks [2402.05238].

For inelastic materials, admissibility extends to dissipation. iCKAN formulates finite-strain inelasticity with the multiplicative decomposition $F = F_e F_i$, dual elastic and inelastic potentials, and a convexification operator applied to the learned inelastic potential so that the Clausius–Plank inequality reduces to a non-negative form $\mathcal{D} = \bar{\Sigma} : \bar{D}_i \ge 0$ [2602.17750]. Finite-strain elastoplastic PANNs similarly define free-energy and yield potentials so that each term entering the reduced dissipation is non-negative by construction [2408.14615]. EUCLID generalizes this logic to generalized standard materials by representing behavior through a convex Helmholtz free energy $\Psi$ and a convex dissipation potential $R$ or $R^\ast$, using convexity to guarantee stability and thermodynamic consistency [2211.04453].

This emphasis on hard constraints narrows the admissible hypothesis space. The literature repeatedly argues that such restriction is not a limitation but a precondition for robust extrapolation, sparse term selection, and physically interpretable parameter recovery [2310.06872], [2210.02202].

## 3. Representations used for discovery

The main representational choices can be organized by the object being discovered: coefficients in a predefined library, symbolic expressions generated from a grammar or genetic program, or trainable neural potentials that are later sparsified or symbolified.

| Representation | Core mechanism | Representative papers |
|---|---|---|
| Sparse library regression | Select a small subset of predefined constitutive features and fit coefficients | [2310.06872], [2305.16362], [2507.10196] |
| Symbolic generation | Generate candidate energy expressions by grammar or genetic programming | [2402.04263], [2402.05238] |
| Physics-augmented neural potentials | Learn $W$, $\Psi$, $\omega$, or hardening potentials with architectural constraints | [2210.02202], [2408.14615], [2602.17750], [2505.07801] |
| Full-field unsupervised inverse discovery | Infer constitutive law from equilibrium residuals without stress labels | [2202.04916], [2211.04453] |

In sparse library approaches, the dictionary is usually built from classical constitutive motifs. Constitutive neural networks in one influential formulation use eight functional building blocks in $(I_1-3)$ and $(I_2-3)$, including linear, quadratic, and exponential terms, so that neo-Hookean, Blatz–Ko, Mooney–Rivlin, Yeoh, and Demiray-type models become special cases of the same architecture [2210.02202]. Best-in-class modeling extends this idea to a library of sixteen building blocks, adding anisotropic terms in $I_4$ and $I_5$ and then performing bottom-up densification rather than top-down pruning [2404.06725]. Supervised EUCLID for brain tissue constructs an invariant library with generalized Mooney–Rivlin monomials and a $\log(I_2/3)$ feature, together with a dense generalized Ogden stretch library using fixed exponents $\alpha_i \in \{-100,\ldots,-0.01,0.01,\ldots,100\}$ [2305.16362].

Symbolic approaches broaden the hypothesis class. Formal grammars generate large libraries of admissible hyperelastic laws in Polish notation while biasing toward objectivity, isotropy, normalization, and growth properties [2402.04263]. Symbolic regression with genetic programming searches over invariant-based, principal-stretch-based, and normal-strain-based formulas and selects compact expressions by combining normalized fit error with expression-complexity control [2402.05238]. iCKAN occupies an intermediate position: the learned functions are represented by trainable B-splines inside a Kolmogorov–Arnold architecture, then converted to closed-form symbolic expressions by a post-training symbolification step [2602.17750].

A further representational distinction concerns linear versus nonlinear dependence on parameters. Principal-stretch Ogden-type networks with fixed exponents are linear in the weights and therefore reduce to convex least-squares problems with unique global minima for the chosen feature set [2310.06872]. In contrast, invariant-based networks with exponentials or free exponents induce non-convex optimization landscapes with multiple local minima [2310.06872], [2404.06725]. Much of the recent optimization literature is a response to this specific distinction.

## 4. Optimization, sparsity, and model selection

Sparsity control is central because constitutive models are expected to remain interpretable, robust in extrapolation, and computationally deployable. A standard objective is
$$
\min_\theta \sum_i \|y_i - f(x_i;\theta)\|^2 + \lambda \|\theta\|_p^p,
$$
with $p=2$ for ridge, $p=1$ for lasso, $p=0$ for explicit cardinality, and $0<p<1$ for fractional penalties [2310.06872].

A recurring conclusion is that the choice of penalty qualitatively changes the discovery outcome. In constitutive neural network studies, $L_2$ or ridge regularization is reported as unsuitable for model discovery because it stabilizes coefficients without inducing exact zeros; $L_1$ or lasso promotes sparsity but introduces strong bias; and $L_0$ provides the most transparent handle on the trade-off between interpretability and predictability, simplicity and accuracy, and bias and variance [2310.06872]. The same emphasis on non-smooth sparsity motivates later work devoted specifically to algorithms for minimizing objectives of the form $f(w)+\alpha\|w\|_1$ in constitutive discovery [2507.10196].

For quadratic mismatch functions, coordinate descent implements the classical LASSO, while LARS computes the entire piecewise-linear regularization path and identifies the critical values of $\alpha$ at which coefficients enter or leave the active set [2507.10196]. For non-quadratic or non-convex constitutive mappings, ISTA and pathwise ISTA are proposed as practical proximal-gradient schemes for obtaining approximate regularization paths [2507.10196]. These methods are particularly relevant when the model outputs depend nonlinearly on weights through exponentials or chain-rule stress maps.

Non-convexity has also led to a shift from top-down sparsification to bottom-up model growth. Best-in-class modeling starts from the best single-term model, then iteratively adds the term that most decreases the objective, thereby converting a combinatorial search such as $2^{16}=65{,}536$ possible term combinations into a sequence of smaller nonlinear fits [2404.06725]. Closely related guidance appears in the constitutive-network study, which recommends exact subset enumeration for one-term and two-term models and a bottom-up “densify” strategy for larger libraries [2310.06872].

Normalization is another technical theme. In multi-test hyperelastic fitting, stress residuals are normalized by the maximum recorded stress in each loading modality to prevent one mode from dominating the objective [2310.06872]. Supervised EUCLID for brain tissue similarly scales and concatenates uniaxial and torsional regression systems with empirically chosen weights $r_{UT}=0.3$ and $r_{ST}=1$ to equalize signal magnitudes before applying non-negative Lasso and Pareto-based model selection [2305.16362].

For full-field and history-dependent problems, optimization is inseparable from differentiable state updates and PDE solves. ADiMU backpropagates through return mapping, Newton iterations, and vectorized finite-element assembly, thereby enabling model updating for conventional, hybrid, and neural models without introducing extra hyperparameters beyond those intrinsic to the selected architecture and optimizer [2505.07801]. This suggests that automated material model discovery is increasingly converging with differentiable scientific computing rather than remaining a standalone sparse-regression problem.

## 5. Data modalities, benchmark domains, and representative discoveries

The field now covers a broad spectrum of datasets. Synthetic stress–deformation data remain the standard vehicle for verification because they allow exact benchmarking against known ground-truth models and parameters [2310.06872], [2402.05238]. Experimental applications include human brain gray and white matter, human brain cortex, VHB viscoelastic polymers, rubber, porcine skin, arteries, mild steel under cyclic loading, and soft-matter systems such as artificial meat [2305.16362], [2602.17750], [2404.06725], [2408.14615].

Several representative discoveries recur across studies. In human brain tissue, multiple approaches converge on strong $I_2$ dependence. Constitutive-network experiments on real brain data report that the best unregularized fit in the screened window often selects the Blatz–Ko term with $w_5 \approx 0.84$ and $w_1 \approx 0$, while the best one-term and two-term $L_0$ models are quadratic or exponential functions of $(I_2-3)$ [2310.06872]. Best-in-class modeling likewise identifies one-term and two-term models centered on $(I_2-3)$ and $(I_2-3)^2$ for gray and white matter [2404.06725]. Supervised EUCLID applied to 81 human-brain specimens finds that the most frequently discovered model is one-term Ogden, followed by two-term Ogden and mixed Ogden-plus-$(I_2-3)^2$ forms, with average standardized mean squared error about $0.17$ across specimens [2305.16362]. Symbolic regression on human brain cortex reaches a related conclusion in the invariant representation: all discovered invariant-based models are functions of $I_2$ only, including an optimal compact form
$$
W(I_2)=0.0170[\exp(27.91[I_2-3])-1]
$$
[2402.05238].

For viscoelastic and thermo-viscoelastic polymers, iCKAN demonstrates that symbolic inelastic discovery can recover explicit elastic and inelastic potentials from sequential finite-strain data. On VHB 4910, a two-branch Maxwell-like architecture discovers exponential-plus-polynomial elastic potentials and convexified inelastic potentials with training and test NMSE of approximately $7.6\times10^{-5}$ and $1.3\times10^{-3}$, while the symbolified versions retain similar accuracy [2602.17750]. On VHB 4905, temperature dependence is learned as explicit cubic or piecewise-quadratic functions $g(\theta)$ inside the elastic potentials, with training and test NMSE approximately $2.2\times10^{-4}$ and $6.5\times10^{-4}$ [2602.17750].

For cyclic metal plasticity at finite strain, physics-augmented neural networks discover hardening potentials that outperform classical Armstrong–Frederick and Ohno–Wang parameter fitting on both synthetic and experimental data. In the reported experimental mild-steel case, the 4NN model achieves best, mean, and standard-deviation losses of $0.0541$, $0.0598$, and $0.0036$, compared with $0.1547$, $10.031$, and $15.976$ for Armstrong–Frederick parameter fitting [2408.14615]. ADiMU further generalizes this differentiable-update paradigm to local stress–strain discovery and global full-field discovery for conventional, hybrid, and neural constitutive models, including elasto-plasticity and GRU-based surrogates with millions of parameters [2505.07801].

At the laboratory interface, not all automation concerns constitutive laws themselves. The AMT GUI exemplifies a neighboring workflow in which predictive surrogates and particle swarm optimization are used to recommend the next experiments from materials datasets without requiring programming expertise [2311.13808]. This is not constitutive discovery in the narrow mechanics sense, but it illustrates a broader trend: materials-model automation increasingly links model fitting, experiment design, and closed-loop data acquisition.

## 6. Limitations, controversies, and likely directions

The literature is explicit about unresolved limitations. One is combinatorial complexity. Exact $L_0$ subset search is tractable only for small libraries or very low cardinalities; larger spaces require heuristics such as forward densification, clustering, or latent-space search [2310.06872], [2404.06725]. A second is non-convexity: invariant-based exponential libraries, fractional penalties, neural hardening laws, and symbolic-generation pipelines all admit multiple local minima, so global optimality is typically not guaranteed [2310.06872], [2507.10196], [2402.04263].

A third limitation is identifiability under restricted loading. Multiple papers note that limited loading modes cannot identify all parameters; lack of shear can preclude inference of shear-related terms, and moderate-strain training data may be insufficient to excite the nonlinear structure of Ogden-type models, causing the discovery algorithm to settle on accurate surrogates rather than the literal ground-truth form [2310.06872], [2402.04263]. This also underlies some apparent controversies in the literature, such as whether brain tissue is best represented by invariant-based or stretch-based laws. The published evidence does not establish a universal winner; rather, it shows that different representations can fit the same dataset well while emphasizing different aspects of asymmetry, convexity, or extrapolation [2305.16362], [2402.05238].

A fourth issue concerns convexity, ellipticity, and extrapolation. Several methods bias toward polyconvexity or convexity but do not guarantee it globally [2310.06872], [2402.04263]. Stretch-based symbolic models for brain cortex can lose convexity under large deformations even when they pass training-range Hessian checks [2402.05238]. Best-in-class modeling explicitly notes that convexity and polyconvexity are not enforced [2404.06725]. This suggests that physical admissibility remains a spectrum: some frameworks guarantee it by architecture, while others rely on curated libraries and posterior checks.

Future directions described across the papers are comparatively aligned. They include extending hyperelastic discovery to viscoelasticity, plasticity, damage, growth, remodeling, anisotropy, orthotropy, and temperature-dependent behavior [2602.17750], [2408.14615], [2211.04453]. They also include uncertainty quantification through Bayesian or probabilistic regularization, stronger convexity guarantees through attribute grammars or convex architectures, and tighter integration with differentiable finite-element solvers and inverse design loops [2310.06872], [2402.04263], [2505.07801]. A plausible implication is that the field is moving toward unified pipelines in which constitutive structure, experiment design, and downstream simulation are co-optimized rather than treated as separate stages.

Automated material model discovery therefore sits at the intersection of continuum mechanics, sparse optimization, symbolic computation, and differentiable programming. Its distinctive claim is not merely that models can be fit automatically, but that the form of the constitutive law itself can be selected, constrained, and interpreted from data in a way that remains compatible with the mathematical and thermodynamic structure of materials theory [2210.02202], [2602.17750].

Source: https://www.emergentmind.com/topics/automated-material-model-discovery