Papers
Topics
Authors
Recent
Search
2000 character limit reached

MatPROV: Provenance Graph Dataset for Synthesis

Updated 9 July 2026
  • MatPROV is a dataset that models material synthesis procedures as PROV-DM-compliant causal graphs capturing complex, non-linear workflows.
  • It uses large language models to extract and structure synthesis steps from literature, overcoming limitations of rigid, linear schemas.
  • The graph-based design supports automated synthesis planning, process optimization, and causal reasoning for advanced materials research.

MatPROV is a publicly released dataset of material-synthesis procedures cast as directed provenance graphs, built on the PROV-DM standard and extracted from scientific literature using LLMs. It was introduced to represent synthesis knowledge as flexible, graph-structured causal records rather than as rigid, domain-specific schemas or linear sequences of operations. In this formulation, materials, tools, operations, and conditions are encoded in PROV-DM-compliant graphs that capture structural complexities and causal relationships among materials, operations, and conditions through visually intuitive directed graphs. The resulting representation is machine-interpretable and is intended to support applications such as automated synthesis planning, process optimization, machine learning, and causal reasoning in materials research (Tsuruta et al., 1 Sep 2025).

1. Motivation and conceptual scope

MatPROV was created in response to two limitations identified in existing public datasets of synthesis recipes. First, many prior resources rely on rigid, domain-specific schemas, such as fixed JSON fields like “Metal_Source” or “Reaction_Time,” which do not generalize across materials classes. Second, many formulations assume that synthesis procedures are linear sequences of operations, which cannot represent branching, converging or re-entrant workflows common in real-world labs. MatPROV addresses these limits by adopting PROV-DM, described as a W3C provenance standard, to encode synthesis procedures as graph-structured causal records (Tsuruta et al., 1 Sep 2025).

This design choice makes topology central rather than incidental. By using provenance, MatPROV can capture arbitrary topology, including branching and merging of intermediate steps, while retaining compatibility with a community standard. The emphasis on interoperability and extensibility is intrinsic to the project: no hand-crafted schema is forced beyond what PROV-JSONLD already provides, and the graph preserves whichever condition fields were reliably parsed from the source. A plausible implication is that MatPROV is designed less as a fixed benchmark artifact than as an extensible representation layer for materials process informatics (Tsuruta et al., 1 Sep 2025).

A common misconception is to treat MatPROV as merely a conversion of laboratory text into ordered step lists. In the published formulation, a synthesis route is not fundamentally a sentence sequence but a directed provenance graph whose structure explicitly distinguishes inputs, operations, intermediates, and products. This makes the framework particularly suited to cases where causal structure cannot be recovered from textual order alone (Zhang et al., 27 May 2026).

2. Formal graph model and PROV-DM compliance

PROV-DM defines three primary classes of nodes—Entity (E)(E), Activity (A)(A), and Agent (Ag)(Ag)—and relations among them. In MatPROV, Agent is omitted and the representation focuses on Entity and Activity. The core sets and relations are given as follows:

E:={e1,e2,… }E := \{e_1, e_2, \dots\}

A:={a1,a2,… }A := \{a_1, a_2, \dots\}

with relations

used(a,e)⊆A×Eused(a,e) \subseteq A \times E

wasGeneratedBy(e,a)⊆E×AwasGeneratedBy(e,a) \subseteq E \times A

wasDerivedFrom(e2,e1)⊆E×EwasDerivedFrom(e_2,e_1) \subseteq E \times E

where “Activity aa uses Entity ee,” “Entity (A)(A)0 is generated by Activity (A)(A)1,” and “Entity (A)(A)2 is derived from (A)(A)3” respectively (Tsuruta et al., 1 Sep 2025).

In the graph-theoretic restatement used in later work, a MatPROV record is a directed, typed provenance graph

(A)(A)4

where (A)(A)5 denotes material entities, (A)(A)6 tool entities, and (A)(A)7 activity nodes. Usage edges satisfy (A)(A)8, and generation edges satisfy (A)(A)9 (Zhang et al., 27 May 2026).

The node and edge inventory is concise:

Component Meaning Examples from the specification
Entity nodes Materials or tools precursors, intermediates, final products; experimental apparatus
Activity nodes Experimental operations “mixing,” “sintering”
Usage / Generation edges Input and output provenance links used(Activity, Entity); wasGeneratedBy(Entity, Activity)

Every synthesis step is a small DAG whose nodes are either Entity nodes representing precursors, intermediates or final products with type="material", or experimental apparatus with type="tool", and Activity nodes representing experimental operations with labels in gerund form. Edges of type “Usage” correspond to used(Activity, Entity), and edges of type “Generation” correspond to wasGeneratedBy(Entity, Activity) (Tsuruta et al., 1 Sep 2025).

Serialization uses PROV-JSONLD. Every node or edge carries zero or more parameter attributes drawn from ten synthesis parameters—temperature, duration, pressure, mass, length, purity, concentration, rotation, atmosphere, and form—attached according to PROV-DM rules. Parameters characterizing an operation go on Activity nodes, while parameters describing an object’s attributes go on Entity nodes (Tsuruta et al., 1 Sep 2025).

Later formalization makes the induced process semantics explicit. Precursors are materials with outgoing usage edges but no incoming generation edge; intermediates have both incoming generation and outgoing usage; products have incoming generation but no outgoing usage. Causal precedence between activities is induced by material flow: if (Ag)(Ag)0 and (Ag)(Ag)1, then (Ag)(Ag)2 must precede (Ag)(Ag)3. The full partial order over activities is obtained by transitive closure of these constraints and then realized via a topological sort, with original document ordering used to break ties (Zhang et al., 27 May 2026).

An illustrative example in the original dataset description encodes sealing a copper sample in a silica tube before annealing as a graph in which the material “Cu” and the tool “silica tube” are both used by an Activity node labeled “sealing,” which generates a new material Entity labeled “sealed sample.” In simplified form, this is rendered as “Cu, silica tube —[sealing]→ sealed sample” (Tsuruta et al., 1 Sep 2025).

3. Corpus construction and extraction pipeline

MatPROV’s 2 367 procedures were extracted from 1 568 open-access papers in Starrydata2, described as a curated database of functional-materials properties covering thermoelectrics, magnets, and batteries under CC BY 4.0. The extraction pipeline begins by downloading PDFs and converting them to XML via GROBID v0.8.2, retaining only body text. GPT-4o mini, with temperature (Ag)(Ag)4 and a few-hundred-token prompt, is then used to identify and extract paragraphs describing actual synthesis steps, excluding pure characterization or citations of prior work (Tsuruta et al., 1 Sep 2025).

For each extracted text block, an LLM, empirically o4-mini, is prompted with the PROV-JSONLD schema: how to create nodes with "@type": "Entity" or "@type": "Activity", how to label materials, tools, and gerund-form operations, how to create “Usage” and “Generation” edges, and how to attach parameters. In zero-shot mode, the model achieves reasonable accuracy; one-shot prompting with a single in-context example of a complex graph, exemplified by DOI 10.1002/advs.201901598, further improves F1. Post-processing then merges identical nodes across steps, adds mandatory JSON-LD "@context" fields via rules, and ensures connectivity (Tsuruta et al., 1 Sep 2025).

The broader graph-mining description adds several operational constraints. The automated pipeline parses each JSON-LD document into the graph (Ag)(Ag)5, mapping JSON-LD “entity,” “activity,” “wasUsedBy,” and “wasGeneratedBy” into the formal node and edge sets. It filters out any record lacking at least one precursor and one activity, validates consistency of typed fields, and discards highly incomplete or obviously contradictory extractions (Zhang et al., 27 May 2026).

Validation combines manual annotation and automated matching. A domain expert manually annotated 44 procedures from 30 papers in PROV-JSONLD to form a ground truth. Matching between LLM outputs and ground truth is carried out by string-similarity on the top-level “label” field, defined as composition plus key operation, via difflib.SequenceMatcher, establishing one-to-one pairs for evaluation (Tsuruta et al., 1 Sep 2025).

4. Dataset characteristics and empirical quality

The released corpus contains 2 367 procedures from 1 568 papers. Of these, 98.1 % are DAGs, 0.2 % are cyclic, and 1.7 % have isolated nodes. The graph-size distribution peaks at 5–20 nodes, with some graphs containing more than 30 nodes. The material-type distribution is reported as thermoelectrics 43.7 %, magnets 33.2 %, and batteries 11.4 %. Elemental coverage includes all non-noble and non-artificial elements, but is skewed toward Te, Bi, Sb, Pb, Se, and Sn on the thermoelectrics side, and Fe, Co, Ni, and Mn on the magnetic side (Tsuruta et al., 1 Sep 2025).

Frequent activity co-occurrence exposes domain-specific “operation backbones.” For thermoelectrics, the typical workflow is described as weighing (Ag)(Ag)6 mixing (Ag)(Ag)7 pressing (Ag)(Ag)8 sintering. For wet-chemistry magnets, the corresponding pattern is dissolving (Ag)(Ag)9 adding E:={e1,e2,… }E := \{e_1, e_2, \dots\}0 washing E:={e1,e2,… }E := \{e_1, e_2, \dots\}1 drying. This suggests that MatPROV can capture both recurring procedural motifs and differences in synthetic regime across materials classes (Tsuruta et al., 1 Sep 2025).

Extraction accuracy is measured at two levels. The structural level reports precision, recall, and F1 over correctly identified nodes and their connectivity. The parametric level reports precision, recall, and F1 over key-value parameter pairs on correctly extracted nodes. On a 25-paper test set, with five papers as one-shot examples and five runs each, zero-shot prompting gives the following metrics: for o4-mini, collection rate E:={e1,e2,… }E := \{e_1, e_2, \dots\}2, structural F1 E:={e1,e2,… }E := \{e_1, e_2, \dots\}3, and parametric F1 E:={e1,e2,… }E := \{e_1, e_2, \dots\}4; for GPT-4.1, collection rate E:={e1,e2,… }E := \{e_1, e_2, \dots\}5, structural F1 E:={e1,e2,… }E := \{e_1, e_2, \dots\}6, and parametric F1 E:={e1,e2,… }E := \{e_1, e_2, \dots\}7. One-shot prompting with the best example yielded structural +6.4 pp and parametric +7.6 pp improvement over zero-shot for o4-mini. Edge-level precision/recall and node-level breakdowns are reported to confirm that advanced reasoning capabilities are needed to capture complex branching correctly (Tsuruta et al., 1 Sep 2025).

These statistics clarify both the strength and the current noise profile of the resource. The high DAG proportion indicates that the intended provenance structure is usually recovered, but the presence of cyclic graphs and isolated nodes shows that the release is not restricted to perfectly canonical extractions. A plausible implication is that MatPROV is useful both as a knowledge base and as a testbed for improving literature-to-graph extraction systems (Tsuruta et al., 1 Sep 2025).

5. Role in reasoning benchmarks and process informatics

MatPROV later became the basis of MatProcBench, a provenance-grounded benchmark for process reasoning. MatProcBench instantiates seven multiple-choice tasks directly on top of MatPROV graphs: A1 route retrieval, A2 missing-step identification, A3 next-activity prediction, B1 condition prediction, B2 full condition-set prediction, C tool selection, and D process ordering. Distractor options are sampled from the corpus of MatPROV records so that they are locally plausible but globally incorrect. Because each gold answer can be deterministically read off the provenance graph, the benchmark provides an exact evaluation of a model’s ability to exploit causal and variable-bearing structure (Zhang et al., 27 May 2026).

Evaluation is performed under four splits—random, material-type, year, and dual-OOD, defined as year + type—to probe generalization under scientific distribution shift. The strictest setting is described as a strict dual-OOD split that combines temporal and material-class shift. In that setting, the ProvMind framework, introduced as a process-memory reasoning framework, retrieves analogous training processes, converts them into provenance-aware option-level compatibility scores, and uses a LLM for constrained final decision making. ProvMind achieves 52.84\% accuracy on the dual-OOD split, outperforming prompting, retrieval-augmented and supervised fine-tuning baselines (Zhang et al., 27 May 2026).

The formalism used in this downstream work makes clear why MatPROV is suitable for reasoning tasks that are awkward in flat text representations. Retrieval scoring combines text, graph-structure, and provenance-heuristic views, and process-order validity is expressed through causal consistency over usage-generation pairs. This suggests that MatPROV is not only a repository of extracted procedures but also a substrate for symbolic-neural reasoning over routes, variables, tools, and causality (Zhang et al., 27 May 2026).

6. Limitations, extensions, and name ambiguity

The published limitations are explicit. Scale is described as modest relative to text-mined collections, and future work is expected to expand beyond Starrydata2. Automated extraction still incurs errors, and fine-tuning LLMs or ensemble methods may raise recall and precision. The current graphs represent discrete procedural steps and conditions, but continuous dynamics, mechanistic kinetics, or thermodynamics are not encoded. The provenance schema is fixed to PROV-JSONLD, and richer ontologies such as Crystal Structure and Reaction Mechanisms are not yet integrated (Tsuruta et al., 1 Sep 2025).

Several extensions are proposed. Adding Agents, for example catalysts or surfactants treated separately, or introducing more detailed equipment hierarchies could enrich the provenance semantics. Other directions include improving extraction precision via joint information-extraction and ontology alignment, integrating physically based simulation or thermodynamic constraints atop the discrete graph structure, and expanding to other domains such as organic synthesis and biological protocols by generalizing the node and edge schema (Tsuruta et al., 1 Sep 2025).

The name “MatPROV” is also used in an unrelated line of work in formal methods. In that usage, MatPROV denotes a matroid-based automated prover for projective geometry, implemented as an external tool and integrated into Coq as a tactic through a plugin that calls the standalone C prover Bip and imports generated proof scripts. That system is based on saturation using matroid rules and targets projective incidence geometry, not materials synthesis provenance graphs (Magaud, 2021). This separate usage can cause bibliographic ambiguity, but it is conceptually distinct from the provenance-grounded materials dataset introduced in 2025 (Tsuruta et al., 1 Sep 2025).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to MatPROV.