---
title: 'MatPROV: Provenance Graph Dataset for Synthesis'
url: https://www.emergentmind.com/topics/matprov
type: topic
---

# MatPROV: Provenance Graph Dataset for Synthesis

MatPROV is a publicly released dataset of material-synthesis procedures cast as directed provenance graphs, built on the PROV-DM standard and extracted from scientific literature using large language models. It was introduced to represent synthesis knowledge as flexible, graph-structured causal records rather than as rigid, domain-specific schemas or linear sequences of operations. In this formulation, materials, tools, operations, and conditions are encoded in PROV-DM-compliant graphs that capture structural complexities and causal relationships among materials, operations, and conditions through visually intuitive directed graphs. The resulting representation is machine-interpretable and is intended to support applications such as automated synthesis planning, process optimization, machine learning, and causal reasoning in materials research [2509.01042].

## 1. Motivation and conceptual scope

MatPROV was created in response to two limitations identified in existing public datasets of synthesis recipes. First, many prior resources rely on rigid, domain-specific schemas, such as fixed JSON fields like “Metal_Source” or “Reaction_Time,” which do not generalize across materials classes. Second, many formulations assume that synthesis procedures are linear sequences of operations, which cannot represent branching, converging or re-entrant workflows common in real-world labs. MatPROV addresses these limits by adopting PROV-DM, described as a W3C provenance standard, to encode synthesis procedures as graph-structured causal records [2509.01042].

This design choice makes topology central rather than incidental. By using provenance, MatPROV can capture arbitrary topology, including branching and merging of intermediate steps, while retaining compatibility with a community standard. The emphasis on interoperability and extensibility is intrinsic to the project: no hand-crafted schema is forced beyond what PROV-JSONLD already provides, and the graph preserves whichever condition fields were reliably parsed from the source. A plausible implication is that MatPROV is designed less as a fixed benchmark artifact than as an extensible representation layer for materials process informatics [2509.01042].

A common misconception is to treat MatPROV as merely a conversion of laboratory text into ordered step lists. In the published formulation, a synthesis route is not fundamentally a sentence sequence but a directed provenance graph whose structure explicitly distinguishes inputs, operations, intermediates, and products. This makes the framework particularly suited to cases where causal structure cannot be recovered from textual order alone [2605.28487].

## 2. Formal graph model and PROV-DM compliance

PROV-DM defines three primary classes of nodes—Entity $(E)$, Activity $(A)$, and Agent $(Ag)$—and relations among them. In MatPROV, Agent is omitted and the representation focuses on Entity and Activity. The core sets and relations are given as follows:

$$
E := \{e_1, e_2, \dots\}
$$

$$
A := \{a_1, a_2, \dots\}
$$

with relations

$$
used(a,e) \subseteq A \times E
$$

$$
wasGeneratedBy(e,a) \subseteq E \times A
$$

$$
wasDerivedFrom(e_2,e_1) \subseteq E \times E
$$

where “Activity $a$ uses Entity $e$,” “Entity $e$ is generated by Activity $a$,” and “Entity $e_2$ is derived from $e_1$” respectively [2509.01042].

In the graph-theoretic restatement used in later work, a MatPROV record is a directed, typed provenance graph

$$
G = (V^m \cup V^t \cup V^a,\; E^u \cup E^g),
$$

where $V^m$ denotes material entities, $V^t$ tool entities, and $V^a$ activity nodes. Usage edges satisfy $E^u \subseteq (V^m \cup V^t) \times V^a$, and generation edges satisfy $E^g \subseteq V^a \times V^m$ [2605.28487].

The node and edge inventory is concise:

| Component | Meaning | Examples from the specification |
|---|---|---|
| Entity nodes | Materials or tools | precursors, intermediates, final products; experimental apparatus |
| Activity nodes | Experimental operations | “mixing,” “sintering” |
| Usage / Generation edges | Input and output provenance links | used(Activity, Entity); wasGeneratedBy(Entity, Activity) |

Every synthesis step is a small DAG whose nodes are either Entity nodes representing precursors, intermediates or final products with `type="material"`, or experimental apparatus with `type="tool"`, and Activity nodes representing experimental operations with labels in gerund form. Edges of type “Usage” correspond to `used(Activity, Entity)`, and edges of type “Generation” correspond to `wasGeneratedBy(Entity, Activity)` [2509.01042].

Serialization uses PROV-JSONLD. Every node or edge carries zero or more parameter attributes drawn from ten synthesis parameters—temperature, duration, pressure, mass, length, purity, concentration, rotation, atmosphere, and form—attached according to PROV-DM rules. Parameters characterizing an operation go on Activity nodes, while parameters describing an object’s attributes go on Entity nodes [2509.01042].

Later formalization makes the induced process semantics explicit. Precursors are materials with outgoing usage edges but no incoming generation edge; intermediates have both incoming generation and outgoing usage; products have incoming generation but no outgoing usage. Causal precedence between activities is induced by material flow: if $(a_i \to v)\in E^g$ and $(v \to a_j)\in E^u$, then $a_i$ must precede $a_j$. The full partial order over activities is obtained by transitive closure of these constraints and then realized via a topological sort, with original document ordering used to break ties [2605.28487].

An illustrative example in the original dataset description encodes sealing a copper sample in a silica tube before annealing as a graph in which the material “Cu” and the tool “silica tube” are both used by an Activity node labeled “sealing,” which generates a new material Entity labeled “sealed sample.” In simplified form, this is rendered as “Cu, silica tube —[sealing]→ sealed sample” [2509.01042].

## 3. Corpus construction and extraction pipeline

MatPROV’s 2 367 procedures were extracted from 1 568 open-access papers in Starrydata2, described as a curated database of functional-materials properties covering thermoelectrics, magnets, and batteries under CC BY 4.0. The extraction pipeline begins by downloading PDFs and converting them to XML via GROBID v0.8.2, retaining only body text. GPT-4o mini, with temperature $= 0.0$ and a few-hundred-token prompt, is then used to identify and extract paragraphs describing actual synthesis steps, excluding pure characterization or citations of prior work [2509.01042].

For each extracted text block, an LLM, empirically o4-mini, is prompted with the PROV-JSONLD schema: how to create nodes with `"@type": "Entity"` or `"@type": "Activity"`, how to label materials, tools, and gerund-form operations, how to create “Usage” and “Generation” edges, and how to attach parameters. In zero-shot mode, the model achieves reasonable accuracy; one-shot prompting with a single in-context example of a complex graph, exemplified by DOI `10.1002/advs.201901598`, further improves F1. Post-processing then merges identical nodes across steps, adds mandatory JSON-LD `"@context"` fields via rules, and ensures connectivity [2509.01042].

The broader graph-mining description adds several operational constraints. The automated pipeline parses each JSON-LD document into the graph $G$, mapping JSON-LD “entity,” “activity,” “wasUsedBy,” and “wasGeneratedBy” into the formal node and edge sets. It filters out any record lacking at least one precursor and one activity, validates consistency of typed fields, and discards highly incomplete or obviously contradictory extractions [2605.28487].

Validation combines manual annotation and automated matching. A domain expert manually annotated 44 procedures from 30 papers in PROV-JSONLD to form a ground truth. Matching between LLM outputs and ground truth is carried out by string-similarity on the top-level “label” field, defined as composition plus key operation, via `difflib.SequenceMatcher`, establishing one-to-one pairs for evaluation [2509.01042].

## 4. Dataset characteristics and empirical quality

The released corpus contains 2 367 procedures from 1 568 papers. Of these, 98.1 % are DAGs, 0.2 % are cyclic, and 1.7 % have isolated nodes. The graph-size distribution peaks at 5–20 nodes, with some graphs containing more than 30 nodes. The material-type distribution is reported as thermoelectrics 43.7 %, magnets 33.2 %, and batteries 11.4 %. Elemental coverage includes all non-noble and non-artificial elements, but is skewed toward Te, Bi, Sb, Pb, Se, and Sn on the thermoelectrics side, and Fe, Co, Ni, and Mn on the magnetic side [2509.01042].

Frequent activity co-occurrence exposes domain-specific “operation backbones.” For thermoelectrics, the typical workflow is described as weighing $\rightarrow$ mixing $\rightarrow$ pressing $\rightarrow$ sintering. For wet-chemistry magnets, the corresponding pattern is dissolving $\rightarrow$ adding $\rightarrow$ washing $\rightarrow$ drying. This suggests that MatPROV can capture both recurring procedural motifs and differences in synthetic regime across materials classes [2509.01042].

Extraction accuracy is measured at two levels. The structural level reports precision, recall, and F1 over correctly identified nodes and their connectivity. The parametric level reports precision, recall, and F1 over key-value parameter pairs on correctly extracted nodes. On a 25-paper test set, with five papers as one-shot examples and five runs each, zero-shot prompting gives the following metrics: for o4-mini, collection rate $0.832 \pm 0.059$, structural F1 $0.771 \pm 0.030$, and parametric F1 $0.748 \pm 0.053$; for GPT-4.1, collection rate $0.930$, structural F1 $0.697$, and parametric F1 $0.595$. One-shot prompting with the best example yielded structural +6.4 pp and parametric +7.6 pp improvement over zero-shot for o4-mini. Edge-level precision/recall and node-level breakdowns are reported to confirm that advanced reasoning capabilities are needed to capture complex branching correctly [2509.01042].

These statistics clarify both the strength and the current noise profile of the resource. The high DAG proportion indicates that the intended provenance structure is usually recovered, but the presence of cyclic graphs and isolated nodes shows that the release is not restricted to perfectly canonical extractions. A plausible implication is that MatPROV is useful both as a knowledge base and as a testbed for improving literature-to-graph extraction systems [2509.01042].

## 5. Role in reasoning benchmarks and process informatics

MatPROV later became the basis of MatProcBench, a provenance-grounded benchmark for process reasoning. MatProcBench instantiates seven multiple-choice tasks directly on top of MatPROV graphs: A1 route retrieval, A2 missing-step identification, A3 next-activity prediction, B1 condition prediction, B2 full condition-set prediction, C tool selection, and D process ordering. Distractor options are sampled from the corpus of MatPROV records so that they are locally plausible but globally incorrect. Because each gold answer can be deterministically read off the provenance graph, the benchmark provides an exact evaluation of a model’s ability to exploit causal and variable-bearing structure [2605.28487].

Evaluation is performed under four splits—random, material-type, year, and dual-OOD, defined as year + type—to probe generalization under scientific distribution shift. The strictest setting is described as a strict dual-OOD split that combines temporal and material-class shift. In that setting, the ProvMind framework, introduced as a process-memory reasoning framework, retrieves analogous training processes, converts them into provenance-aware option-level compatibility scores, and uses a language model for constrained final decision making. ProvMind achieves 52.84\% accuracy on the dual-OOD split, outperforming prompting, retrieval-augmented and supervised fine-tuning baselines [2605.28487].

The formalism used in this downstream work makes clear why MatPROV is suitable for reasoning tasks that are awkward in flat text representations. Retrieval scoring combines text, graph-structure, and provenance-heuristic views, and process-order validity is expressed through causal consistency over usage-generation pairs. This suggests that MatPROV is not only a repository of extracted procedures but also a substrate for symbolic-neural reasoning over routes, variables, tools, and causality [2605.28487].

## 6. Limitations, extensions, and name ambiguity

The published limitations are explicit. Scale is described as modest relative to text-mined collections, and future work is expected to expand beyond Starrydata2. Automated extraction still incurs errors, and fine-tuning LLMs or ensemble methods may raise recall and precision. The current graphs represent discrete procedural steps and conditions, but continuous dynamics, mechanistic kinetics, or thermodynamics are not encoded. The provenance schema is fixed to PROV-JSONLD, and richer ontologies such as Crystal Structure and Reaction Mechanisms are not yet integrated [2509.01042].

Several extensions are proposed. Adding Agents, for example catalysts or surfactants treated separately, or introducing more detailed equipment hierarchies could enrich the provenance semantics. Other directions include improving extraction precision via joint information-extraction and ontology alignment, integrating physically based simulation or thermodynamic constraints atop the discrete graph structure, and expanding to other domains such as organic synthesis and biological protocols by generalizing the node and edge schema [2509.01042].

The name “MatPROV” is also used in an unrelated line of work in formal methods. In that usage, MatPROV denotes a matroid-based automated prover for projective geometry, implemented as an external tool and integrated into Coq as a tactic through a plugin that calls the standalone C prover Bip and imports generated proof scripts. That system is based on saturation using matroid rules and targets projective incidence geometry, not materials synthesis provenance graphs [2107.05493]. This separate usage can cause bibliographic ambiguity, but it is conceptually distinct from the provenance-grounded materials dataset introduced in 2025 [2509.01042].

Source: https://www.emergentmind.com/topics/matprov