---
title: Procedural Graph
url: https://www.emergentmind.com/topics/procedural-graph
type: topic
---

# Procedural Graph

A **procedural graph** is a graph-based representation of a procedure, process, executable program, or generative construction in which nodes denote actions, entities, operations, stages, parts, or intermediate states, and edges encode relations such as temporal precedence, data flow, dependency, grounding, transformation, spatial assembly, or conditional control flow. The term is not associated with a single formalism. Depending on the domain, a procedural graph may be a drainage forest for terrain synthesis, a heterogeneous entity graph for multimodal instructions, a typed DAG for material programs, a partial-order task graph inferred from demonstrations, or a business-process graph containing actors, gateways, constraints, and typed flows. Its common feature is the explicit representation of procedural structure rather than reliance on an unstructured sequence, image, mesh, or document embedding.

## 1. Conceptual scope and formal characteristics

Procedural graphs represent processes whose meaning depends on organization, execution, and transformation. A graph may encode a total sequence, a partial order, a branching workflow, a directed acyclic computation, or an evolving interaction structure. In procedural comprehension, for example, the relevant information includes which entities persist across steps, which actions transform them, and how textual and visual evidence correspond [2204.02566]. In task-graph learning, an edge can indicate that one action is a prerequisite of another, while independent actions remain unordered and may be executed in either order [2406.01486].

A generic procedural graph can be described as a structured object containing nodes, edges, attributes, and execution semantics:

$$
G=(V,E,\mathcal{A}_V,\mathcal{A}_E),
$$

where $V$ denotes procedural elements, $E$ denotes relations, and the attribute sets encode parameters, semantic labels, conditions, or physical properties. This notation is a general abstraction rather than a universal definition: individual systems use substantially different schemas.

The principal graph properties are:

- **Directedness**: edges may encode execution order, data dependencies, prerequisite relations, or state transformations.
- **Acyclicity**: many procedural programs and workflows are DAGs because they are evaluated in dependency order; some task systems assume acyclicity, while other procedural domains may require repetition or loops.
- **Typed nodes and edges**: nodes may represent actions, ingredients, locations, operators, parts, actors, gateways, or constraints; edges may represent sequence flow, condition flow, constraint flow, port connections, typed mates, or causal transitions.
- **Attributes**: nodes and edges can carry parameters, geometry, semantic types, reliability records, conditions, quantities, or visual features.
- **Executable semantics**: a graph may be directly evaluated, compiled into code, interpreted by a domain-specific generator, or used as an explicit reasoning structure.
- **Partial ordering**: procedures often admit multiple valid executions. A graph can therefore encode prerequisite relations without prescribing one unique sequence.

This distinguishes procedural graphs from ordinary concept graphs and document embeddings. A concept graph may represent co-occurrence or semantic relatedness, whereas a procedural graph represents how elements are organized for execution. The distinction is explicit in sparse directed Joint Concept Interaction Graphs, where direction is added to concept relations to preserve the flow of instructional documents [2402.03957].

## 2. Major graph formalisms

### Entity-centered multimodal graphs

Procedural Multimodal Documents pair textual instructions with images. The Temporal-Modal Entity Graph represents noun phrases, textual tokens, image regions, and instruction or image-level `[CLS]` nodes. Edges encode temporal evolution across steps, intra-modal relations, cross-modal grounding, and temporal or modal changes in relations [2204.02566].

Visual entities are obtained with Faster R-CNN, with at most 36 detected objects per image. Textual entities are identified from nouns and noun phrases using a POS tagger. Temporal visual links are approximated by Euclidean feature distance with threshold $\lambda_t=7$, while cross-modal links are based on grounding-box IoU with threshold $\lambda_m=0.5$. Relation types are injected into graph attention through learned temporal and modal biases.

This formulation is designed for visual cloze, visual coherence, and visual ordering. On RecipeQA, TMEG reports an average accuracy of 69.73, and on CraftQA an average of 49.54. Removing temporal encoding reduces the average to 67.63 on RecipeQA and 46.95 on CraftQA, indicating the importance of entity evolution. These values characterize one specific model and dataset rather than a general property of procedural graphs.

### Semantic and flow graphs for instructional question generation

Graph-guided question generation combines local AMR graphs with global action-flow graphs [2401.13594]. AMR graphs represent the semantics of individual instructions, including actions, arguments, locations, durations, quantities, manner, purpose, and modifiers. Action-flow graphs represent cross-step dependencies, object transformations, ingredient composition, intermediate mixtures, and temporal order.

The two graph levels serve different functions. AMR supports exhaustive selection of answer-bearing roles and subgraphs, while the flow graph supports questions about composition, preceding and succeeding actions, and action ordering. This separation allows semantic content selection to be graph-controlled while natural-language realization is delegated to AMR-to-question models or LLMs.

### Latent text-to-graph representations

Unsupervised Learning of Graph from Recipes learns a heterogeneous graph containing actions, ingredients, and locations without graph-edge annotations [2401.12088]. A BERT-based entity identifier proposes nodes, while a graph-structure encoder produces relation scores and converts them into a continuous adjacency matrix using a log-domain Sinkhorn–Knopp procedure. A two-layer GCN encodes the graph, and a Transformer decoder reconstructs the original recipe.

The objective combines node-type classification, graph-to-text reconstruction, and adjacency sparsity:

$$
\mathcal{L}_{tot}=\mathcal{L}_{gse}+\mathcal{L}_{gen}+\lambda\|\mathbf{A}\|_1.
$$

The graph is learned incrementally as instructions are processed. Persistent nodes represent recurring entities, including intermediate products such as *dough* or *mixture*. The approach is graph-unsupervised rather than annotation-free: the entity identifier uses sentence-level annotations from the Now You’re Cooking dataset.

### Procedural document graphs

PAGED defines procedural graphs in a business-process style [2408.03630]. Nodes include actors, actions, start and end nodes, exclusive gateways, inclusive gateways, parallel gateways, data constraints, and action constraints. Edges include sequence flows, condition flows, and constraint flows. Gateways occur in branch–merge pairs and encode XOR, OR, and AND execution semantics.

The benchmark contains 3,394 document–graph pairs, 36,537 actions, 22,775 actors, 7,024 exclusive gateways, 1,204 inclusive gateways, 2,050 parallel gateways, 36,438 sequence flows, 10,598 condition flows, and 5,807 constraint flows. Its evaluations show a distinction between textual element identification and graph construction: LLMs are comparatively effective at finding actions and actors, but gateway and flow construction remains difficult.

A related multi-agent framework, text2flow, separates structural and logical refinement [2601.19170]. A Graph Builder generates an initial graph, a Simulation Agent tests reachability and executable paths, and a Semantic Agent checks gateway logic against linguistic cues such as “otherwise,” “also,” and “at the same time.” Structural feedback is derived from 10,000 simulation trials per graph, while selected feedback is injected into subsequent prompts for at most two refinement iterations.

### Task graphs and partial-order activity models

Task Graph Maximum Likelihood learns weighted DAGs from action sequences rather than gold graph edges [2406.01486]. The graph contains key-step nodes augmented with artificial start and end nodes. An edge $K_i\rightarrow K_j$ means that $K_j$ is a prerequisite of $K_i$, so edge orientation is opposite to ordinary execution direction.

The model defines a continuous feasibility score based on the total edge weight from an unobserved action to already observed prerequisites. The resulting Task Graph Maximum Likelihood loss increases weights associated with observed predecessor relations and suppresses weights from future actions to the same history. Direct Optimization learns one adjacency matrix per procedure, while Task Graph Transformer predicts graph structure from textual or video embeddings.

On CaptainCook4D, Direct Optimization achieves precision 86.4, recall 89.7, and $F_1=87.8$, compared with $F_1=71.1$ for MSG$^2$. The learned graphs support previous-step identification, optional-step detection, mistake detection, missing-step prediction, future-step prediction, and online mistake detection. Their interpretation remains statistical: frequent execution order may be mistaken for necessary precedence.

## 3. Procedural graphs as executable programs

### Material node graphs

Procedural material graphs are directed acyclic multigraphs whose nodes are image operators and whose edges connect typed input and output slots [2207.01044]. Nodes can represent noise and pattern generators, filters, transforms, color operations, normal and height processing, and material-output operators. Parameters may be scalars, vectors, or variable-length arrays.

MatFormer factorizes generation into node types, heterogeneous parameters, and typed edges. Node sequences are serialized under traversal orders such as back-to-front breadth-first traversal, reverse traversal, front-to-back breadth-first traversal, and random topological ordering. Parameters are quantized into 32 levels, while edge generation uses a pointer mechanism over dynamically constructed input and output slots. Sampling masks prevent invalid parameter choices, wrong port directions, occupied inputs, and cycles.

The graph representation is executable: nodes are evaluated in topological order, unconnected inputs receive operator-specific defaults, and special output nodes define material channels. MatFormer reports its best performance with the canonical back-to-front ordering $\pi_r$, obtaining graph-statistics error 0.046 and FID 48.6.

Subsequent systems condition procedural material-graph generation on images, text, or partial graphs. “Generating Procedural Materials from Text or Image Prompts” uses CLIP-based conditioning and factorizes graph generation into node, edge, and global parameter sequences [2304.13172]. It applies semantic validity masks, candidate ranking, and differentiable optimization of continuous parameters. The cleaned dataset contains 4,667 graph topologies and 466,700 parameterized instances. The model reports a validity rate above 90 percent for generated graphs and supports image-conditioned, text-conditioned, unconditional, and partial-graph generation.

VLMaterial instead represents Blender material graphs as executable Python programs generated by a vision-language model [2501.18623]. A LLaVA-NeXT system with a CLIP ViT-L/14 encoder and LLaMA 3 8B decoder emits Blender code that creates nodes, assigns parameters, and links sockets. Validity is determined operationally by Blender execution and rendering tests. The curated dataset contains 1,640 usable programs, and augmentation produces more than 550,000 image–program pairs.

MultiMat extends this direction by supplying both compact textual graph programs and visual graph representations with intermediate node outputs [2509.22151]. It uses a Qwen-based vision-language model and incrementally validates generated Substance Designer nodes through transpilation and execution. Its constrained tree search rejects invalid continuations and backtracks by discarding $2^{i-1}$ recently generated nodes at backtracking iteration $i$. On unconditional generation, MultiMat Graph reports KID 2.365 and Node Error Ratio 15.024; on conditional generation without parameter optimization, it reports DreamSim 36.609, CLIP similarity 67.907, and style loss 3.178.

### Procedural geometry and 3D asset graphs

ProcGen3D treats a procedural graph as an intermediate representation between an image and a generator-specific 3D asset [2511.07142]. Nodes contain structural points or control locations with attributes such as coordinates, radii, and semantic classes. Edges encode branches, segments, members, or other structural relations and may contain force-related or semantic attributes.

The system tokenizes edges rather than isolated nodes:

$$
\tau(e_i)=\big(v_a,\mathcal{A}_{v_a},v_b,\mathcal{A}_{v_b},\mathcal{A}_{e_i}\big).
$$

DFS ordering is used for cacti and trees, while BFS ordering is used for bridges. An OPT-350M transformer predicts edge tokens conditioned on an RGB image. Monte Carlo Tree Search then explores graph continuations using a silhouette-overlap reward. The graph is decoded by the original procedural generator, which supplies detailed geometry, materials, textures, and category-specific structure.

Proc3D introduces the Procedural Compact Graph for language-driven, editable 3D modeling [2601.12234]. Its nodes encode primitives, transformations, geometry operations, arithmetic relations, switches, joins, and exposed parameters. Its edges encode geometry, scalar, vector, Boolean, and transform dependencies. The representation is interpreted into Blender, Unity3D, Substance 3D, or other runtimes. In reported comparisons, PCG achieves an 89 percent compile rate with GPT-4o, uses an average of 702 tokens, and reports edit times of approximately 0.01 seconds.

Procedura models an assembly graph whose nodes are named parametric CSG parts and whose typed edges are mates [2608.26238]. Mate types include bolt-pattern, peg-socket, seat-face, flange, tab-slot, press-fit, lip-rabbet, snap-tab, key, revolute, prismatic, and spherical. Deterministic frame-based placement solves a new part’s transform from existing interface frames:

$$
T_j=F_i\cdot\Delta(\phi)\cdot F_j^{-1}.
$$

A part is admitted only after compile, mate, and connectivity checks pass. This gives the graph machine-checkable assembly semantics in addition to geometric structure.

## 4. Procedural graphs for simulation and generative environments

### Terrain drainage graphs

In graph-based terrain generation, each terrain tile is also a graph node [2210.14496]. Nodes store land height, water height, total height, constraint values, constraint strengths, moisture, slope, drainage, and temporary gorge information. Each node selects at most one lower neighboring node from its four-connected Von Neumann neighborhood. The resulting graph is a forest of directed drainage trees rooted at local minima.

Drainage propagates recursively:

$$
D(t)=\mu_t+k_d\sum_{u\in\operatorname{Trib}(t)}D(u),
$$

where $k_d=0.68$ in the reported experiments. Fluvial erosion uses:

$$
\Delta h=k_eD^ns^m,
$$

with $k_e=0.5$, $n=1$, and $m=2$. Gorge formation connects local minima through graph-traversed paths, and constraint strengths are weakened along gorge paths so that carved channels persist.

This graph is not an abstract post-processing structure. It is dynamically rebuilt from the current height field at every simulation tick. The system uses graph traversal to propagate rainfall, identify tributaries and minima, select gorge paths, and apply erosion. The principal operations are approximately $O(N)$ per iteration, excluding implementation-dependent gorge-search overhead.

### Reinforcement-learning graph generation

G-PCGRL treats an extended adjacency matrix as the state of a Markov decision process [2407.10483]. Diagonal entries store node types or empty-node symbols, while off-diagonal entries encode undirected edges. A reinforcement-learning agent toggles edges until type-based adjacency constraints are satisfied.

The graph-specific representations are graph-narrow and graph-wide. Graph-narrow lets the environment select an edge and gives the agent only “toggle” or “do nothing” actions. Graph-wide provides the entire matrix and lets the agent choose both the edge position and toggle operation. PPO policies use multilayer perceptrons with layers $128\rightarrow256\rightarrow128$.

Constraints are declarative and existential, such as requiring every Source to connect to a Converter and every Converter to connect to a Source and Pool. The reward gives positive feedback for adding missing permitted edges or removing incorrect edges, negative feedback for the reverse, and a terminal-validity bonus. The method supports controllable node counts and node-type compositions but does not directly support directed edges, exact edge counts, acyclicity, global connectivity, arbitrary motifs, or weighted edges.

### Procedure graphs as evolving world models

ProPlay uses an attributed directed graph whose nodes are abstract procedures and whose edges represent causal transitions between task stages [2606.12780]. Each transition has a reliability-record embedding that accumulates successful, task-conditioned experience:

$$
c_{ij}^{k}=\sum_{r=1}^{k}\mathbb{I}\big((p_i,p_j)\in W^r\big)R^r\phi(d_{\tau_r}).
$$

Before an episode, known transitions are ranked by similarity between the current task description and their reliability records. An LLM generates a procedure-level preplay path, which is injected as soft guidance rather than enforced as a hard plan. After execution, productive trajectory prefixes are abstracted into procedures, failures are stored separately, and the graph is refined.

On ScienceWorld, ProPlay reports success rate 37.4 and average score 70.2, compared with 32.2 and 64.8 for the variant without transitions. The graph is therefore used as both memory and planning substrate. Its edges are hypotheses about reusable stage transitions, not deterministic environment dynamics.

## 5. Extraction, learning, and generation methodologies

Procedural graph systems differ primarily in how graphs are obtained and how validity is enforced.

### Rule-based and parser-based extraction

Rule-based systems typically identify actions, entities, or sentence relations and then construct graph edges with manually specified patterns. PAGED reports that earlier systems can identify local elements but perform poorly on inclusive gateways, parallel structures, constraints, and typed flows [2408.03630]. Their limitations arise because graph construction requires nonlocal logical organization rather than only span extraction.

AMR and flow-graph pipelines provide a more structured alternative. They parse local semantic roles and global dependencies separately, then use graph transformations and traversals to derive question semantics [2401.13594]. Their reliability depends on parser quality and manually designed transformations, particularly for ambiguous roles such as AMR `:ARG2`.

### Differentiable structure learning

Differentiable methods relax discrete graph construction into continuous matrices or edge weights. Recipe graph induction uses Sinkhorn-normalized soft adjacency matrices and reconstruction loss [2401.12088]. Task Graph Maximum Likelihood directly optimizes continuous edge weights using sequence likelihood [2406.01486]. Both approaches permit gradient-based learning without enumerating all discrete graph structures.

The principal caveat is semantic identifiability. A graph that reconstructs text or predicts observed order may not contain the intended causal relations. Multiple graph structures can support similar outputs, and frequent execution order can be mistaken for necessary dependency.

### Autoregressive graph generation

Material and procedural-program systems serialize graphs into sequences and generate node types, parameters, edges, or programs autoregressively. MatFormer factorizes node, parameter, and edge generation; conditional material systems use pointer networks and graph-aware parameter conditioning [2207.01044] [2304.13172]. ProcGen3D tokenizes edges, while VLMaterial and Proc3D use Python or compact procedural languages as intermediate representations [2501.18623] [2511.07142] [2601.12234].

Serialization order is a central modeling choice. Canonical or category-appropriate orders reduce ambiguity, while random topological orderings can degrade performance. Edge-based tokenization preserves endpoint and relation information, whereas compact programs expose high-level operations and parameters directly.

### Constrained decoding and external verification

Procedural graph generation frequently combines learned distributions with deterministic validity checks. Material systems mask invalid ports, parameters, cycles, and occupied inputs. MultiMat validates each intermediate graph by transpiling and executing it. Procedura uses compile, mate, and connectivity gates. text2flow uses simulation to identify unreachable or dead-end paths and semantic agents to diagnose gateway inconsistencies.

External verification is particularly important because procedural validity is stronger than syntactic well-formedness. A generated graph may compile but produce an empty material, disconnected assembly, unreachable workflow, or semantically incorrect branch structure.

### Search and refinement

Search methods address the mismatch between model likelihood and executable or visual quality. Material generation uses candidate sampling, rendering-based ranking, differentiable parameter optimization, MCMC, or constrained tree search. ProcGen3D uses MCTS with silhouette overlap to steer graph generation toward image alignment. Procedural-world systems use graph-based preplay and post-episode refinement.

These methods improve output selection but increase computation. For ProcGen3D, MCTS increases average inference from seconds to 24 minutes for trees and 42 minutes for pine trees. In material generation, candidate rendering and optimization similarly create a trade-off between quality and latency.

## 6. Applications, evaluation, and limitations

Procedural graphs support several classes of application:

- **Procedural comprehension**: answering questions about entities, ordering, transformations, and multimodal correspondence.
- **Question-generation data synthesis**: exhaustive generation of local and temporal QA pairs from graph structure [2401.13594].
- **Document similarity**: comparing instructional documents using concept order and directed sparse graphs [2402.03957].
- **Activity understanding**: predicting prerequisites, detecting mistakes, identifying missing steps, and anticipating future actions [2406.01486] [2502.17753].
- **Material authoring**: generating editable node graphs from images, text, partial graphs, or multimodal graph context.
- **3D modeling**: producing generator-compatible assets, parametric assemblies, and natural-language-editable procedural models.
- **Terrain generation**: simulating drainage, erosion, and gorge formation through graph traversal.
- **Game-content generation**: generating skill trees and game economies under declarative constraints.
- **World modeling**: organizing successful experience into reusable task-stage transitions.
- **Procedural stylometry**: characterizing creator-specific workflow topology and generating persona-conditioned procedures [2608.24369].

Evaluation is correspondingly heterogeneous. Graph extraction uses element, gateway, flow, and constraint F1; graph similarity uses accuracy and F1; task graphs use edge precision, recall, and $F_1$; materials use FID, KID, style loss, SWD, CLIP, DreamSim, graph statistics, and validity; 3D systems use Chamfer Distance, LPIPS, CLIP similarity, compile rate, and edit latency; terrain systems report runtime, interactivity, variety, realism, and independence.

Several limitations recur across domains:

- **Graph-construction ambiguity**: multiple graphs may explain the same text, image, or sequence.
- **Approximate semantics**: inferred directions, grounding links, object identities, and causal dependencies may be noisy.
- **Serialization dependence**: sequence order can introduce arbitrary biases and affect generation quality.
- **Domain-specific schemas**: procedural graphs often require a specialized ontology, generator, interpreter, or constraint language.
- **Incomplete validity guarantees**: execution checks establish practical validity but do not guarantee semantic correctness.
- **Sparse or biased data**: observed workflows may reflect common practice rather than all valid procedures.
- **Scalability**: graph size increases sequence length, action-space complexity, search cost, or simulation burden.
- **Limited handling of repetition and concurrency**: many systems assume DAGs, repetition-free sequences, or implicit rather than explicit parallelism.
- **Topology–appearance mismatch**: a visually similar output need not recover the original graph, and a structurally valid graph may produce poor visual or physical results.
- **Human interpretability**: explicit graphs are more inspectable than latent representations, but large generated graphs can remain difficult to understand.

The unifying research problem is therefore not simply graph generation or graph extraction. It is the construction of **valid, interpretable, executable, and domain-faithful procedural structures** from text, images, demonstrations, or design constraints. Across multimodal comprehension, program synthesis, activity modeling, terrain simulation, document analysis, and 3D authoring, procedural graphs provide an intermediate representation in which ordering, dependencies, transformations, constraints, and editable parameters become explicit. Their effectiveness depends on aligning graph semantics with the execution system that ultimately interprets them.

Source: https://www.emergentmind.com/topics/procedural-graph