---
title: 'AceWGS: LLM Catalyst Design Framework'
url: https://www.emergentmind.com/topics/acewgs
type: topic
---

# AceWGS: LLM Catalyst Design Framework

AceWGS is an open-source framework leveraging large language models (LLMs) to accelerate catalyst design for the water-gas shift (WGS) reaction, a core process in low-temperature hydrogen production. The system bridges gaps in existing AI-based catalyst discovery by supporting end-to-end workflows that integrate textual and numerical information, enable multi-modal query interactions, and facilitate inverse catalyst design rooted in chemical theory. AceWGS is architected for extensibility, enabling adaptation across catalysis domains where literature comprehension and data-driven optimization intersect [2503.05607].

## 1. System Architecture and Components

AceWGS employs a modular architecture combining natural language processing (NLP), database management, retrieval-augmented generation (RAG), and numerical optimization. At the interaction layer, a Python 3.11/Tkinter graphical interface captures free-form researcher queries. A rule-based “Switch” module assigns each query to one of four processing features: (1) general LLM-based Q&A, (2) structured WGS literature database extraction via a Pandas agent, (3) passage-level article comprehension, or (4) AI-aided inverse catalyst design.

Relevant literature (82 PDF journal articles) undergoes dual indexing. High-level metadata populates a Pandas DataFrame for structured queries, while full texts are embedded (OllamaEmbeddings, "mxbai-embed-large") and stored in a FAISS vector database for semantic chunk retrieval. The context comprehension pipeline employs a VectorRetriever to select top-k relevant text windows for integration into LLM prompts. For catalyst candidate identification, a GUI collects design constraints, an inverse modeling engine (integrating a neural forward predictor and particle swarm optimization, PSO) explores the composition–preparation–condition space, and a prompt-guided LLM translates results into human-readable synthesis instructions.

## 2. Data Curation and Processing Pipelines

Textual information flows through a multi-stage ingestion and indexing pipeline. PDFMiner extracts raw text, LangChain’s RecursiveCharacterTextSplitter generates overlapping chunks (1,000-character length, 150-character overlap), which are embedded (768–1,024 dimensions) and indexed in FAISS for real-time similarity search. Query-time retrieval returns the top-N most relevant passages for RAG prompting.

The numerical pipeline manages a tabular database of 8,908 catalyst records with 99 features, encompassing composition percentages, preparation methods (binary-encoded), and reaction conditions. Candidate vectors $x \in \mathbb{R}^n$ are normalized before evaluation by the forward model $f_\theta(x)$, which predicts CO conversion. PSO operates within researcher-defined bounds, sampling $x$ values and applying theoretical and empirical constraints during search.

## 3. Inverse Modeling and Mathematical Formalism

The inverse design module targets the compositional and operational variables underpinning WGS catalyst performance:

- Metal/promoter/support fractions: $m \in [0,1]^d$, $\sum_k m_k = 1$
- Preparation method: $b \in \{0,1\}^p$ (one-hot encoding)
- Continuous conditions: $c \in \mathbb{R}^k$ (temperature, space velocity, feed)

The forward model maps $x$ to predicted CO conversion $\hat{y} = f_\theta(x) \in [0,100]$, constrained by a thermodynamic equilibrium bound $\hat{y}_{\text{eq}} = g(x)$. Training loss incorporates both RMSE and a soft constraint enforcing physical plausibility:
$$
L(\theta) = \text{RMSE}(\hat{y}, y_{\text{true}}) + \lambda \cdot \text{ReLU}(\hat{y} - \hat{y}_{\text{eq}})
$$

The inverse design task is:
$$
\max_x\; f_\theta(x) \\
\text{subject to:} \;
0 \leq m_k \leq 1, \; \sum_k m_k = 1; \;
b \in \{0,1\}^p \text{ (one-hot)}; \;
c_{\text{lower}} \leq c \leq c_{\text{upper}}; \;
f_\theta(x) \leq g(x)
$$
PSO updates particles by canonical rules:
$$
v_i(t+1) = w v_i(t) + c_1 \cdot \mathrm{rand}() \cdot (p\mathrm{best}_i - x_i(t)) + c_2 \cdot \mathrm{rand}() \cdot (\mathrm{gbest} - x_i(t)) \\
x_i(t+1) = x_i(t) + v_i(t+1)
$$
The objective function $J(x)$ incorporates penalties:
$$
J(x) = f_\theta(x) - \mu \max[0, f_\theta(x) - g(x)] - \sum_k \gamma_k\, \mathrm{ReLU}(m_k - 1) - \ldots
$$

## 4. Retrieval-Augmented LLM Integration

AceWGS relies on Ollama-hosted local LLMs (Llama3-8B, Llama3.1, Llama2, Gemma2) orchestrated by LangChain for multi-step prompting and tool use. Task-specific prompt templates enforce factuality and suitability for domain-specific questions. For general queries, the LLM is prompted as an "expert in catalysis and AI." Structured extraction tasks dispatch DataFrame-filtering instructions to the pandas agent, while comprehension prompts embed retrieved document chunks and chemistry questions. Inverse modeling output is translated to lab-usable instructions in plain scientific English.

All LLM inference procedures use zero-shot or few-shot prompting, with temperature=0, top_k=10, and top_p=0.5 to bias toward correctness. No full model fine-tuning was performed in the reported implementation.

## 5. Case Study: Accelerating Low-Temperature WGS Catalyst Design

A demonstration objective was to design a noble-metal WGS catalyst achieving $\geq 95\%$ CO conversion at $T < 300^\circ$C. Database extraction surfaced two relevant 2021 studies (R71: Pt/α-MoC, IWI; R51: Au/α-MoC, ALD). Article comprehension of R71 detailed a multi-step synthesis: α-MoC preparation by thermolysis/reduction, followed by incipient wetness impregnation of Pt, drying, and reduction/activation at $623–863$ K.

General query mode proposed a Pt–Au alloy on α-MoC with tunable size distributions. The inverse design module output an optimal recipe: Pt 4.26 wt%, Au 3.09 wt%, α-MoC 92.64 wt%; reaction conditions at 200°C, feed: CO 0.1%, H₂O 6.18%, CO₂ 5%, H₂ 0.15%, N₂ 88.57%, W/F = 1 mg·min/mL, 1 h on stream, with predicted CO conversion $95.07 \pm 0.79\%$. Single-metal α-MoC catalysts typically yielded 85–90% at these conditions, indicating a 5–10% improvement.

## 6. Open-Source Implementation and Adaptability

AceWGS is implemented entirely in Python with open-source components. Core modules include:

| Feature            | Code Module                   | Functionality Description                                   |
|--------------------|------------------------------|------------------------------------------------------------|
| General query      | feature1_general_query.py     | LLM Q&A chain                                              |
| Data extraction    | feature2_dataframe_agent.py   | DataFrame agent, executes queries on structured metadata    |
| Article comprehension | feature3_rag_comprehension.py | Chunk/retrieve text, RAG LLM chain                       |
| Inverse design     | feature4_inverse.py           | GUI collector, PSO optimization, LLM report synthesis      |

Embeddings are produced by OllamaEmbeddings ("mxbai-embed-large"). All databases and interfaces—Pandas, PDFMiner-processed full texts, FAISS retrieval—are swappable for new catalysis domains. LangChain’s "Switch" logic and prompting infrastructure support rapid domain adaptation.

AceWGS demonstrates the potential of retrieval-augmented LLMs to integrate unstructured text, structured data, and numerical optimization, enabling seamless, cross-modal catalyst discovery while remaining accessible and extensible for broader catalysis and chemical design research [2503.05607].

Source: https://www.emergentmind.com/topics/acewgs