Retro-Expert: Interpretable Retrosynthesis
- Retro-Expert is an interpretable framework that uses collaborative reasoning to predict reactants and provide natural-language chemical explanations.
- It integrates specialized shallow reasoning models with LLM-driven critical reasoning, achieving top-1 accuracies up to 66% and improved interpretability.
- The approach modularly preserves expert knowledge and employs reinforcement learning to optimize decision policies beyond static pattern matching.
Retro-Expert denotes, in its most specific use, the interpretable retrosynthesis framework introduced in "Retro-Expert: Collaborative Reasoning for Interpretable Retrosynthesis" (Li et al., 14 Aug 2025). In that formulation, retrosynthesis is treated not as static pattern matching but as collaborative reasoning that combines specialized models, a LLM, and reinforcement learning, with the system producing both reactant predictions and natural-language explanations grounded in chemical logic. In a broader arXiv usage, the term also appears descriptively for architectures that preserve, retrieve, or route expert behavior through external memory, modular experts, or expert banks rather than relying on a single monolithic model (Cervera, 15 Mar 2026, Zelezetsky et al., 26 Aug 2025, Zhu et al., 26 May 2026, Lu et al., 2 Jul 2026).
1. Definition and conceptual scope
Retro-Expert in chemistry is an interpretable retrosynthesis framework for inferring reactant molecules from a product molecule while exposing a reasoning path that identifies reaction types, reaction centers, candidate reactants, and the chemical rationale for choosing among them (Li et al., 14 Aug 2025). The framework is motivated by three limitations attributed to prior retrosynthesis systems: static pattern matching, black-box decision processes, and the lack of expert-aligned explanations. Its central claim is that retrosynthesis is naturally a reasoning process resembling the workflow of a synthetic chemist.
A common simplification is to treat Retro-Expert as an LLM-only retrosynthesis predictor. The framework is not presented that way. It combines specialized models that perform shallow reasoning, an LLM that performs critical-generative reasoning, and Knowledge-Guided Policy Optimization that optimizes the interpretable decision policy (Li et al., 14 Aug 2025). This division of labor is important because the specialized models provide a high-recall chemical decision space, while the LLM navigates that space and can generate alternatives when the supplied candidates are inadequate.
Outside chemistry, the same label has been used more loosely for systems that retrofit expert knowledge into retrieval-centered or routing-centered infrastructures. This suggests that "Retro-Expert" has become a family-resemblance term for architectures that preserve expert modularity and make expert behavior queryable at inference time rather than fully absorbing it into a single parametric model (Cervera, 15 Mar 2026, Zelezetsky et al., 26 Aug 2025, Zhu et al., 26 May 2026, Lu et al., 2 Jul 2026).
2. Collaborative reasoning architecture in retrosynthesis
The chemical Retro-Expert framework has three components. First, specialized models perform shallow reasoning to construct a high-quality chemical decision space. Second, an LLM performs critical reasoning to generate predictions and the corresponding interpretable reasoning path. Third, reinforcement learning optimizes the interpretable decision policy (Li et al., 14 Aug 2025).
In the main instantiation, the specialized models are T5Chem for reaction type prediction and GraphRetro for reaction center localization and reactant prediction. The LLM is Qwen2.5-7B-Instruct. Inputs include the product’s standard SMILES, atom-mapped SMILES, and IUPAC name. For each subtask, the specialized model produces a top- candidate set,
and the resulting chemical decision space is
In the main experiments, the subtasks are reaction type, reaction center, and reactant prediction, so the decision space is explicitly factorized rather than left implicit in end-to-end sequence generation (Li et al., 14 Aug 2025).
The LLM is prompted to reason in a structured format. It receives the product description, the most likely reaction type, the most likely reaction center, and a reactant candidate set. It outputs a > ... segment containing the reasoning process and an <answer>...</answer> segment containing only the predicted reactants. This format is not merely cosmetic: later reward computation includes an explicit format term, and the reasoning path is treated as part of the learned object (Li et al., 14 Aug 2025).
3. Chemical decision space, critical reasoning, and generation
Retro-Expert’s inference behavior is best understood as navigation over the decision space . The LLM constructs a reasoning path
which corresponds to one choice per subtask, together with a natural-language explanation and final reactant prediction (Li et al., 14 Aug 2025). The paper describes this as critical-generative reasoning. The critical component evaluates reaction type, reaction center, and reactant candidates against chemical logic; the generative component allows the model to produce a new reactant set if the candidate set is inadequate.
That distinction matters empirically. Retro-Expert is not restricted to selecting from specialized-model outputs. When all specialized model candidates are wrong, it can still generate novel correct reactants in of such cases (Li et al., 14 Aug 2025). This makes it a meta-reasoner rather than a reranker over static options.
The ablation study shows that the full multi-dimensional decision space is necessary. Using only reactant candidates yields top-1 accuracy; adding reaction type yields ; adding reaction center yields 0; and using reaction type, reaction center, and reactant candidates together yields 1 (Li et al., 14 Aug 2025). The sharp increase when reaction center information is included indicates that center localization carries much of the mechanistic structure that raw candidate lists do not provide.
The framework therefore departs from the common pattern in LLM retrosynthesis systems where explanations, when present, are post hoc. Here the reasoning path is part of the inference object and part of the training signal. A plausible implication is that interpretability and accuracy are coupled through the staged decision representation rather than traded against one another.
4. Knowledge-Guided Policy Optimization
Retro-Expert uses Knowledge-Guided Policy Optimization, based on Group Relative Policy Optimization, to optimize a policy 2 over outputs 3 conditioned on the query 4, where 5 is the product, 6 is the decision space, and 7 denotes external knowledge (Li et al., 14 Aug 2025). The objective is
8
with 9 in the reported training configuration (Li et al., 14 Aug 2025).
The reward is explicitly multi-stage: 0 where 1 scores the correctness of each subtask decision, 2 scores final reactant correctness, and 3 scores compliance with the required output format. The reported coefficients are 4, 5, and 6 (Li et al., 14 Aug 2025). The weighting makes intermediate reasoning correctness more important than format and even more heavily weighted than final-answer correctness alone.
Training uses a curated 9k-sample subset from USPTO-50K. The paper emphasizes a reward-hacking issue: specialized models often place the correct candidate at Top-1, so a naïve policy can learn a positional shortcut. To prevent this, the correct candidate is shuffled among the top three positions with a 7 probability distribution during training (Li et al., 14 Aug 2025). This is a concrete example of aligning the optimization target with reasoning rather than superficial regularities in the prompt.
5. Performance, expert alignment, and failure modes
On USPTO-50K, Retro-Expert reports a top-1 accuracy of 8, compared with 9 for Qwen2.5-7B-Instruct, 0 for Gemini-2.5-preview, 1 for GPT-3.5, and 2 for GPT-4o under the same decision-space assistance setting (Li et al., 14 Aug 2025). It also reports BLEU 3, Levenshtein distance 4, validity 5, MACCS 6, RDK 7, and Morgan 8 (Li et al., 14 Aug 2025).
When used collaboratively with specialized retrosynthesis models, Retro-Expert improves top-1 accuracy across all reported backbones: LocalRetro from 9 to 0, GLN from 1 to 2, GraphRetro from 3 to 4, RetroPrime from 5 to 6, Graph2Edits from 7 to 8, Retroformer from 9 to 0, and UAlign from 1 to 2 (Li et al., 14 Aug 2025). These gains indicate that the framework functions as a collaborative meta-reasoner rather than merely replacing specialized models.
On the ChemBench out-of-distribution benchmark, Retro-Expert reaches 3 top-1 accuracy, compared with 4 for Qwen2.5-7B-Instruct, 5 for ChemLLM, and 6 for DeepSeek-R1 (Li et al., 14 Aug 2025). The OOD result is central to the paper’s argument that the learned policy is not only memorizing frequent transformation templates.
Interpretability is evaluated both automatically and by chemists. Under GPT-4o evaluation, Retro-Expert scores 7 in Mechanism Accuracy, 8 in Factual Correctness, and 9 in Logical Consistency; under human evaluation by three synthetic organic chemists, it scores 0, 1, and 2 on the same metrics, each higher than the base Qwen2.5-7B system (Li et al., 14 Aug 2025). The framework therefore aims to bridge predictive performance and expert trust, not only performance and post hoc explanation.
The paper also reports wet-lab corroboration. It describes the first reported synthesis of 3-(2-ethoxyphenyl)thiophene via Suzuki coupling and a novel Jones oxidation route for 1-(4-ethoxyphenyl)ethanone as cases where Retro-Expert’s predictions were successfully executed in the lab (Li et al., 14 Aug 2025).
The reported limitations are chemically specific. Failure cases include molecules with multiple similar reactive sites and reactions where a broad reaction type admits multiple possible transformation pathways (Li et al., 14 Aug 2025). These are not generic LLM hallucinations; they are failures of fine-grained site discrimination and path ranking inside a chemically plausible neighborhood.
6. Broader uses of “Retro-Expert” in adjacent literatures
Several arXiv papers use "Retro-Expert" descriptively for modular systems that preserve expert structure and expose it through retrieval, routing, or expert selection. The usages are not identical, but they share a retrofit pattern: externalize expertise, keep it queryable, and add a coordination mechanism above it.
| Usage | Core mechanism | Paper |
|---|---|---|
| Expert knowledge preservation | Retrieval-centered expert knowledge infrastructure with multimodal capture, vector storage, and a conversational interface | (Cervera, 15 Mar 2026) |
| Offline RL with scarce demonstrations | Associative Memory Buffer populated by expert trajectories and queried during training and evaluation | (Zelezetsky et al., 26 Aug 2025) |
| Memory-constrained MoE inference | Router fine-tuning to boost short-horizon expert reuse and cache locality | (Zhu et al., 26 May 2026) |
| Replay-free continual ECG deployment | Frozen backbone plus per-source expert bank and lightweight router with top-2 margin fusion | (Lu et al., 2 Jul 2026) |
In "Expert Mind," the term is explicitly generalized into a retrieval-centered expert knowledge infrastructure in which LLMs sit on top of a curated, continuously updated expert memory (Cervera, 15 Mar 2026). In "Re:Frame," the idea appears as retroactive use of expert trajectories stored in an Associative Memory Buffer; using as few as 60 expert trajectories, corresponding to 3 of a 6000-trajectory dataset, improves a Decision Transformer baseline by up to 4 normalized points in three of four D4RL settings (Zelezetsky et al., 26 Aug 2025). In "ReMoE," a “Retro-Expert” router is effectively one that favors recently used experts; the reported system improves expert reuse by 5 while maintaining downstream task performance (Zhu et al., 26 May 2026). In replay-free continual ECG deployment, the analogous pattern is a frozen expert bank over ECGFounder features: source-aware expert selection reaches 6 Macro-F1, whereas an autonomous MLP router with top-2 margin fusion reaches 7, leaving autonomous source inference as the main bottleneck (Lu et al., 2 Jul 2026).
Across these uses, a consistent motif is visible. This suggests that "Retro-Expert" has come to denote not one architecture but a design stance: preserve expert structure, expose it through memory or modular experts, and learn when or how to consult it. In chemistry, that stance yields an interpretable collaborative reasoning system for retrosynthesis; in other domains, it yields retrieval systems, expert-memory RL, locality-aware MoE routing, or replay-free expert banks.