Determine the real frontier of plausible and diverse retrosynthetic reactions

Determine the real frontier of chemical plausibility and diversity for reactions generated in one-to-many single-step retrosynthesis, so that future model-development efforts can be guided by attainable frontier values beyond those observed among the evaluated models.

Background

Single-step retrosynthesis is intrinsically one-to-many: a target molecule may admit multiple independent, chemically plausible disconnections. Consequently, ChemCensor-based Top-K benchmarks do not establish an absolute upper bound on the quality and diversity of possible predictions. The paper evaluates combinations of general-purpose LLMs, conventional SSRS models, and C3LM variants, but the best combined scores obtained from these models represent only the highest detectable frontier so far, rather than the true frontier.

The authors explicitly identify the unknown frontier as useful for navigating subsequent research. They note that the evaluated model pool still leaves room for improvement, particularly on metrics sensitive to both reaction diversity and plausibility.

References

While C3LM-LFM2-RFT-CC-NR achieves results on the challenging benchmark, the real frontier of chemical plausibility and diversity of generated reactions remains unknown due to the one-to-many nature of the ChemCensor-based benchmarks.

Training Chemical Plausibility-Aware Large Language Models for Single-Step Retrosynthesis  (2608.18940 - Zagribelnyy et al., 19 Aug 2026) in Section 4, subsection “Frontier of plausibility and diversity”