CataLM: Empowering Catalyst Design Through Large Language Models
Abstract: The field of catalysis holds paramount importance in shaping the trajectory of sustainable development, prompting intensive research efforts to leverage AI in catalyst design. Presently, the fine-tuning of open-source LLMs has yielded significant breakthroughs across various domains such as biology and healthcare. Drawing inspiration from these advancements, we introduce CataLM Cata}lytic LLM), a LLM tailored to the domain of electrocatalytic materials. Our findings demonstrate that CataLM exhibits remarkable potential for facilitating human-AI collaboration in catalyst knowledge exploration and design. To the best of our knowledge, CataLM stands as the pioneering LLM dedicated to the catalyst domain, offering novel avenues for catalyst discovery and development.
- Scibert: A pretrained language model for scientific text. arXiv preprint arXiv:1903.10676.
- Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
- Instructmol: Multi-modal integration for building a versatile and reliable molecular assistant in drug discovery.
- Large language model enhanced corpus of co2 reduction electrocatalysts and synthesis procedures. Scientific Data, 11(1):347.
- Matchat: A large language model and application service platform for materials science. Chinese Physics B, 32(11):118104.
- Electra: Pre-training text encoders as discriminators rather than generators. arXiv preprint arXiv:2003.10555.
- What would it take for renewably powered electrosynthesis to displace petrochemical processes? Science, 364(6438):eaav3506.
- Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805.
- Autodive: An integrated onsite scientific literature annotation tool. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations), pages 76–85.
- The pile: An 800gb dataset of diverse text for language modeling. arXiv preprint arXiv:2101.00027.
- Revisiting electrocatalyst design by a knowledge graph of cu-based catalysts for co 2 reduction. ACS Catalysis, 13:8525–8534.
- Neural network training method for materials science based on multi-source databases. Scientific Reports, 12(1):15326.
- Matscibert: A materials domain language model for text mining and information extraction. npj Computational Materials, 8(1):102.
- Lora: Low-rank adaptation of large language models.
- Commentary: The materials project: A materials genome approach to accelerating materials innovation. APL materials, 1(1).
- Protein function prediction as approximate semantic entailment. Nature Machine Intelligence, pages 1–9.
- A universal model for accurately predicting the formation energy of inorganic compounds. Science China Materials, 66(1):343–351.
- Progress and challenges toward the rational design of oxygen electrocatalysts based on a descriptor approach. Advanced Science, 7(1):1901614.
- Pymupdf. Available at http://pymupdf.readthedocs.io/en/latest/.
- Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692.
- Swarm intelligence for new materials. Computational Materials Science, 214:111699.
- LudiWang (2023). The extended corpus of CO2 reduction electrocatalysts and synthesis procedures.
- Biogpt: generative pre-trained transformer for biomedical text generation and mining. Briefings in bioinformatics, 23(6):bbac409.
- Towards the computational design of solid catalysts. Nature chemistry, 1(1):37–46.
- Open, A. (2022). Introducing chatgpt. open ai.
- OpenAI, R. (2023). Gpt-4 technical report. arxiv 2303.08774. View in Article, 2(5).
- A liver cancer question-answering system based on next-generation intelligence and the large model med-palm 2. International Journal of Computer Science and Information Technology, 2(1):28–35.
- Toolllm: Facilitating large language models to master 16000+ real-world apis. arXiv preprint arXiv:2307.16789.
- Improving language understanding by generative pre-training.
- Language models are unsupervised multitask learners. OpenAI blog, 1(8):9.
- Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of machine learning research, 21(140):1–67.
- Materials design and discovery with high-throughput density functional theory: the open quantum materials database (oqmd). Jom, 65:1501–1509.
- Combining theory and experiment in electrocatalysis: Insights into materials design. Science, 355(6321):eaad4998.
- A perovskite oxide optimized for oxygen evolution catalysis from molecular orbital principles. Science, 334(6061):1383–1385.
- A corpus of co2 electrocatalytic reduction process extracted from the scientific literature. Scientific Data, 10(1):175.
- Self-instruct: Aligning language models with self-generated instructions. arXiv preprint arXiv:2212.10560.
- Pmc-llama: Further finetuning llama on medical papers. arXiv preprint arXiv:2304.14454.
- Lu–h–n phase diagram from first-principles calculations. Chinese Physics Letters, 40(5):057401.
- Large language models as master key: Unlocking the secrets of materials science with gpt.
- Doctorglm: Fine-tuning your chinese doctor is not a herculean task. arXiv preprint arXiv:2304.01097.
- Huatuogpt, towards taming language model to be a doctor. arXiv preprint arXiv:2305.15075.
- Chatgpt chemistry assistant for text mining and prediction of mof synthesis. arXiv preprint arXiv:2306.11296.
- Chatgpt chemistry assistant for text mining and the prediction of mof synthesis. Journal of the American Chemical Society, 145(32):18048–18062. PMID: 37548379.
Paper Prompts
Sign up for free to create and run prompts on this paper.