Large Language Models with Retrieval-Augmented Generation for Zero-Shot Disease Phenotyping (2312.06457v1)

Published 11 Dec 2023 in cs.AI, cs.CL, and cs.IR

Abstract: Identifying disease phenotypes from electronic health records (EHRs) is critical for numerous secondary uses. Manually encoding physician knowledge into rules is particularly challenging for rare diseases due to inadequate EHR coding, necessitating review of clinical notes. LLMs offer promise in text understanding but may not efficiently handle real-world clinical documentation. We propose a zero-shot LLM-based method enriched by retrieval-augmented generation and MapReduce, which pre-identifies disease-related text snippets to be used in parallel as queries for the LLM to establish diagnosis. We show that this method as applied to pulmonary hypertension (PH), a rare disease characterized by elevated arterial pressures in the lungs, significantly outperforms physician logic rules ($F_1$ score of 0.62 vs. 0.75). This method has the potential to enhance rare disease cohort identification, expanding the scope of robust clinical research and care gap identification.

PDF HTML Abstract

Summarize PDF Markdown Bookmark Chat (Pro)

References (25)

Authors (12)

Will E. Thompson (4 papers)
David M. Vidmar (1 paper)
Jessica K. De Freitas (2 papers)
John M. Pfeifer (1 paper)
Brandon K. Fornwalt (5 papers)
Ruijun Chen (12 papers)
Gabriel Altay (14 papers)
Kabir Manghnani (3 papers)
Andrew C. Nelsen (1 paper)
Kellie Morland (1 paper)
Martin C. Stumpe (22 papers)
Riccardo Miotto (7 papers)

Citations (6)

View on Semantic Scholar

Large Language Models with Retrieval-Augmented Generation for Zero-Shot Disease Phenotyping (2312.06457v1)

Related Papers