Consensus, dissensus and synergy between clinicians and specialist foundation models in radiology report generation (2311.18260v3)

Published 30 Nov 2023 in eess.IV, cs.CL, cs.CV, and cs.LG

Abstract: Radiology reports are an instrumental part of modern medicine, informing key clinical decisions such as diagnosis and treatment. The worldwide shortage of radiologists, however, restricts access to expert care and imposes heavy workloads, contributing to avoidable errors and delays in report delivery. While recent progress in automated report generation with vision-LLMs offer clear potential in ameliorating the situation, the path to real-world adoption has been stymied by the challenge of evaluating the clinical quality of AI-generated reports. In this study, we build a state-of-the-art report generation system for chest radiographs, $\textit{Flamingo-CXR}$, by fine-tuning a well-known vision-language foundation model on radiology data. To evaluate the quality of the AI-generated reports, a group of 16 certified radiologists provide detailed evaluations of AI-generated and human written reports for chest X-rays from an intensive care setting in the United States and an inpatient setting in India. At least one radiologist (out of two per case) preferred the AI report to the ground truth report in over 60$\%$ of cases for both datasets. Amongst the subset of AI-generated reports that contain errors, the most frequently cited reasons were related to the location and finding, whereas for human written reports, most mistakes were related to severity and finding. This disparity suggested potential complementarity between our AI system and human experts, prompting us to develop an assistive scenario in which Flamingo-CXR generates a first-draft report, which is subsequently revised by a clinician. This is the first demonstration of clinician-AI collaboration for report writing, and the resultant reports are assessed to be equivalent or preferred by at least one radiologist to reports written by experts alone in 80$\%$ of in-patient cases and 60$\%$ of intensive care cases.

Authors (26)

Ryutaro Tanno (36 papers)
David G. T. Barrett (16 papers)
Andrew Sellergren (8 papers)
Sumedh Ghaisas (2 papers)
Sumanth Dathathri (14 papers)
Abigail See (9 papers)
Johannes Welbl (20 papers)
Karan Singhal (26 papers)
Shekoofeh Azizi (23 papers)
Tao Tu (45 papers)
Mike Schaekermann (20 papers)
Rhys May (3 papers)
Roy Lee (2 papers)
SiWai Man (2 papers)
Zahra Ahmed (2 papers)
Sara Mahdavi (2 papers)
Danielle Belgrave (6 papers)
Vivek Natarajan (40 papers)
Shravya Shetty (21 papers)
Pushmeet Kohli (116 papers)

Citations (9)

View on Semantic Scholar

Summary

We haven't generated a summary for this paper yet.

Summarize Now

Tweets

https://twitter.com/RyutaroTanno/status/1749506026699714599

https://twitter.com/s0f1ra/status/1749891255159382118

https://twitter.com/DrStarson/status/1749976883700400596

https://twitter.com/joejanizek/status/1883624157956686169

Consensus, dissensus and synergy between clinicians and specialist foundation models in radiology report generation (2311.18260v3)

Summary

Related Papers

Tweets