Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
41 tokens/sec
GPT-4o
59 tokens/sec
Gemini 2.5 Pro Pro
41 tokens/sec
o3 Pro
7 tokens/sec
GPT-4.1 Pro
50 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

Multi-domain Clinical Natural Language Processing with MedCAT: the Medical Concept Annotation Toolkit (2010.01165v2)

Published 2 Oct 2020 in cs.CL, cs.AI, and cs.LG

Abstract: Electronic health records (EHR) contain large volumes of unstructured text, requiring the application of Information Extraction (IE) technologies to enable clinical analysis. We present the open-source Medical Concept Annotation Toolkit (MedCAT) that provides: a) a novel self-supervised machine learning algorithm for extracting concepts using any concept vocabulary including UMLS/SNOMED-CT; b) a feature-rich annotation interface for customising and training IE models; and c) integrations to the broader CogStack ecosystem for vendor-agnostic health system deployment. We show improved performance in extracting UMLS concepts from open datasets (F1:0.448-0.738 vs 0.429-0.650). Further real-world validation demonstrates SNOMED-CT extraction at 3 large London hospitals with self-supervised training over ~8.8B words from ~17M clinical records and further fine-tuning with ~6K clinician annotated examples. We show strong transferability (F1 > 0.94) between hospitals, datasets, and concept types indicating cross-domain EHR-agnostic utility for accelerated clinical and research use cases.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (18)
  1. Zeljko Kraljevic (11 papers)
  2. Thomas Searle (10 papers)
  3. Anthony Shek (7 papers)
  4. Lukasz Roguski (6 papers)
  5. Kawsar Noor (5 papers)
  6. Daniel Bean (8 papers)
  7. Aurelie Mascio (4 papers)
  8. Leilei Zhu (2 papers)
  9. Amos A Folarin (13 papers)
  10. Angus Roberts (13 papers)
  11. Rebecca Bendayan (7 papers)
  12. Mark P Richardson (1 paper)
  13. Robert Stewart (19 papers)
  14. Wai Keong Wong (5 papers)
  15. Zina Ibrahim (17 papers)
  16. Richard JB Dobson (19 papers)
  17. Anoop D Shah (1 paper)
  18. James T Teo (3 papers)
Citations (139)