Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
80 tokens/sec
GPT-4o
59 tokens/sec
Gemini 2.5 Pro Pro
43 tokens/sec
o3 Pro
7 tokens/sec
GPT-4.1 Pro
50 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

Improving Precancerous Case Characterization via Transformer-based Ensemble Learning (2212.05150v1)

Published 10 Dec 2022 in cs.LG

Abstract: The application of NLP to cancer pathology reports has been focused on detecting cancer cases, largely ignoring precancerous cases. Improving the characterization of precancerous adenomas assists in developing diagnostic tests for early cancer detection and prevention, especially for colorectal cancer (CRC). Here we developed transformer-based deep neural network NLP models to perform the CRC phenotyping, with the goal of extracting precancerous lesion attributes and distinguishing cancer and precancerous cases. We achieved 0.914 macro-F1 scores for classifying patients into negative, non-advanced adenoma, advanced adenoma and CRC. We further improved the performance to 0.923 using an ensemble of classifiers for cancer status classification and lesion size named entity recognition (NER). Our results demonstrated the potential of using NLP to leverage real-world health record data to facilitate the development of diagnostic tests for early cancer prevention.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (8)
  1. Yizhen Zhong (4 papers)
  2. Jiajie Xiao (2 papers)
  3. Thomas Vetterli (2 papers)
  4. Mahan Matin (1 paper)
  5. Ellen Loo (1 paper)
  6. Jimmy Lin (208 papers)
  7. Richard Bourgon (2 papers)
  8. Ofer Shapira (5 papers)
Citations (1)