Papers
Topics
Authors
Recent
Search
2000 character limit reached

Identifying and Aligning Medical Claims Made on Social Media with Medical Evidence

Published 18 May 2024 in cs.CL and cs.SI | (2405.11219v1)

Abstract: Evidence-based medicine is the practice of making medical decisions that adhere to the latest, and best known evidence at that time. Currently, the best evidence is often found in the form of documents, such as randomized control trials, meta-analyses and systematic reviews. This research focuses on aligning medical claims made on social media platforms with this medical evidence. By doing so, individuals without medical expertise can more effectively assess the veracity of such medical claims. We study three core tasks: identifying medical claims, extracting medical vocabulary from these claims, and retrieving evidence relevant to those identified medical claims. We propose a novel system that can generate synthetic medical claims to aid each of these core tasks. We additionally introduce a novel dataset produced by our synthetic generator that, when applied to these tasks, demonstrates not only a more flexible and holistic approach, but also an improvement in all comparable metrics. We make our dataset, the Expansive Medical Claim Corpus (EMCC), available at https://zenodo.org/records/8321460

Authors (2)
Definition Search Book Streamline Icon: https://streamlinehq.com
References (50)
  1. Knowledge Graph Based Synthetic Corpus Generation for Knowledge-Enhanced Language Model Pre-training. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3554–3565, Online. Association for Computational Linguistics.
  2. BERT Goes Brrr: A Venture Towards the Lesser Error in Classifying Medical Self-Reporters on Twitter. In Proceedings of the Sixth Social Media Mining for Health (#SMM4H) Workshop and Shared Task, pages 58–64, Mexico City, Mexico. Association for Computational Linguistics.
  3. Contextual String Embeddings for Sequence Labeling. In Proceedings of the 27th International Conference on Computational Linguistics, pages 1638–1649, Santa Fe, New Mexico, USA. Association for Computational Linguistics.
  4. COMETA: A Corpus for Medical Entity Linking in the Social Media. Publisher: arXiv Version Number: 2.
  5. Seventy-Five Trials and Eleven Systematic Reviews a Day: How Will We Ever Keep Up? PLoS Medicine, 7(9):e1000326.
  6. Erdenebileg Batbaatar and Keun Ho Ryu. 2019. Ontology-Based Healthcare Named Entity Recognition from Twitter Messages Using a Recurrent Neural Network Approach. International Journal of Environmental Research and Public Health, 16(19):3628.
  7. #WhyWeTweetMH: Understanding Why People Use Twitter to Discuss Mental Health Problems. Journal of Medical Internet Research, 19(4):e107.
  8. DreamDrug - A crowdsourced NER dataset for detecting drugs in darknet markets. In Proceedings of the Seventh Workshop on Noisy User-generated Text (W-NUT 2021), pages 137–157, Online. Association for Computational Linguistics.
  9. Robert M Bond and R Kelly Garrett. 2023. Engagement with fact-checked posts on Reddit. PNAS Nexus, 2(3):pgad018.
  10. Improving reference prioritisation with PICO recognition. BMC Medical Informatics and Decision Making, 19(1):256.
  11. Samir Chabou and Michal Iglewski. 2018. Combination of conditional random field with a rule based method in the extraction of PICO elements. BMC Medical Informatics and Decision Making, 18(1):128.
  12. Using Social Media for Actionable Disease Surveillance and Outbreak Management: A Systematic Literature Review. PLOS ONE, 10(10):e0139701.
  13. Mining Social Media Data for Biomedical Signals and Health-Related Behavior. Annual Review of Biomedical Data Science, 3(1):433–458.
  14. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. ArXiv:1810.04805 [cs].
  15. FuzzyBIO: A Proposal for Fuzzy Representation of Discontinuous Entities. In Proceedings of the 12th International Workshop on Health Text Mining and Information Analysis, pages 77–82, Louhi. Association for Computational Linguistics.
  16. Learning Dense Representations for Entity Retrieval. In Proceedings of the 23rd Conference on Computational Natural Language Learning (CoNLL), pages 528–537, Hong Kong, China. Association for Computational Linguistics.
  17. Xiaoli Huang and Jimmy Lin. 2006. Evaluation of PICO as a Knowledge Representation for Clinical Questions. American Medical Informatics Association, 2006:359.
  18. Gautier Izacard and Edouard Grave. 2021. Leveraging Passage Retrieval with Generative Models for Open Domain Question Answering. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 874–880, Online. Association for Computational Linguistics.
  19. Kalervo Jirvelin and Jaana Kekiiliinen. 2017. IR evaluation methods for retrieving highly relevant documents. ACM SIGIR Forum, 51(2).
  20. SpanBERT: Improving Pre-training by Representing and Predicting Spans. Transactions of the Association for Computational Linguistics, 8:64–77.
  21. Cadec: A corpus of adverse drug event annotations. Journal of Biomedical Informatics, 55:73–81.
  22. Dense Passage Retrieval for Open-Domain Question Answering. ArXiv:2004.04906 [cs].
  23. CTRL: A Conditional Transformer Language Model for Controllable Generation. ArXiv:1909.05858 [cs].
  24. Automatic classification of sentences to support Evidence Based Medicine. BMC Bioinformatics, 12(S2):S5.
  25. Automating Biomedical Evidence Synthesis: RobotReviewer. In Proceedings of ACL 2017, System Demonstrations, pages 7–12, Vancouver, Canada. Association for Computational Linguistics.
  26. Trialstreamer data.
  27. BERTweet: A pre-trained language model for English Tweets. Publisher: arXiv Version Number: 2.
  28. The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only. Publisher: arXiv Version Number: 1.
  29. Glove: Global Vectors for Word Representation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1532–1543, Doha, Qatar. Association for Computational Linguistics.
  30. AILAB-Udine@SMM4H 22: Limits of Transformers and BERT Ensembles. Publisher: arXiv Version Number: 1.
  31. Scientific Claim Verification with VerT5erini.
  32. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer. Publisher: arXiv Version Number: 3.
  33. MasonNLP+ at SemEval-2023 Task 8: Extracting Medical Questions, Experiences and Claims from Social Media using Knowledge-Augmented Pre-trained Language Models. ArXiv:2304.13875 [cs].
  34. Exploring the limits of transfer learning with a unified text-to-text transformer. 21:5485–5551.
  35. Stephen Robertson and Hugo Zaragoza. 2009. The Probabilistic Relevance Framework: BM25 and Beyond. Foundations and Trends® in Information Retrieval, 3(4):333–389.
  36. Muhammad Saaiq and Bushra Ashraf. 2017. Modifying “Pico” Question into “Picos” Model for More Robust and Reproducible Presentation of the.
  37. Autobots Ensemble: Identifying and Extracting Adverse Drug Reaction from Tweets Using Transformer Based Pipelines. In Proceedings of the Fifth Social Media Mining for Health Applications Workshop & Shared Task, pages 104–109, Barcelona, Spain. Association for Computational Linguistics.
  38. David Samuel and Milan Straka. 2021. ÚFAL at MultiLexNorm 2021: Improving Multilingual Lexical Normalization by Fine-tuning ByT5. In Proceedings of the Seventh Workshop on Noisy User-generated Text (W-NUT 2021), pages 483–492, Online. Association for Computational Linguistics.
  39. The PICO strategy for the research question construction and evidence search. Revista Latino-Americana de Enfermagem, 15(3):508–511.
  40. Data and systems for medication-related text classification and concept normalization from Twitter: insights from the Social Media Mining for Health (SMM4H)-2017 shared task. Journal of the American Medical Informatics Association, 25(10):1274–1283.
  41. Evidence-based Fact-Checking of Health-related Claims. In Findings of the Association for Computational Linguistics: EMNLP 2021, pages 3499–3512, Punta Cana, Dominican Republic. Association for Computational Linguistics.
  42. Correcting Diacritics and Typos with a ByT5 Transformer Model. Applied Sciences, 12(5):2636.
  43. EBM+: Advancing Evidence-Based Medicine via two level automatic identification of Populations, Interventions, Outcomes in medical literature. Artificial Intelligence in Medicine, 108:101949.
  44. A Contrastive Framework for Neural Text Generation. ArXiv:2202.06417 [cs].
  45. SciFact-Open: Towards open-domain scientific claim verification. ArXiv:2210.13777 [cs].
  46. RedHOT: A Corpus of Annotated Medical Questions, Experiences, and Claims on Social Media. ArXiv:2210.06331 [cs].
  47. Deep neural networks ensemble for detecting medication mentions in tweets. Journal of the American Medical Informatics Association, 26(12):1618–1626.
  48. ByT5: Towards a Token-Free Future with Pre-trained Byte-to-Byte Models. Transactions of the Association for Computational Linguistics, 10:291–306.
  49. Adverse drug reaction detection on social media with deep linguistic features. Journal of Biomedical Informatics, 106:103437.
  50. The PsyTAR dataset: From patients generated narratives to a corpus of adverse drug events and effectiveness of psychiatric medications. Data in Brief, 24:103838.

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.

Collections

Sign up for free to add this paper to one or more collections.

Tweets

Sign up for free to view the 1 tweet with 0 likes about this paper.