Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
80 tokens/sec
GPT-4o
59 tokens/sec
Gemini 2.5 Pro Pro
43 tokens/sec
o3 Pro
7 tokens/sec
GPT-4.1 Pro
50 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

MeDiaQA: A Question Answering Dataset on Medical Dialogues (2108.08074v1)

Published 18 Aug 2021 in cs.CL and cs.AI

Abstract: In this paper, we introduce MeDiaQA, a novel question answering(QA) dataset, which constructed on real online Medical Dialogues. It contains 22k multiple-choice questions annotated by human for over 11k dialogues with 120k utterances between patients and doctors, covering 150 specialties of diseases, which are collected from haodf.com and dxy.com. MeDiaQA is the first QA dataset where reasoning over medical dialogues, especially their quantitative contents. The dataset has the potential to test the computing, reasoning and understanding ability of models across multi-turn dialogues, which is challenging compared with the existing datasets. To address the challenges, we design MeDia-BERT, and it achieves 64.3% accuracy, while human performance of 93% accuracy, which indicates that there still remains a large room for improvement.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (5)
  1. Huqun Suri (1 paper)
  2. Qi Zhang (785 papers)
  3. Wenhua Huo (1 paper)
  4. Yan Liu (420 papers)
  5. Chunsheng Guan (1 paper)
Citations (4)