---
title: 'UIT-ViCoQA: Vietnamese Conversational QA in Healthcare'
url: https://www.emergentmind.com/topics/uit-vicoqa
type: topic
---

# UIT-ViCoQA: Vietnamese Conversational QA in Healthcare

Searching arXiv for UIT-ViCoQA and closely related Vietnamese QA resources.
UIT-ViCoQA is a Vietnamese corpus for conversational machine comprehension, introduced for research on conversational question answering over healthcare texts. It is built from health news articles and is designed to evaluate systems that answer a current question by using both the source passage and the prior question–answer history in a multi-turn dialogue. The corpus contains 10,000 questions with answers over 2,000 conversations about health news articles, and the benchmark paper evaluates several neural baselines while emphasizing that performance remains well below human levels [2105.01542].

## 1. Definition and task formulation

UIT-ViCoQA is framed as a corpus for **conversational machine comprehension / conversational machine reading comprehension** in Vietnamese. The underlying task is not independent single-turn QA. Instead, a model must answer the current question using the source passage together with the preceding dialogue turns. The paper illustrates this dependency explicitly: to answer \(Q2\), the answerer needs the passage plus \(Q1\) and \(A1\); to answer \(Q3\), the answerer needs the passage plus \((Q1, A1)\) and \((Q2, A2)\). In the paper’s own wording, “The chain of question-answer pairs Q1-A1, Q2-A2 is the history of the conversation” [2105.01542].

The dataset is in Vietnamese and is focused on the healthcare/news domain. Its source material consists of health news articles collected from VnExpress, specifically from `https://vnexpress.net/suc-khoe`, which the paper describes as one of the most-read online newspapers in Vietnam. Each conversation is grounded in one reading article and contains five question-answer pairs, so the corpus structure is fixed at five turns per conversation. This makes the benchmark explicitly history-aware and context-dependent rather than a collection of isolated article-question pairs.

The paper is high-level rather than formally axiomatized. It does not define the task through a compact probabilistic equation, but it does describe the ingredients in notation-like prose: a conversation reading comprehension task consists of a reading passage as context \(C\), the conversation history \(H\) including multiple question-answer pairs, and the generated answers \(A\). In the data-collection description, each turn also includes a supporting span \(S\), because the answerer first selects evidence from the article and then submits a natural answer.

## 2. Source documents and corpus construction

The corpus was created in three phases. First, the authors collected health news articles from VnExpress using **Scrapy**. Second, they built an annotation tool that supports two-person conversation creation over a given article. Third, they hired a team of annotators to create the data through that tool [2105.01542].

The corpus-level statistics reported in the paper are concise and important.

| Aspect | Value |
|---|---|
| Number of passages | 2,000 |
| Number of questions | 10,000 |
| Passage length | 404.1 words on average |
| Question length | 9.4 words on average |
| Answer length | 9.7 words on average |

The dataset is split randomly into **70% train, 15% development, and 15% test**. Because the corpus contains 2,000 conversations and each conversation has five questions, this suggests approximately 1,400 training conversations, 300 development conversations, and 300 test conversations, corresponding to approximately 7,000, 1,500, and 1,500 questions. This implication follows directly from the stated split and the fixed five-turn structure.

In addition to the main split, the authors sampled **100 articles from the development set** to form an **analysis set**, following CoQA. This reflects the paper’s explicit methodological alignment with CoQA-style conversational QA, although UIT-ViCoQA is substantially smaller and domain-specific.

## 3. Annotation workflow and answer representation

For each conversation, there were **two different annotators** with distinct roles: a **questioner** and an **answerer**. The questioner asks a question. The answerer reads the article, selects a supporting text span from the article, and then submits a natural answer. Each turn therefore contains three aligned objects: the question \(Q\), the supporting span \(S\), and the answer \(A\) [2105.01542].

This design is central to how UIT-ViCoQA should be understood. The answers are evidence-grounded, because the answerer first selects a span from the source article, but the final output is a natural-language response rather than necessarily a verbatim extraction. Accordingly, the benchmark is closer in spirit to CoQA-style conversational QA than to strictly extractive SQuAD-style MRC.

The annotation guidelines impose several constraints. Answers must be extracted from the article, so unanswerable questions are not allowed. Questioners are encouraged to ask questions using synonyms, opposite words, and coreference. Answerers are instructed to provide short answers and to limit the introduction of new words beyond the article content; at the same time, the paper states that answers should be full answers with complete texts, correct syntax, and punctuation. The resulting answer format is therefore best described as extractively grounded but natural/free-form short answers.

The paper also describes an automatic validation heuristic. After the answerer submits an answer, the annotation system compares the answer with the asked question at character level. If the given answer matches about **70%** with the asked question, it is considered a valid answer; otherwise, the answerer must submit another answer. The paper reports this mechanism as part of the workflow, but it does not formalize it further. Formal inter-annotator agreement statistics, such as Cohen’s kappa, are not reported.

## 4. Conversation structure, linguistic phenomena, and corpus profile

Conversation structure is not incidental in UIT-ViCoQA; it is the defining property of the benchmark. Later questions may depend on earlier turns through ellipsis, paraphrase, and coreference. The linguistic analysis in the paper quantifies this directly: **20.6%** of questions involve explicit coreference, **5.8%** implicit coreference, and **73.6%** no coreference. In the relation between question and passage, **47.6%** are lexical match, **48.0%** paraphrasing, and **4.4%** pragmatic [2105.01542].

The question-type distribution further characterizes the corpus. The largest category is **What** at **32.6%**, followed by **How many** at **17.2%**, **Who** at **9.0%**, **Why** at **7.8%**, **How** at **7.6%**, **Which** at **7.0%**, **Yes/No** at **6.6%**, **Others** at **5.6%**, **Where** at **4.0%**, and **When** at **2.6%**. This distribution matters because it shapes both modeling difficulty and the interpretation of aggregate benchmark scores.

Compared with CoQA, the paper states that UIT-ViCoQA is much smaller and less diverse in domain, but its passages, questions, and answers are longer on average. The comparison table gives **UIT-ViCoQA** as health domain, 2,000 passages, 10,000 questions, passage length 404.1, question length 9.4, and answer length 9.7, whereas **CoQA** is reported as diverse domains, 8,399 passages, 127,000 questions, passage length 271.0, question length 5.5, and answer length 2.7. A plausible implication is that Vietnamese conversational QA in this corpus combines longer textual context with more expansive answer strings, which partly explains the large gap between token-overlap and exact-match evaluation.

## 5. Baselines, preprocessing, and evaluation protocol

The benchmark evaluates four systems: **DrQA**, **SDNet**, **FlowQA**, and **GraphFlow**. DrQA is used via its Document Reader to extract answer spans for questions. SDNet is described as a contextual attention-based model built on the idea of DrQA but with a mechanism to extract the context of the conversation. FlowQA and GraphFlow are the two FLOW models; the paper states that the FLOW mechanism allows models to encode the history of the conversation comprehensively and integrates well the latent semantic of the conversation history [2105.01542].

Before model fitting, the preprocessing and representation steps are explicit. The authors remove special characters and stop words, segment sentences into words using **Underthesea**, and transform texts into vectors using Vietnamese **fastText** embeddings from Grave et al. The embedding dimension is **300**. The paper does not provide architecture-specific hyperparameters, optimizer schedules, or training epochs for these baselines.

Evaluation uses **Exact Match (EM)** and **F1-score**. The paper defines them informally rather than with explicit equations. EM measures exact matching of prediction answers with original answers, while F1 measures overlap between predicted and correct answers. The benchmark results are reported on both development and test sets, with human performance included for comparison.

| Model | Test EM | Test F1 |
|---|---:|---:|
| DrQA | 13.50 | 37.71 |
| SDNet | 15.60 | 40.50 |
| FlowQA | 12.53 | 45.27 |
| GraphFlow | 14.73 | 45.16 |
| Human performance | 38.66 | 76.18 |

The best model by **test F1** is **FlowQA with 45.27%**, slightly ahead of **GraphFlow with 45.16%**. The best **test EM** is **SDNet with 15.60%**. The abstract highlights that the best model is **30.91 points behind human performance in F1**, indicating substantial room for improvement.

## 6. Empirical findings, error analysis, and limitations

The headline empirical result is that conversation-aware architectures outperform a strong non-conversational reader on F1, but exact match remains difficult. The paper explicitly attributes the large F1–EM gap to the prevalence of free-form answers and to linguistic ambiguity in Vietnamese interrogatives [2105.01542].

The answer-type analysis on the development set divides predicted answers into **matching answers** at **16.73%**, **free-form answers** at **59.93%**, and **wrong answers** at **23.27%**. The concentration in the free-form category is used by the authors to explain why F1 and EM differ so substantially: a prediction may overlap strongly with the gold answer without reproducing the exact canonical string.

The worked error analysis shows that **FlowQA** and **GraphFlow** generally provide the most relevant answers when previous turns matter, but all models fail on some ambiguous questions. The paper identifies Vietnamese interrogative ambiguity as a language-specific challenge, citing short elliptical questions such as “Cụ thể?” and “Nguy cơ là gì?” as examples that can be interpreted in more than one way.

The paper also analyzes performance by question type. A question is counted as correctly answered if **F1-score > 70%**. On the development set, the ratios of correct answers are **35.12** for What, **14.55** for How many, **11.70** for Who, **8.19** for How, **7.02** for Which, **6.86** for Why, **5.02** for Yes/No, **4.35** for When, **3.18** for Where, and **4.01** for Others. Since What questions are also the most common, the benchmark exhibits a familiar interaction between class frequency and apparent competence.

The paper makes its limitations explicit. UIT-ViCoQA is relatively small compared with English conversational QA datasets; it is limited to the health-news domain; models still struggle to understand conversational history and contextual meaning; and the corpus exhibits a large F1–EM gap due to answer variability and Vietnamese interrogative phenomena. Future work proposed by the authors includes increasing the quantity and quality of the dataset and experimenting with deep learning and transfer learning using pre-trained language models such as **BERT**, **XLM-style cross-lingual models**, and **PhoBERT**.

## 7. Position in Vietnamese QA research and naming distinctions

UIT-ViCoQA occupies a specific position within the Vietnamese QA landscape. Earlier Vietnamese resources cited alongside it include **UIT-ViQuAD** for Vietnamese extractive MRC on Wikipedia, **UIT-ViNewsQA** for extractive MRC on health news, and **ViMMRC** for Vietnamese multiple-choice MRC on primary-school textbooks. Against these, UIT-ViCoQA is distinguished by its move from single-turn reading comprehension to **multi-turn conversational QA** [2105.01542].

This role is corroborated by later survey-style discussion in the VLSP 2021 ViMRC paper, which refers to **ViCoQA** as a Vietnamese corpus for **conversational reading comprehension** and contrasts it with **UIT-ViQuAD 2.0**, a single-turn extractive benchmark with answerable and unanswerable questions [2203.11400]. In that ecosystem, UIT-ViCoQA serves as the conversation-history counterpart to span-extraction benchmarks.

A recurrent source of confusion is the similarity between **UIT-ViCoQA** and **UIT-ViCoV19QA**. These are separate datasets. **UIT-ViCoV19QA** is a Vietnamese community-based generative QA benchmark in the COVID-19 public-health domain, built from trusted FAQ sources and formulated as single-turn answer generation with multiple paraphrased references, not as conversational machine reading comprehension [2209.06668]. The lexical resemblance between the names can obscure a substantive difference in task design: UIT-ViCoQA centers on article-grounded, five-turn, history-aware dialogue, whereas UIT-ViCoV19QA centers on single-turn generative answering over FAQ-style COVID-19 information.

The benchmark’s broader significance lies in establishing Vietnamese conversational QA as a concrete evaluation problem in a high-interest domain. The paper presents the corpus, the baseline results, and the linguistic analysis as evidence that healthcare dialogue in Vietnamese requires not only passage understanding but also robust modeling of conversational flow, paraphrase, and coreference. The dataset is available for research purposes at `http://nlp.uit.edu.vn/datasets/` [2105.01542].

Source: https://www.emergentmind.com/topics/uit-vicoqa