---
title: 'SeaReader: Multi-Doc Medical QA'
url: https://www.emergentmind.com/topics/seareader
type: topic
---

# SeaReader: Multi-Doc Medical QA

Searching arXiv for the SeaReader paper and closely related QA papers to ground citations.
arxiv_search(query="SeaReader MedQA 1802.10279", max_results=5, sort_by="relevance")
arxiv_search(query="MedQA R-NET neural reasoner iterative attention medical exam question answering", max_results=10, sort_by="relevance")
SeaReader is a large-scale, end-to-end reading comprehension system designed to answer multiple-choice medical exam questions by reading, integrating, and reasoning over many documents. It was introduced for MedQA, a question-answering task built from the National Medical Licensing Examination in China, and is formulated as a modular LSTM-based architecture with a dual-path attention mechanism that simultaneously performs focused reading within individual documents and evidence integration across documents. The system targets the central difficulties of clinical question answering: professional language, heterogeneous question styles, large-scale retrieval, non-verbatim evidence, and multi-fact reasoning over distributed textual support [1802.10279].

## 1. MedQA and the problem formulation

SeaReader was developed for MedQA, in which each problem consists of a question, possibly a brief case description, five candidate answers, and access to a large document collection. The system must select the best answer using the documents. The questions are drawn from the National Medical Licensing Examination in China and were filtered to remove duplicates and incompletes and to ensure that validation and test do not contain near-duplicates.

| Item | Value |
|---|---:|
| Training problems | 222,323 |
| Validation problems | 6,446 |
| Test problems | 6,405 |
| Average question length | 27.4 words |
| Average answer length | 4.2 words |
| Candidate answers per problem | 5 |

The document collection was extracted from 32 medical publications, including textbooks, reference books, guidebooks, and exam preparation books, and segmented into paragraph-level documents with metadata such as tags, book title, and chapter. The collection contains 243,712 documents, with average document length 43.2 words and an average of 3.8 metadata tags per document [1802.10279].

MedQA includes multiple question types with distinct reasoning profiles. A1 questions are single-statement best choice items and constitute 36.8% of the dataset. B1 questions, 11.2%, require selecting the best compatible choice from a shared candidate set. A2 questions, 35.3%, are case-summary best choice items with brief medical records. A3/A4 questions, 16.7%, are case-group best choice problems with shared information among multiple questions. This distribution makes the task materially different from cloze-style or single-passage span extraction: the system must often determine the *best* answer rather than recover a text span, and must do so under large retrieval and aggregation burdens.

Several characteristics make the benchmark a proxy for clinical reading comprehension rather than ordinary passage QA. The language is domain-specific; question styles span diagnosis, treatment, mechanism, and examination strategy; answer selection is comparative; and the relevant evidence is frequently distributed across multiple retrieved documents rather than stated verbatim in one passage. The paper’s Parkinson’s disease treatment example illustrates the intended setting: one document lists commonly used drugs, while another states that phenanthrene is used with caution in elderly patients, so the correct answer depends on integrating medication facts with the patient’s age condition [1802.10279].

## 2. End-to-end pipeline from retrieval to answer selection

SeaReader begins with document retrieval. For each problem, a Lucene-based BM25 retriever is run for the question and for each candidate answer, and the intersection of the two result sets is taken. Top-$N$ documents are then kept per candidate answer, with diversity enforced by balancing book and source types. In the typical configuration, top-10 documents are retained per candidate answer, yielding 50 documents per problem [1802.10279].

The question and a candidate answer are concatenated into a statement $S$, and each retrieved document $D_n$ is processed relative to that statement. Word embeddings are passed through shared BiLSTM context encoders to obtain contextual token representations for $S$ and for every $D_n$. This transforms the original multiple-choice problem into a statement-document matching problem conditioned on each answer option.

The central computation is then split into two complementary attention paths. The question-centric path extracts per-document information relevant to the question and aligns document content to question tokens. The document-centric path injects question information into each document and then performs cross-document attention, so that evidence can be compared and aggregated across retrieved documents. These paths are followed by a reasoning layer that applies word-level gating, reads over the gated sequences with a BiLSTM, and compresses each document into a per-document support vector via sequence-level max-pooling.

The final decision stage suppresses irrelevant documents with document-level gating and fuses the remaining support using both max-pooling and mean-pooling across documents. A one-layer feed-forward scorer then produces a scalar score for each candidate answer, and a softmax over the five candidate scores yields the final answer probabilities. This design makes SeaReader a candidate-scoring architecture rather than a span extractor, which aligns directly with the multiple-choice structure of medical examinations [1802.10279].

## 3. Dual-path attention and formal specification

Let $S \in \mathbb{R}^{L_Q \times d}$ denote the contextual embeddings of the concatenated statement, and let $D_n \in \mathbb{R}^{L_D \times d}$ denote the contextual embeddings of the $n$-th document. SeaReader constructs a matching matrix for each document:

$$
M_n(i,j) = S(i) \odot D_n(j).
$$

In the question-centric path, attention is applied column-wise over document tokens for each question token:

$$
\alpha_n(i,j) = \operatorname{softmax}(M_n(i,1), \ldots, M_n(i,L_D))(j),
$$

$$
R_n^Q(i) = \sum_{j=1}^{L_D} \alpha_n(i,j) D_n(j).
$$

For each token $S(i)$, $R_n^Q(i)$ summarizes the most relevant positions in document $D_n$ aligned to that token. This realizes focused per-document reading conditioned on the statement [1802.10279].

The document-centric path first applies row-wise attention to read related information in $S$ into each document token, summarized notationally as $R_n^D$. It then performs cross-document attention over concatenated document-question representations:

$$
M'_{mn}(i,j) = (D_m(i) \oplus R_m^D(i)) \odot (D_n(j) \oplus R_n^D(j)),
$$

$$
\beta_{mn}(i,j) =
\operatorname{softmax}(M'_{m1}(i,1), \ldots, M'_{m1}(i,L_D), \ldots, M'_{mN}(i,1), \ldots, M'_{mN}(i,L_D))_n(j),
$$

$$
R_m^{'D}(i) = \sum_{n=1}^{N} \sum_{j=1}^{L_D} \beta_{mn}(i,j) (D_n(j) \oplus R_n^D(j)).
$$

This cross-document attention extracts evidence from other documents conditional on token $i$ in document $m$, enabling multi-document support aggregation rather than isolated passage matching.

SeaReader supplements attention with matching features $F$ computed from the question-document matching matrix. Two alternatives are described. The CNN extractor uses two convolutional layers with max-pooling between and after layers, with the second layer employing dilated convolution to preserve word-level resolution. The lighter alternative uses row-wise and column-wise max-pooling and mean-pooling over the matching matrix. These features are combined with the attention reads in the reasoning module.

At the output layer, if $z_k$ is the score assigned to candidate answer $k$, then the final probability is

$$
P(y = k \mid Q, \{D\}) = \operatorname{softmax}(z_1,\ldots,z_5)(k),
$$

with base objective

$$
L_{\text{task}} = - \log P(y = y^* \mid Q, \{D\}).
$$

The architecture also admits interpretability-oriented regularized variants, including a gating importance penalty, noisy gating, and $L_{21}$-regularized delta embeddings. These additions are not separate modules layered on top of the model’s predictions; they are part of the training formulation used to expose token and document salience [1802.10279].

## 4. Training regime and interpretability mechanisms

SeaReader uses 200-dimensional word embeddings trained with skip-gram on all training questions and the entire document collection. Unseen test words are mapped to zero vectors. To balance fit and generalization, the system refines embeddings through an $L_{21}$-regularized delta term:

$$
w = w_{\text{skip-gram}} + w_{\Delta},
$$

with loss augmented by

$$
\text{loss} = \text{loss}_{\text{task}} + c \sum_{i=1}^{n}\left(\sum_{j=1}^{d} w_{\Delta ij}^2\right)^{1/2}.
$$

The contextual and reasoning BiLSTMs use hidden size 128 per direction; dropout is 0.2; the decision layer is a single-layer feed-forward scorer; and optimization uses Adam with $\epsilon = 10^{-6}$ and an exponentially decayed learning rate. Implementation is in TensorFlow. Batch size is 15, corresponding to approximately 750 documents per batch, described as the practical maximum on a single GPU [1802.10279].

Because computation scales rapidly with both document count and sequence length, documents are truncated to 100 words before entering SeaReader. The paper states that this slightly reduces accuracy but substantially speeds up training. Retrieval typically keeps top-10 documents per candidate answer, although performance continues to improve up to top-20, at increasing computational cost.

Interpretability is treated as an architectural objective rather than only a post hoc visualization step. First, SeaReader employs gating with an importance penalty, in which average gate values above threshold $t = 0.7$ are penalized, making gate magnitude a proxy for token or document importance:

$$
\text{loss} = \text{loss}_{\text{task}} + c \cdot \max\left(\frac{1}{L}\sum_{i=1}^L g(i) - t, 0\right).
$$

Second, noisy gating adds Gaussian noise after gating,

$$
X_{\text{out}} = \text{gate}(X_{\text{in}})X_{\text{in}} + \sigma(0,s),
$$

so that high-gate features remain visible over noise and low-gate features are deemphasized. Third, attention visualization is provided for the column-wise question-centric attention, the row-wise attention, and the cross-document attention, allowing inspection of matching patterns and evidence aggregation. The paper reports that the $L_{21}$-regularized delta embeddings surface logical connectives and common diagnostic terms as the most modified words, offering another lens on task salience [1802.10279].

## 5. Empirical performance on MedQA

SeaReader substantially improves over the compared baselines on MedQA. On the validation and test sets, the single model achieves 73.9% and 73.6% accuracy, respectively, while the ensemble reaches 75.8% and 75.3%. The paper reports test accuracies of 59.3% for Iterative Attention, 53.0% for Neural Reasoner, and 64.5% for R-NET. A human passing score is given as 60.0% (360/600), as a reference point rather than a like-for-like model comparison [1802.10279].

| System | Validation | Test |
|---|---:|---:|
| SeaReader | 73.9% | 73.6% |
| SeaReader (ensemble) | 75.8% | 75.3% |
| Iterative Attention | — | 59.3% |
| Neural Reasoner | — | 53.0% |
| R-NET | — | 64.5% |

Performance remains strong across MedQA categories. On the test set, SeaReader obtains 0.754 on A1, 0.738 on B1, 0.737 on A2, and 0.707 on A3/A4. The corresponding R-NET numbers are 0.671, 0.617, 0.663, and 0.618. The largest relative gains appear in B1 and A3/A4, where information is mixed or shared across questions and documents. This suggests that the dual-path attention is especially useful when evidence must be integrated across multiple textual contexts rather than extracted from a single, self-contained passage.

The effect of retrieval depth is particularly informative. With top-1 document per candidate, test accuracy is 57.8%, and the relevant document ratio is 0.90. At top-5, accuracy rises to 71.7% despite the relevant ratio dropping to 0.54. At top-10, SeaReader reaches 73.6% with relevant ratio 0.46, and at top-20 it reaches 74.4% with relevant ratio 0.29. Performance therefore improves as more documents are supplied even though the fraction of relevant documents decreases. A plausible implication is that SeaReader’s cross-document filtering and fusion are effective enough to convert retrieval breadth into additional evidence rather than only additional noise [1802.10279].

## 6. Limitations, comparative position, and example reasoning behavior

SeaReader’s main limitations are tied to retrieval, truncation, and difficult inference chains. Accuracy depends on the retrieval stage’s ability to surface relevant documents. As the number of documents increases, the fraction of relevant ones decreases, and gains eventually show diminishing returns because of noise. Document truncation to 100 tokens improves training efficiency but can remove late-occurring evidence or long-range dependencies. Some clinical questions require nuanced reasoning over many facts, and although cross-document attention helps, extremely long chains and subtle causal mechanisms remain challenging. Rare specialist terminology and atypical cases may also be poorly represented in the embeddings, even with $L_{21}$ delta adjustments [1802.10279].

Within the reading-comprehension literature, SeaReader differs from span-extraction systems such as BiDAF, AoA Reader, and R-NET in three explicit ways. It operates over many retrieved documents rather than a pre-selected paragraph, uses dual-path attention to couple intra-document reading with inter-document integration, and produces candidate-level scores instead of answer spans. Relative to Iterative Attention and Neural Reasoner, its explicit cross-document attention and modular gating provide stronger multi-document evidence integration on real-world, non-cloze medical questions.

The Parkinson’s disease example in the paper illustrates the architecture’s intended behavior. One retrieved document states that commonly used drugs include phenanthrene, amantadine, levodopa, and compound levodopa. Another states that phenanthrene is mainly suitable for those with obvious tremble but is more used in young patients and should be used with caution in elderly patients. SeaReader’s question-centric attention highlights the treatment options relative to the candidate answers, while the document-centric path integrates the age condition from the question into each document and aggregates the cautionary evidence across documents. Document-level gating then emphasizes the caution against phenanthrene for elderly patients together with the broader effectiveness of levodopa-based treatments, and the pooled support favors the levodopa option [1802.10279].

SeaReader therefore occupies a specific position in large-scale machine reading: it is a multi-document, multiple-choice clinical QA model whose principal innovation is not merely higher-capacity encoding, but an explicit decomposition of reading into per-document alignment and cross-document evidence synthesis. Its interpretability mechanisms, including gating penalties, noisy gating, delta embeddings, and attention visualization, further distinguish it as an attempt to make multi-document decision formation inspectable rather than only accurate.

Source: https://www.emergentmind.com/topics/seareader