SeaReader: Multi-Doc Medical QA
- SeaReader is a large-scale, end-to-end reading comprehension system designed for multiple-choice medical exams by integrating text from numerous documents.
- It employs a modular LSTM-based architecture with dual-path attention that performs both focused question-centric reading and cross-document evidence synthesis.
- Empirical results on MedQA show significant performance gains over baseline models, with interpretability features like gating and delta embeddings enhancing insight into its reasoning.
Searching arXiv for the SeaReader paper and closely related QA papers to ground citations. arxiv_search(query="SeaReader MedQA (Zhang et al., 2018)", max_results=5, sort_by="relevance") arxiv_search(query="MedQA R-NET neural reasoner iterative attention medical exam question answering", max_results=10, sort_by="relevance") SeaReader is a large-scale, end-to-end reading comprehension system designed to answer multiple-choice medical exam questions by reading, integrating, and reasoning over many documents. It was introduced for MedQA, a question-answering task built from the National Medical Licensing Examination in China, and is formulated as a modular LSTM-based architecture with a dual-path attention mechanism that simultaneously performs focused reading within individual documents and evidence integration across documents. The system targets the central difficulties of clinical question answering: professional language, heterogeneous question styles, large-scale retrieval, non-verbatim evidence, and multi-fact reasoning over distributed textual support (Zhang et al., 2018).
1. MedQA and the problem formulation
SeaReader was developed for MedQA, in which each problem consists of a question, possibly a brief case description, five candidate answers, and access to a large document collection. The system must select the best answer using the documents. The questions are drawn from the National Medical Licensing Examination in China and were filtered to remove duplicates and incompletes and to ensure that validation and test do not contain near-duplicates.
| Item | Value |
|---|---|
| Training problems | 222,323 |
| Validation problems | 6,446 |
| Test problems | 6,405 |
| Average question length | 27.4 words |
| Average answer length | 4.2 words |
| Candidate answers per problem | 5 |
The document collection was extracted from 32 medical publications, including textbooks, reference books, guidebooks, and exam preparation books, and segmented into paragraph-level documents with metadata such as tags, book title, and chapter. The collection contains 243,712 documents, with average document length 43.2 words and an average of 3.8 metadata tags per document (Zhang et al., 2018).
MedQA includes multiple question types with distinct reasoning profiles. A1 questions are single-statement best choice items and constitute 36.8% of the dataset. B1 questions, 11.2%, require selecting the best compatible choice from a shared candidate set. A2 questions, 35.3%, are case-summary best choice items with brief medical records. A3/A4 questions, 16.7%, are case-group best choice problems with shared information among multiple questions. This distribution makes the task materially different from cloze-style or single-passage span extraction: the system must often determine the best answer rather than recover a text span, and must do so under large retrieval and aggregation burdens.
Several characteristics make the benchmark a proxy for clinical reading comprehension rather than ordinary passage QA. The language is domain-specific; question styles span diagnosis, treatment, mechanism, and examination strategy; answer selection is comparative; and the relevant evidence is frequently distributed across multiple retrieved documents rather than stated verbatim in one passage. The paper’s Parkinson’s disease treatment example illustrates the intended setting: one document lists commonly used drugs, while another states that phenanthrene is used with caution in elderly patients, so the correct answer depends on integrating medication facts with the patient’s age condition (Zhang et al., 2018).
2. End-to-end pipeline from retrieval to answer selection
SeaReader begins with document retrieval. For each problem, a Lucene-based BM25 retriever is run for the question and for each candidate answer, and the intersection of the two result sets is taken. Top- documents are then kept per candidate answer, with diversity enforced by balancing book and source types. In the typical configuration, top-10 documents are retained per candidate answer, yielding 50 documents per problem (Zhang et al., 2018).
The question and a candidate answer are concatenated into a statement , and each retrieved document is processed relative to that statement. Word embeddings are passed through shared BiLSTM context encoders to obtain contextual token representations for and for every . This transforms the original multiple-choice problem into a statement-document matching problem conditioned on each answer option.
The central computation is then split into two complementary attention paths. The question-centric path extracts per-document information relevant to the question and aligns document content to question tokens. The document-centric path injects question information into each document and then performs cross-document attention, so that evidence can be compared and aggregated across retrieved documents. These paths are followed by a reasoning layer that applies word-level gating, reads over the gated sequences with a BiLSTM, and compresses each document into a per-document support vector via sequence-level max-pooling.
The final decision stage suppresses irrelevant documents with document-level gating and fuses the remaining support using both max-pooling and mean-pooling across documents. A one-layer feed-forward scorer then produces a scalar score for each candidate answer, and a softmax over the five candidate scores yields the final answer probabilities. This design makes SeaReader a candidate-scoring architecture rather than a span extractor, which aligns directly with the multiple-choice structure of medical examinations (Zhang et al., 2018).
3. Dual-path attention and formal specification
Let denote the contextual embeddings of the concatenated statement, and let denote the contextual embeddings of the -th document. SeaReader constructs a matching matrix for each document:
In the question-centric path, attention is applied column-wise over document tokens for each question token:
0
For each token 1, 2 summarizes the most relevant positions in document 3 aligned to that token. This realizes focused per-document reading conditioned on the statement (Zhang et al., 2018).
The document-centric path first applies row-wise attention to read related information in 4 into each document token, summarized notationally as 5. It then performs cross-document attention over concatenated document-question representations:
6
7
8
This cross-document attention extracts evidence from other documents conditional on token 9 in document 0, enabling multi-document support aggregation rather than isolated passage matching.
SeaReader supplements attention with matching features 1 computed from the question-document matching matrix. Two alternatives are described. The CNN extractor uses two convolutional layers with max-pooling between and after layers, with the second layer employing dilated convolution to preserve word-level resolution. The lighter alternative uses row-wise and column-wise max-pooling and mean-pooling over the matching matrix. These features are combined with the attention reads in the reasoning module.
At the output layer, if 2 is the score assigned to candidate answer 3, then the final probability is
4
with base objective
5
The architecture also admits interpretability-oriented regularized variants, including a gating importance penalty, noisy gating, and 6-regularized delta embeddings. These additions are not separate modules layered on top of the model’s predictions; they are part of the training formulation used to expose token and document salience (Zhang et al., 2018).
4. Training regime and interpretability mechanisms
SeaReader uses 200-dimensional word embeddings trained with skip-gram on all training questions and the entire document collection. Unseen test words are mapped to zero vectors. To balance fit and generalization, the system refines embeddings through an 7-regularized delta term:
8
with loss augmented by
9
The contextual and reasoning BiLSTMs use hidden size 128 per direction; dropout is 0.2; the decision layer is a single-layer feed-forward scorer; and optimization uses Adam with 0 and an exponentially decayed learning rate. Implementation is in TensorFlow. Batch size is 15, corresponding to approximately 750 documents per batch, described as the practical maximum on a single GPU (Zhang et al., 2018).
Because computation scales rapidly with both document count and sequence length, documents are truncated to 100 words before entering SeaReader. The paper states that this slightly reduces accuracy but substantially speeds up training. Retrieval typically keeps top-10 documents per candidate answer, although performance continues to improve up to top-20, at increasing computational cost.
Interpretability is treated as an architectural objective rather than only a post hoc visualization step. First, SeaReader employs gating with an importance penalty, in which average gate values above threshold 1 are penalized, making gate magnitude a proxy for token or document importance:
2
Second, noisy gating adds Gaussian noise after gating,
3
so that high-gate features remain visible over noise and low-gate features are deemphasized. Third, attention visualization is provided for the column-wise question-centric attention, the row-wise attention, and the cross-document attention, allowing inspection of matching patterns and evidence aggregation. The paper reports that the 4-regularized delta embeddings surface logical connectives and common diagnostic terms as the most modified words, offering another lens on task salience (Zhang et al., 2018).
5. Empirical performance on MedQA
SeaReader substantially improves over the compared baselines on MedQA. On the validation and test sets, the single model achieves 73.9% and 73.6% accuracy, respectively, while the ensemble reaches 75.8% and 75.3%. The paper reports test accuracies of 59.3% for Iterative Attention, 53.0% for Neural Reasoner, and 64.5% for R-NET. A human passing score is given as 60.0% (360/600), as a reference point rather than a like-for-like model comparison (Zhang et al., 2018).
| System | Validation | Test |
|---|---|---|
| SeaReader | 73.9% | 73.6% |
| SeaReader (ensemble) | 75.8% | 75.3% |
| Iterative Attention | — | 59.3% |
| Neural Reasoner | — | 53.0% |
| R-NET | — | 64.5% |
Performance remains strong across MedQA categories. On the test set, SeaReader obtains 0.754 on A1, 0.738 on B1, 0.737 on A2, and 0.707 on A3/A4. The corresponding R-NET numbers are 0.671, 0.617, 0.663, and 0.618. The largest relative gains appear in B1 and A3/A4, where information is mixed or shared across questions and documents. This suggests that the dual-path attention is especially useful when evidence must be integrated across multiple textual contexts rather than extracted from a single, self-contained passage.
The effect of retrieval depth is particularly informative. With top-1 document per candidate, test accuracy is 57.8%, and the relevant document ratio is 0.90. At top-5, accuracy rises to 71.7% despite the relevant ratio dropping to 0.54. At top-10, SeaReader reaches 73.6% with relevant ratio 0.46, and at top-20 it reaches 74.4% with relevant ratio 0.29. Performance therefore improves as more documents are supplied even though the fraction of relevant documents decreases. A plausible implication is that SeaReader’s cross-document filtering and fusion are effective enough to convert retrieval breadth into additional evidence rather than only additional noise (Zhang et al., 2018).
6. Limitations, comparative position, and example reasoning behavior
SeaReader’s main limitations are tied to retrieval, truncation, and difficult inference chains. Accuracy depends on the retrieval stage’s ability to surface relevant documents. As the number of documents increases, the fraction of relevant ones decreases, and gains eventually show diminishing returns because of noise. Document truncation to 100 tokens improves training efficiency but can remove late-occurring evidence or long-range dependencies. Some clinical questions require nuanced reasoning over many facts, and although cross-document attention helps, extremely long chains and subtle causal mechanisms remain challenging. Rare specialist terminology and atypical cases may also be poorly represented in the embeddings, even with 5 delta adjustments (Zhang et al., 2018).
Within the reading-comprehension literature, SeaReader differs from span-extraction systems such as BiDAF, AoA Reader, and R-NET in three explicit ways. It operates over many retrieved documents rather than a pre-selected paragraph, uses dual-path attention to couple intra-document reading with inter-document integration, and produces candidate-level scores instead of answer spans. Relative to Iterative Attention and Neural Reasoner, its explicit cross-document attention and modular gating provide stronger multi-document evidence integration on real-world, non-cloze medical questions.
The Parkinson’s disease example in the paper illustrates the architecture’s intended behavior. One retrieved document states that commonly used drugs include phenanthrene, amantadine, levodopa, and compound levodopa. Another states that phenanthrene is mainly suitable for those with obvious tremble but is more used in young patients and should be used with caution in elderly patients. SeaReader’s question-centric attention highlights the treatment options relative to the candidate answers, while the document-centric path integrates the age condition from the question into each document and aggregates the cautionary evidence across documents. Document-level gating then emphasizes the caution against phenanthrene for elderly patients together with the broader effectiveness of levodopa-based treatments, and the pooled support favors the levodopa option (Zhang et al., 2018).
SeaReader therefore occupies a specific position in large-scale machine reading: it is a multi-document, multiple-choice clinical QA model whose principal innovation is not merely higher-capacity encoding, but an explicit decomposition of reading into per-document alignment and cross-document evidence synthesis. Its interpretability mechanisms, including gating penalties, noisy gating, delta embeddings, and attention visualization, further distinguish it as an attempt to make multi-document decision formation inspectable rather than only accurate.