Papers
Topics
Authors
Recent
Search
2000 character limit reached

SeaReader: Multi-Doc Medical QA

Updated 7 July 2026
  • SeaReader is a large-scale, end-to-end reading comprehension system designed for multiple-choice medical exams by integrating text from numerous documents.
  • It employs a modular LSTM-based architecture with dual-path attention that performs both focused question-centric reading and cross-document evidence synthesis.
  • Empirical results on MedQA show significant performance gains over baseline models, with interpretability features like gating and delta embeddings enhancing insight into its reasoning.

Searching arXiv for the SeaReader paper and closely related QA papers to ground citations. arxiv_search(query="SeaReader MedQA (Zhang et al., 2018)", max_results=5, sort_by="relevance") arxiv_search(query="MedQA R-NET neural reasoner iterative attention medical exam question answering", max_results=10, sort_by="relevance") SeaReader is a large-scale, end-to-end reading comprehension system designed to answer multiple-choice medical exam questions by reading, integrating, and reasoning over many documents. It was introduced for MedQA, a question-answering task built from the National Medical Licensing Examination in China, and is formulated as a modular LSTM-based architecture with a dual-path attention mechanism that simultaneously performs focused reading within individual documents and evidence integration across documents. The system targets the central difficulties of clinical question answering: professional language, heterogeneous question styles, large-scale retrieval, non-verbatim evidence, and multi-fact reasoning over distributed textual support (Zhang et al., 2018).

1. MedQA and the problem formulation

SeaReader was developed for MedQA, in which each problem consists of a question, possibly a brief case description, five candidate answers, and access to a large document collection. The system must select the best answer using the documents. The questions are drawn from the National Medical Licensing Examination in China and were filtered to remove duplicates and incompletes and to ensure that validation and test do not contain near-duplicates.

Item Value
Training problems 222,323
Validation problems 6,446
Test problems 6,405
Average question length 27.4 words
Average answer length 4.2 words
Candidate answers per problem 5

The document collection was extracted from 32 medical publications, including textbooks, reference books, guidebooks, and exam preparation books, and segmented into paragraph-level documents with metadata such as tags, book title, and chapter. The collection contains 243,712 documents, with average document length 43.2 words and an average of 3.8 metadata tags per document (Zhang et al., 2018).

MedQA includes multiple question types with distinct reasoning profiles. A1 questions are single-statement best choice items and constitute 36.8% of the dataset. B1 questions, 11.2%, require selecting the best compatible choice from a shared candidate set. A2 questions, 35.3%, are case-summary best choice items with brief medical records. A3/A4 questions, 16.7%, are case-group best choice problems with shared information among multiple questions. This distribution makes the task materially different from cloze-style or single-passage span extraction: the system must often determine the best answer rather than recover a text span, and must do so under large retrieval and aggregation burdens.

Several characteristics make the benchmark a proxy for clinical reading comprehension rather than ordinary passage QA. The language is domain-specific; question styles span diagnosis, treatment, mechanism, and examination strategy; answer selection is comparative; and the relevant evidence is frequently distributed across multiple retrieved documents rather than stated verbatim in one passage. The paper’s Parkinson’s disease treatment example illustrates the intended setting: one document lists commonly used drugs, while another states that phenanthrene is used with caution in elderly patients, so the correct answer depends on integrating medication facts with the patient’s age condition (Zhang et al., 2018).

2. End-to-end pipeline from retrieval to answer selection

SeaReader begins with document retrieval. For each problem, a Lucene-based BM25 retriever is run for the question and for each candidate answer, and the intersection of the two result sets is taken. Top-NN documents are then kept per candidate answer, with diversity enforced by balancing book and source types. In the typical configuration, top-10 documents are retained per candidate answer, yielding 50 documents per problem (Zhang et al., 2018).

The question and a candidate answer are concatenated into a statement SS, and each retrieved document DnD_n is processed relative to that statement. Word embeddings are passed through shared BiLSTM context encoders to obtain contextual token representations for SS and for every DnD_n. This transforms the original multiple-choice problem into a statement-document matching problem conditioned on each answer option.

The central computation is then split into two complementary attention paths. The question-centric path extracts per-document information relevant to the question and aligns document content to question tokens. The document-centric path injects question information into each document and then performs cross-document attention, so that evidence can be compared and aggregated across retrieved documents. These paths are followed by a reasoning layer that applies word-level gating, reads over the gated sequences with a BiLSTM, and compresses each document into a per-document support vector via sequence-level max-pooling.

The final decision stage suppresses irrelevant documents with document-level gating and fuses the remaining support using both max-pooling and mean-pooling across documents. A one-layer feed-forward scorer then produces a scalar score for each candidate answer, and a softmax over the five candidate scores yields the final answer probabilities. This design makes SeaReader a candidate-scoring architecture rather than a span extractor, which aligns directly with the multiple-choice structure of medical examinations (Zhang et al., 2018).

3. Dual-path attention and formal specification

Let S∈RLQ×dS \in \mathbb{R}^{L_Q \times d} denote the contextual embeddings of the concatenated statement, and let Dn∈RLD×dD_n \in \mathbb{R}^{L_D \times d} denote the contextual embeddings of the nn-th document. SeaReader constructs a matching matrix for each document:

Mn(i,j)=S(i)⊙Dn(j).M_n(i,j) = S(i) \odot D_n(j).

In the question-centric path, attention is applied column-wise over document tokens for each question token:

αn(i,j)=softmax⁡(Mn(i,1),…,Mn(i,LD))(j),\alpha_n(i,j) = \operatorname{softmax}(M_n(i,1), \ldots, M_n(i,L_D))(j),

SS0

For each token SS1, SS2 summarizes the most relevant positions in document SS3 aligned to that token. This realizes focused per-document reading conditioned on the statement (Zhang et al., 2018).

The document-centric path first applies row-wise attention to read related information in SS4 into each document token, summarized notationally as SS5. It then performs cross-document attention over concatenated document-question representations:

SS6

SS7

SS8

This cross-document attention extracts evidence from other documents conditional on token SS9 in document DnD_n0, enabling multi-document support aggregation rather than isolated passage matching.

SeaReader supplements attention with matching features DnD_n1 computed from the question-document matching matrix. Two alternatives are described. The CNN extractor uses two convolutional layers with max-pooling between and after layers, with the second layer employing dilated convolution to preserve word-level resolution. The lighter alternative uses row-wise and column-wise max-pooling and mean-pooling over the matching matrix. These features are combined with the attention reads in the reasoning module.

At the output layer, if DnD_n2 is the score assigned to candidate answer DnD_n3, then the final probability is

DnD_n4

with base objective

DnD_n5

The architecture also admits interpretability-oriented regularized variants, including a gating importance penalty, noisy gating, and DnD_n6-regularized delta embeddings. These additions are not separate modules layered on top of the model’s predictions; they are part of the training formulation used to expose token and document salience (Zhang et al., 2018).

4. Training regime and interpretability mechanisms

SeaReader uses 200-dimensional word embeddings trained with skip-gram on all training questions and the entire document collection. Unseen test words are mapped to zero vectors. To balance fit and generalization, the system refines embeddings through an DnD_n7-regularized delta term:

DnD_n8

with loss augmented by

DnD_n9

The contextual and reasoning BiLSTMs use hidden size 128 per direction; dropout is 0.2; the decision layer is a single-layer feed-forward scorer; and optimization uses Adam with SS0 and an exponentially decayed learning rate. Implementation is in TensorFlow. Batch size is 15, corresponding to approximately 750 documents per batch, described as the practical maximum on a single GPU (Zhang et al., 2018).

Because computation scales rapidly with both document count and sequence length, documents are truncated to 100 words before entering SeaReader. The paper states that this slightly reduces accuracy but substantially speeds up training. Retrieval typically keeps top-10 documents per candidate answer, although performance continues to improve up to top-20, at increasing computational cost.

Interpretability is treated as an architectural objective rather than only a post hoc visualization step. First, SeaReader employs gating with an importance penalty, in which average gate values above threshold SS1 are penalized, making gate magnitude a proxy for token or document importance:

SS2

Second, noisy gating adds Gaussian noise after gating,

SS3

so that high-gate features remain visible over noise and low-gate features are deemphasized. Third, attention visualization is provided for the column-wise question-centric attention, the row-wise attention, and the cross-document attention, allowing inspection of matching patterns and evidence aggregation. The paper reports that the SS4-regularized delta embeddings surface logical connectives and common diagnostic terms as the most modified words, offering another lens on task salience (Zhang et al., 2018).

5. Empirical performance on MedQA

SeaReader substantially improves over the compared baselines on MedQA. On the validation and test sets, the single model achieves 73.9% and 73.6% accuracy, respectively, while the ensemble reaches 75.8% and 75.3%. The paper reports test accuracies of 59.3% for Iterative Attention, 53.0% for Neural Reasoner, and 64.5% for R-NET. A human passing score is given as 60.0% (360/600), as a reference point rather than a like-for-like model comparison (Zhang et al., 2018).

System Validation Test
SeaReader 73.9% 73.6%
SeaReader (ensemble) 75.8% 75.3%
Iterative Attention — 59.3%
Neural Reasoner — 53.0%
R-NET — 64.5%

Performance remains strong across MedQA categories. On the test set, SeaReader obtains 0.754 on A1, 0.738 on B1, 0.737 on A2, and 0.707 on A3/A4. The corresponding R-NET numbers are 0.671, 0.617, 0.663, and 0.618. The largest relative gains appear in B1 and A3/A4, where information is mixed or shared across questions and documents. This suggests that the dual-path attention is especially useful when evidence must be integrated across multiple textual contexts rather than extracted from a single, self-contained passage.

The effect of retrieval depth is particularly informative. With top-1 document per candidate, test accuracy is 57.8%, and the relevant document ratio is 0.90. At top-5, accuracy rises to 71.7% despite the relevant ratio dropping to 0.54. At top-10, SeaReader reaches 73.6% with relevant ratio 0.46, and at top-20 it reaches 74.4% with relevant ratio 0.29. Performance therefore improves as more documents are supplied even though the fraction of relevant documents decreases. A plausible implication is that SeaReader’s cross-document filtering and fusion are effective enough to convert retrieval breadth into additional evidence rather than only additional noise (Zhang et al., 2018).

6. Limitations, comparative position, and example reasoning behavior

SeaReader’s main limitations are tied to retrieval, truncation, and difficult inference chains. Accuracy depends on the retrieval stage’s ability to surface relevant documents. As the number of documents increases, the fraction of relevant ones decreases, and gains eventually show diminishing returns because of noise. Document truncation to 100 tokens improves training efficiency but can remove late-occurring evidence or long-range dependencies. Some clinical questions require nuanced reasoning over many facts, and although cross-document attention helps, extremely long chains and subtle causal mechanisms remain challenging. Rare specialist terminology and atypical cases may also be poorly represented in the embeddings, even with SS5 delta adjustments (Zhang et al., 2018).

Within the reading-comprehension literature, SeaReader differs from span-extraction systems such as BiDAF, AoA Reader, and R-NET in three explicit ways. It operates over many retrieved documents rather than a pre-selected paragraph, uses dual-path attention to couple intra-document reading with inter-document integration, and produces candidate-level scores instead of answer spans. Relative to Iterative Attention and Neural Reasoner, its explicit cross-document attention and modular gating provide stronger multi-document evidence integration on real-world, non-cloze medical questions.

The Parkinson’s disease example in the paper illustrates the architecture’s intended behavior. One retrieved document states that commonly used drugs include phenanthrene, amantadine, levodopa, and compound levodopa. Another states that phenanthrene is mainly suitable for those with obvious tremble but is more used in young patients and should be used with caution in elderly patients. SeaReader’s question-centric attention highlights the treatment options relative to the candidate answers, while the document-centric path integrates the age condition from the question into each document and aggregates the cautionary evidence across documents. Document-level gating then emphasizes the caution against phenanthrene for elderly patients together with the broader effectiveness of levodopa-based treatments, and the pooled support favors the levodopa option (Zhang et al., 2018).

SeaReader therefore occupies a specific position in large-scale machine reading: it is a multi-document, multiple-choice clinical QA model whose principal innovation is not merely higher-capacity encoding, but an explicit decomposition of reading into per-document alignment and cross-document evidence synthesis. Its interpretability mechanisms, including gating penalties, noisy gating, delta embeddings, and attention visualization, further distinguish it as an attempt to make multi-document decision formation inspectable rather than only accurate.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to SeaReader.