RareDis Corpus: Gold Standard for Rare Disease NLP
- RareDis Corpus is a public, English biomedical dataset annotated for entities like diseases, rare diseases, signs, symptoms, and anaphors with detailed relation links.
- It comprises 1,041 documents and 9,141 sentences, achieving an inter-annotator agreement of 83.5% F1 for entities and 81.3% for relations.
- Its document-centric design supports complex tasks including nested, discontinuous, and cross-sentence relation detection for improved rare disease analysis.
Searching arXiv for the cited papers to ground the article in the primary source and related later work. RareDis Corpus is an English biomedical corpus drawn from the NORD database and annotated for rare-disease information extraction at entity and relation level (Martínez-deMiguel et al., 2021). It contains 1,041 texts, 9,141 sentences, and 192,041 tokens, with annotations for diseases, rare diseases, signs, symptoms, and anaphors, together with relations such as clinical manifestation, taxonomic, synonymic, acronymic, and anaphoric links (Martínez-deMiguel et al., 2021). More than 5,000 rare diseases and almost 6,000 clinical manifestations are annotated, and the reported inter-annotator agreement reaches 83.5% F1 for entities and 81.3% for relations under exact match criteria, positioning the resource as a public gold-standard corpus for BioNLP on rare diseases (Martínez-deMiguel et al., 2021).
1. Corpus scope and provenance
RareDis Corpus is composed of 1,041 documents, each corresponding to a unique rare disease entry. The corpus statistics are 9,141 sentences and 192,041 tokens. The standard split is 729 training documents, 104 validation documents, and 208 test documents (Martínez-deMiguel et al., 2021).
| Subset | Documents | Sentences |
|---|---|---|
| Training | 729 | 6,451 |
| Validation | 104 | 903 |
| Test | 208 | 1,787 |
| Total | 1,041 | 9,141 |
The texts were sourced from the NORD database, covering 1,200+ rare diseases, and were collected from the seven main descriptive sections for each rare disease. These sections include general discussion, signs and symptoms, causes, diagnosis, and related material. The corpus language is English, and the resource is publicly available with guidelines at https://github.com/isegura/NLP4RARE-CM-UC3M (Martínez-deMiguel et al., 2021).
The resulting design is document-centric rather than sentence-centric: each text corresponds to a disease entry, which makes the corpus suitable for both local extraction tasks and document-level phenomena such as anaphora and cross-sentence relation detection. This is significant in rare-disease NLP, where publicly available annotated resources have been scarce (Martínez-deMiguel et al., 2021).
2. Annotation schema
The entity layer comprises five types. Disease denotes an abnormal condition of an organism and is not necessarily rare. Rare disease denotes a disease affecting a small proportion of the population, following the Orphanet/EU definition. Symptom denotes a patient-experienced, subjective indicator. Sign denotes an objective abnormality observed by clinicians or tests. Anaphor denotes a word or phrase referring to a disease entity, including forms such as “This disease” or “These diseases” (Martínez-deMiguel et al., 2021).
The relation layer comprises six types. produces links a disease to a sign or symptom. increases risk of links a disease to another disorder for which it raises risk. is a marks a taxonomic relation. is acron marks an acronym relation. is synon marks a synonym relation between disease names. anaphora links an anaphor to its antecedent disease (Martínez-deMiguel et al., 2021).
| Layer | Type | Total |
|---|---|---|
| Entity | Disease | 2,348 |
| Entity | Rare Disease | 5,221 |
| Entity | Symptom | 396 |
| Entity | Sign | 5,333 |
| Entity | Anaphor | 1,535 |
| Relation | produces | 5,793 |
| Relation | increases risk of | 245 |
| Relation | is a | 975 |
| Relation | is acron | 288 |
| Relation | is synon | 111 |
| Relation | anaphora | 1,543 |
The most frequent entity type is sign, and the most frequent relation type is produces (Martínez-deMiguel et al., 2021). This distribution reflects the corpus objective: rare diseases are represented together with their clinical manifestations, rather than only as terminological mentions. A common misconception is that RareDis is principally a rare-disease name corpus; in fact, its largest annotation mass is in manifestation and manifestation-linking categories.
3. Annotation workflow and decision rules
Data collection was performed by web scraping from NORD, focusing on the seven main descriptive sections per rare disease. Pre-annotation was then carried out automatically with a dictionary-based procedure using Disease Ontology (DOID), Orphan Rare Disease Ontology (ORDO), and Symptom Ontology (SYMP), with spaCy used for automated entity spotting. This pre-annotation stage automatically found 3,003 diseases, 2,542 rare diseases, and 1,560 symptoms (Martínez-deMiguel et al., 2021).
Manual annotation involved four experts: two with expertise in rare diseases or biomedicine and two with experience in biomedical corpus annotation. Annotation was performed with the BRAT annotation tool. Signs and anaphors were fully annotated manually, and all relations were annotated manually from scratch. The guidelines were developed iteratively after a pilot phase and disagreement analysis, and the final guidelines were released as supplementary material (Martínez-deMiguel et al., 2021).
Several difficult cases were addressed by explicit decision rules. Nested entities were annotated when inner and outer spans belonged to different types, as in “central pain syndrome” as a disease and “pain” as a symptom. When multiple labels were possible, only the more specific label was assigned, with rare disease preferred over disease. Discontinuous entities were marked, as in “malformations of the abdominal wall.” General terms such as “disorder” or “disease” were not annotated unless modified by an adjective, as in “neurological disorder.” Acronyms and synonyms were annotated only with first appearance, except when explicit in the text (Martínez-deMiguel et al., 2021).
The distinction between signs and symptoms was treated as a central methodological issue. Signs were defined as observable and objective; symptoms were defined as subjective and patient-reported. The paper identifies this demarcation as a major source of ambiguity, which is consistent with the lower agreement later observed for the sign category (Martínez-deMiguel et al., 2021).
4. Inter-annotator agreement and quality assessment
Corpus quality was evaluated with inter-annotator agreement on a random sample of 51 texts. The metric was F1-measure rather than Kappa, because Kappa is not optimal for NER and discontinuous spans. Under the exact match criterion, agreement required both mention boundaries and types to match; for relations, it required the identical pair of entities and the same relation type (Martínez-deMiguel et al., 2021).
The agreement study followed a multi-stage process. The pilot annotation yielded an initial average of 62.6%, with sign at 48%. After analysis of disagreements and refinement of the guidelines, the corpus was re-annotated and a final IAA was calculated (Martínez-deMiguel et al., 2021).
The final overall entity IAA is 83.5% F1. By type, the scores are 91.2% for anaphor, 90.9% for symptom, 83.4% for disease, 81.4% for rare disease, and 67.3% for sign. The final overall relation IAA is 81.3% F1. By type, anaphora reaches 90.8%, produces 83.1%, and is synon 60%; increases risk of is reported as having low agreement because it is context-sensitive (Martínez-deMiguel et al., 2021).
These figures are described as similar to or above those of other biomedical corpora (Martínez-deMiguel et al., 2021). Within the corpus itself, the pattern of scores is methodologically informative: anaphora and symptom annotations are comparatively stable, while signs and low-frequency relations remain harder. This suggests that RareDis is both a resource and a benchmark for the edge cases that standard biomedical NER pipelines often treat inadequately, particularly discontinuous spans, nested mentions, and cross-sentence links.
5. Research functions and downstream relevance
RareDis Corpus was intended for training and evaluating BioNLP models for named entity recognition and relation extraction involving rare diseases and their manifestations (Martínez-deMiguel et al., 2021). Because it includes anaphors and anaphora relations, it also supports research in anaphora resolution, a comparatively under-resourced area in biomedical NLP. The corpus further supports improvement of rare disease knowledge bases and can be used to accelerate diagnosis, reduce diagnostic delay, and guide treatment decisions through information extraction from unstructured biomedical texts (Martínez-deMiguel et al., 2021).
The resource is also explicitly positioned as a foundation for complex NER. It supports work on nested entities, discontinuous entities, and cross-sentence relation detection, all of which are difficult settings for conventional sequence-labeling approaches. The scarcity of annotations for low-frequency relation types such as is synon and increases risk of simultaneously makes those categories challenging for supervised learning and useful for methodological evaluation (Martínez-deMiguel et al., 2021).
The paper characterizes RareDis as the first comprehensive, public gold-standard corpus for rare diseases and their clinical manifestations (Martínez-deMiguel et al., 2021). It further states that the resource could open the door to further NLP applications that would facilitate the diagnosis and treatment of rare diseases and improve the quality of life of affected patients (Martínez-deMiguel et al., 2021). A plausible implication is that RareDis functions not only as a supervised-learning dataset but also as an infrastructure layer for machine reading, knowledge base construction, and clinical decision support in low-resource biomedical domains.
6. Relation to later rare-disease NLP resources
RareDis should be distinguished from later NORD-derived resources built for LLM evaluation and retrieval-augmented generation. A later study introduced the ReDis-QA dataset for rare disease question-answering and a NORD-derived corpus termed ReCOP, in which each disease report is split into seven thematic sections or chunks: overview, symptoms, causes, affects, related disorders, diagnosis, and standard therapies (Wang et al., 2024).
The distinction is architectural. RareDis is an annotated corpus with entity spans, relation instances, discontinuous and nested mentions, and anaphora links (Martínez-deMiguel et al., 2021). ReCOP, by contrast, is a property-chunked corpus intended to align retrieval units with question types in rare disease QA, and its use was reported to improve the accuracy of LLMs on ReDis-QA by an average of 8% while guiding them to generate trustworthy answers and explanations traceable to existing literature (Wang et al., 2024).
This later development clarifies the continuing relevance of RareDis. Rare-disease NLP now spans at least two complementary paradigms: fine-grained annotation for extraction and document understanding, exemplified by RareDis, and structured chunk retrieval for grounded generation, exemplified by ReCOP. The two paradigms are not interchangeable, but they address adjacent layers of the same problem space.