PEM-Rel-8K: Multidisciplinary Semantic Relations
- PEM-Rel-8K is a benchmark that defines semantic relations (broader, narrower, same-as, other) between paired research topics from authoritative domain ontologies.
- The dataset is constructed from MeSH, PhySH, and IEEE sources and is used for domain-specific, cross-domain, and multidisciplinary evaluations.
- Fine-tuning large language models on PEM-Rel-8K significantly improves F1 scores, demonstrating its value for ontology construction and taxonomy alignment.
Searching arXiv for the specified paper and closely related benchmark context. PEM-Rel-8K is a modular, multidisciplinary benchmark for predicting semantic relations between pairs of research topics. Introduced in “Leveraging LLMs for Generating Research Topic Ontologies: A Multi-Disciplinary Study” (Aggarwal et al., 28 Aug 2025), it operationalizes a core ontology-generation subproblem as single-label, four-class classification over topic pairs drawn from three widely used academic knowledge organization systems: MeSH for biomedicine, PhySH for physics, and the IEEE Thesaurus for engineering and computer science. The benchmark contains 8,075 relations in total and is designed for domain-specific evaluation, cross-domain transfer, and joint multidisciplinary training. Its central premise is that broader, narrower, same-as, and negative relations are foundational for ontology and taxonomy construction, extension, and alignment.
1. Conceptual scope and task definition
PEM-Rel-8K is organized around a single task: given two research topics and , assign exactly one semantic relation label. The task is defined as a single-label, multi-class classification problem, with four mutually exclusive labels (Aggarwal et al., 28 Aug 2025).
The labels are:
- broader: the first topic subsumes the second.
- narrower: the first topic is subsumed by the second.
- same-as: the two topics are semantically equivalent or interchangeable.
- other: none of the above relations holds.
The paper gives illustrative examples. For broader, databases subsumes distributed databases. For narrower, adaptive signal processing is subsumed by signal processing. For same-as, ontology alignment and ontology matching are treated as equivalent. The inclusion of other is explicitly motivated as a negative class that prevents a classifier from being forced to hallucinate a semantic relation for unrelated pairs.
This label set combines hierarchical structure with equivalence and reliable negatives. The paper states that the first three labels are semantic relations, while other is a negative class. It also notes that broader/narrower map naturally to hierarchy and that same-as captures synonymy or equivalence, making the task closely aligned with SKOS-style knowledge organization. This suggests PEM-Rel-8K is designed not merely as a generic relation-classification corpus, but as a benchmark centered on ontology-relevant relation types.
2. Dataset composition and source ontologies
The benchmark is built from three existing, widely used taxonomies, each corresponding to a different academic domain (Aggarwal et al., 28 Aug 2025). The initials of these sources give the benchmark its name:
- P: PhySH / Physics
- E: Engineering via IEEE
- M: MeSH / Biomedicine
The three component datasets are:
| Component dataset | Source ontology | Domain |
|---|---|---|
| IEEE-Rel-3K | IEEE Thesaurus, RDF v1.02 | Engineering and computer science |
| PhySH-Rel-875 | PhySH, RDF following SKOS, April 2024 | Physics |
| MeSH-Rel-4K | MeSH, January 2025 release | Biomedicine |
These component sets can be used independently or merged into the full benchmark. Their sizes are:
- IEEE-Rel-3K: 3,200
- PhySH-Rel-875: 875
- MeSH-Rel-4K: 4,000
- PEM-Rel-8K: 8,075
The paper describes the aggregate benchmark as “over 8,000 relationships.” In the merged set, the domain proportions are 39.6% IEEE, 10.8% PhySH, and 49.6% MeSH.
Each example effectively consists of Topic A, Topic B, and Label {broader, narrower, same-as, other}. For fine-tuning, topic pairs are cast into an instruction format:
1 |
Classify the relationship between `[TOPIC-A]' and `[TOPIC-B]' |
with the expected output:
1 |
relationship: [RELATIONSHIP-TYPE] |
An example training instance given in the paper is:
1 2 |
user: Classify the relationship between `Biology' and `Genetics' model: relationship: broader |
The benchmark does not provide a full tabular schema with field names, and it does not describe additional normalization procedures such as deduplication rules, case-folding, lemmatization, or URI retention beyond the sampling process.
3. Construction methodology and labeling
PEM-Rel-8K was assembled through extraction from authoritative ontologies, manual validation of ambiguous equivalence candidates, generation of ontology-aware negatives, and standard train/validation/test partitioning (Aggarwal et al., 28 Aug 2025). The construction pipeline is described as:
- select source release/version,
- extract broader/narrower relations from hierarchy properties,
- inspect alt-label or related-concept structures for synonym/equivalence candidates,
- manually validate same-as where necessary,
- generate random unrelated pairs for
other, - split each domain dataset 7:1:2,
- merge all three into PEM-Rel-8K.
IEEE-Rel-3K
The engineering set samples 3,200 relations:
- 800 broader
- 800 narrower
- 800 same-as
- 800 other
Broader and narrower were sampled directly from the IEEE Thesaurus. Same-as required manual supervision because the thesaurus does not explicitly encode synonymy as a dedicated same-as relation; instead it uses skos:prefLabel and skos:altLabel. The paper emphasizes that alternative labels are not always true synonyms. To construct reliable equivalence pairs, three experts manually analyzed topics connected through these label fields and extracted 800 pairs judged to be true synonyms, lexical variants, or near-synonyms. Negative examples were built by randomly generating 800 topic pairs that do not share any semantic link in the original ontology. The example '4G mobile communication' as a skos:altLabel of '5G mobile communication' is used to show why manual review was necessary.
PhySH-Rel-875
The physics set contains 875 relations:
- 250 broader
- 250 narrower
- 125 same-as
- 250 other
This is the only imbalanced subset. The imbalance is explicitly attributed to the difficulty of validating enough trustworthy same-as cases. PhySH also uses skos:altLabel without an explicit synonymy relation. The authors reviewed 608 available skos:altLabel entries, and three experts manually validated 125 as genuine same-as relations. Negative pairs were again generated by random selection subject to the constraint that the topics share no semantic link in the ontology. The example algebraic structure as an skos:altLabel of abstract algebra illustrates that alternative labels are not automatically equivalent.
MeSH-Rel-4K
The biomedical set contains 4,000 relations:
- 1,000 broader
- 1,000 narrower
- 1,000 same-as
- 1,000 other
MeSH uses its own mesh: schema rather than pure SKOS. Topics are mesh:TopicalDescriptor, hierarchy is represented via mesh:broaderDescriptor, and equivalence-like relations are represented via mesh:relatedConcept. Broader and narrower were extracted from mesh:broaderDescriptor; same-as was extracted from mesh:relatedConcept; and 1,000 unrelated topic pairs were selected for other.
Across the benchmark, the strongest quality-control mechanism is expert review of ambiguous equivalence candidates. The paper also stresses that preserving an imbalanced same-as class in PhySH-Rel-875 was preferable to artificially expanding it at lower quality. This suggests that construction decisions were explicitly shaped by annotation reliability rather than by uniform class balancing alone.
4. Experimental role and evaluation protocol
PEM-Rel-8K is the central benchmark used to evaluate whether LLMs can identify semantic relations between research topics and whether such capability transfers across scientific domains (Aggarwal et al., 28 Aug 2025). Its modular structure supports three experimental scenarios:
- domain-specific evaluation
- cross-domain evaluation
- multidisciplinary evaluation
Each domain-specific dataset and the merged benchmark use a train / validation / test = 7:1:2 split. The resulting sizes are:
| Dataset | Train | Validation | Test |
|---|---|---|---|
| IEEE-Rel-3K | 2,240 | 320 | 640 |
| PhySH-Rel-875 | 613 | 87 | 175 |
| MeSH-Rel-4K | 2,800 | 400 | 800 |
| PEM-Rel-8K | 5,653 | 807 | 1,615 |
The study evaluates 12 decoder-only open-weight LLMs ranging from 2.51B to 27.2B parameters:
- mistral-7b-instruct-v0.3
- Mistral-Nemo-Instruct-2407
- Mistral-Small-Instruct-2409
- Llama-3.2-3B-Instruct
- llama-2-7b-chat
- Meta-Llama-3.1-8B-Instruct
- gemma-2b-it
- gemma-2-9b-it
- gemma-2-27b-it
- Phi-3.5-mini-instruct
- phi-4
- zephyr-sft
All models were quantized to 4-bit, fine-tuned with LoRA, and run using KoboldAI for zero-shot settings and Unsloth for fine-tuning.
Three evaluation modes are studied:
Standard zero-shot prompting
A fixed prompt defines the task, explains the four labels, and constrains the output format.
Bidirectional Chain-of-Thought (bCoT)
This stronger zero-shot baseline uses a two-prompt reasoning procedure in both topic orders. The model is asked to define the topics, place them in a sentence, reflect on their relation, and then produce a final classification. The process is repeated with reversed topic order, and a rule-based referee resolves disagreements.
Fine-tuning
Each model is instruction-tuned on one of four datasets:
- IEEE-Rel-3K
- MeSH-Rel-4K
- PhySH-Rel-875
- PEM-Rel-8K
The standardized fine-tuning design yields 12 LLMs × 4 training datasets = 48 fine-tuned models. Across training/testing combinations, the modular design enables 24 train/test combinations and 288 total runs, consisting of 96 zero-shot experiments and 192 fine-tuning experiments. Evaluation uses Precision, Recall, and F1-score. The paper does not provide explicit equations for these metrics, and it does not report few-shot prompting.
5. Empirical performance and error structure
The central empirical result is that fine-tuning on PEM-Rel-8K yields much stronger performance than prompt-only approaches across all disciplines (Aggarwal et al., 28 Aug 2025). The paper reports that Bidirectional CoT consistently outperformed standard prompting, improving average F1 by 5.1% over standard prompting across the four datasets, with standard deviation . On the full PEM-Rel-8K benchmark, the best standard prompting result was F1 = 0.677 using mistral-22b, while the best bidirectional CoT result was F1 = 0.730 using mistral-7b.
Fine-tuning is substantially stronger. When trained and tested on full PEM-Rel-8K, the best result is obtained by gemma-27b:
- F1 = 0.935
- Precision = 0.935
- Recall = 0.935
Other strong full-benchmark fine-tuned results are:
- gemma-9b: F1 = 0.926
- mistral-22b: F1 = 0.922
- phi-4: F1 = 0.918
- zephyr-7b: F1 = 0.915
Domain-specific and multidisciplinary results
Best in-domain fine-tuned models are:
- IEEE-Rel-3K: gemma-27b, F1 = 0.989
- MeSH-Rel-4K: phi-4, F1 = 0.917
- PhySH-Rel-875: gemma-9b, F1 = 0.936
When models are trained on full PEM-Rel-8K and tested on individual domains, the best reported results are:
- IEEE-Rel-3K: gemma-27b, F1 = 0.973
- MeSH-Rel-4K: gemma-27b, F1 = 0.908
- PhySH-Rel-875: gemma-27b, F1 = 0.925
The paper highlights that models trained on full PEM-Rel-8K were, on average, only 1.2% lower in F1 than the best discipline-specific models. This is one of the benchmark’s most consequential findings: a single multidisciplinary training set can nearly match separate per-domain tuning.
Cross-domain transfer
Cross-domain transfer is a primary design goal of PEM-Rel-8K. The best cross-domain results are:
- Test on IEEE-Rel-3K, trained on PhySH-Rel-875: phi-4, F1 = 0.947
- Test on MeSH-Rel-4K, trained on PhySH-Rel-875: gemma-27b, F1 = 0.877
- Test on PhySH-Rel-875, trained on MeSH-Rel-4K: phi-4, F1 = 0.864
- Test on full PEM-Rel-8K, best single-domain-trained transfer result: trained on MeSH-Rel-4K, phi-4, F1 = 0.919
The aggregate transfer result is that the best cross-domain model for each discipline was on average only 5.1% F1 below the best in-domain model, with standard deviation , and cross-domain fine-tuned models improved by 16.1 percentage points over the best zero-shot strategy, with standard deviation .
Per-relation performance and errors
For models fine-tuned and evaluated on PEM-Rel-8K, the paper reports that other is the easiest class and same-as is the hardest. For gemma-27b, the per-label F1 scores are:
- broader: 0.935
- narrower: 0.922
- other: 0.966
- same-as: 0.916
Averaged across all 12 fine-tuned models on PEM-Rel-8K:
- broader: 0.904
- narrower: 0.893
- other: 0.943
- same-as: 0.885
The main reported confusion pattern is between hierarchical relations and equivalence. Two examples are emphasized:
- In MeSH, “microsporea” is more specific than “microsporidians”, but all top models predicted
same-as. - In IEEE, “Nanoscale technology” is more specific than “Nanotechnology”, yet all five top models treated them as
same-as.
This suggests that high lexical overlap remains a major source of error, especially when hierarchy and near-equivalence are difficult to separate lexically.
6. Significance, applications, and limitations
According to the authors, PEM-Rel-8K is the first large-scale benchmark specifically designed to study LLM-based identification of semantic relationships between research topics across multiple scientific disciplines (Aggarwal et al., 28 Aug 2025). Its significance lies in three properties: it operationalizes a core ontology-learning subtask, spans multiple disciplines, and enables transfer-learning analysis rather than only within-domain evaluation.
The benchmark’s principal strengths are explicitly identified in the paper:
- multidisciplinary coverage across engineering/computer science, physics, and biomedicine;
- modular design, allowing per-domain, merged, and transfer experiments;
- grounding in authoritative ontologies, namely IEEE, PhySH, and MeSH;
- coverage of hierarchy, equivalence, and negatives rather than only positive hierarchical links;
- human quality control for ambiguous same-as candidates;
- strong empirical usefulness, especially under fine-tuning.
These features support several applications. The paper frames PEM-Rel-8K as useful for semantic relation prediction, ontology generation and expansion, taxonomy alignment and integration, and broader research knowledge organization. Better topic ontologies can support repositories, digital libraries, search engines, recommender systems, analytics dashboards, conversational agents, and research profiling systems. Because the benchmark includes same-as, broader, narrower, and other, it is suited to workflows where both hierarchical and equivalence structure must be inferred without forcing a semantic relation onto every pair.
The limitations are also explicit. Domain coverage is restricted to three STEM fields, and the paper notes that future work should extend to additional STEM areas and later to the social sciences, humanities, and linguistics. The same-as label is inherently difficult because equivalence is interpreted inconsistently across ontologies and thesauri. PhySH-Rel-875 is imbalanced because reliable equivalence cases were scarce. Positive transfer results are shown only among relatively related STEM domains, so transfer to conceptually distant disciplines remains open. The benchmark addresses pairwise relation prediction, not end-to-end ontology induction, conflict resolution, logical consistency enforcement, or structural optimization. Finally, the paper does not fully document field-level schema details or normalization rules.
PEM-Rel-8K therefore occupies a specific methodological position. It is not a full ontology-construction system, but a benchmark for one of its central inference primitives: assigning ontology-relevant semantic relations between topic pairs. Its reported results indicate that this primitive can be learned effectively in a multidisciplinary setting, and that the learned signal transfers nontrivially across domains.