---
title: 'PEM-Rel-8K: Multidisciplinary Semantic Relations'
url: https://www.emergentmind.com/topics/pem-rel-8k
type: topic
---

# PEM-Rel-8K: Multidisciplinary Semantic Relations

Searching arXiv for the specified paper and closely related benchmark context.
PEM-Rel-8K is a modular, multidisciplinary benchmark for predicting semantic relations between pairs of research topics. Introduced in “Leveraging Large Language Models for Generating Research Topic Ontologies: A Multi-Disciplinary Study” [2508.20693], it operationalizes a core ontology-generation subproblem as single-label, four-class classification over topic pairs drawn from three widely used academic knowledge organization systems: MeSH for biomedicine, PhySH for physics, and the IEEE Thesaurus for engineering and computer science. The benchmark contains **8,075 relations** in total and is designed for domain-specific evaluation, cross-domain transfer, and joint multidisciplinary training. Its central premise is that broader, narrower, same-as, and negative relations are foundational for ontology and taxonomy construction, extension, and alignment.

## 1. Conceptual scope and task definition

PEM-Rel-8K is organized around a single task: given two research topics \(t_A\) and \(t_B\), assign exactly one semantic relation label. The task is defined as a **single-label, multi-class classification problem**, with four mutually exclusive labels [2508.20693].

The labels are:

- **broader**: the first topic subsumes the second.
- **narrower**: the first topic is subsumed by the second.
- **same-as**: the two topics are semantically equivalent or interchangeable.
- **other**: none of the above relations holds.

The paper gives illustrative examples. For **broader**, `databases` subsumes `distributed databases`. For **narrower**, `adaptive signal processing` is subsumed by `signal processing`. For **same-as**, `ontology alignment` and `ontology matching` are treated as equivalent. The inclusion of **other** is explicitly motivated as a negative class that prevents a classifier from being forced to hallucinate a semantic relation for unrelated pairs.

This label set combines hierarchical structure with equivalence and reliable negatives. The paper states that the first three labels are semantic relations, while `other` is a negative class. It also notes that broader/narrower map naturally to hierarchy and that same-as captures synonymy or equivalence, making the task closely aligned with SKOS-style knowledge organization. This suggests PEM-Rel-8K is designed not merely as a generic relation-classification corpus, but as a benchmark centered on ontology-relevant relation types.

## 2. Dataset composition and source ontologies

The benchmark is built from three existing, widely used taxonomies, each corresponding to a different academic domain [2508.20693]. The initials of these sources give the benchmark its name:

- **P**: PhySH / Physics
- **E**: Engineering via IEEE
- **M**: MeSH / Biomedicine

The three component datasets are:

| Component dataset | Source ontology | Domain |
|---|---|---|
| IEEE-Rel-3K | IEEE Thesaurus, RDF v1.02 | Engineering and computer science |
| PhySH-Rel-875 | PhySH, RDF following SKOS, April 2024 | Physics |
| MeSH-Rel-4K | MeSH, January 2025 release | Biomedicine |

These component sets can be used independently or merged into the full benchmark. Their sizes are:

- **IEEE-Rel-3K**: **3,200**
- **PhySH-Rel-875**: **875**
- **MeSH-Rel-4K**: **4,000**
- **PEM-Rel-8K**: **8,075**

The paper describes the aggregate benchmark as “over 8,000 relationships.” In the merged set, the domain proportions are **39.6%** IEEE, **10.8%** PhySH, and **49.6%** MeSH.

Each example effectively consists of **Topic A**, **Topic B**, and **Label** \(\in\) {`broader`, `narrower`, `same-as`, `other`}. For fine-tuning, topic pairs are cast into an instruction format:

```text
Classify the relationship between `[TOPIC-A]' and `[TOPIC-B]'
```

with the expected output:

```text
relationship: [RELATIONSHIP-TYPE]
```

An example training instance given in the paper is:

```text
user: Classify the relationship between `Biology' and `Genetics'
model: relationship: broader
```

The benchmark does not provide a full tabular schema with field names, and it does not describe additional normalization procedures such as deduplication rules, case-folding, lemmatization, or URI retention beyond the sampling process.

## 3. Construction methodology and labeling

PEM-Rel-8K was assembled through extraction from authoritative ontologies, manual validation of ambiguous equivalence candidates, generation of ontology-aware negatives, and standard train/validation/test partitioning [2508.20693]. The construction pipeline is described as:

1. select source release/version,
2. extract broader/narrower relations from hierarchy properties,
3. inspect alt-label or related-concept structures for synonym/equivalence candidates,
4. manually validate same-as where necessary,
5. generate random unrelated pairs for `other`,
6. split each domain dataset **7:1:2**,
7. merge all three into PEM-Rel-8K.

### IEEE-Rel-3K

The engineering set samples **3,200** relations:

- **800 broader**
- **800 narrower**
- **800 same-as**
- **800 other**

Broader and narrower were sampled directly from the IEEE Thesaurus. Same-as required manual supervision because the thesaurus does **not explicitly encode synonymy as a dedicated same-as relation**; instead it uses `skos:prefLabel` and `skos:altLabel`. The paper emphasizes that alternative labels are not always true synonyms. To construct reliable equivalence pairs, **three experts manually analyzed** topics connected through these label fields and extracted **800 pairs** judged to be true synonyms, lexical variants, or near-synonyms. Negative examples were built by randomly generating **800 topic pairs that do not share any semantic link in the original ontology**. The example `'4G mobile communication'` as a `skos:altLabel` of `'5G mobile communication'` is used to show why manual review was necessary.

### PhySH-Rel-875

The physics set contains **875** relations:

- **250 broader**
- **250 narrower**
- **125 same-as**
- **250 other**

This is the only imbalanced subset. The imbalance is explicitly attributed to the difficulty of validating enough trustworthy same-as cases. PhySH also uses `skos:altLabel` without an explicit synonymy relation. The authors reviewed **608 available `skos:altLabel` entries**, and **three experts** manually validated **125** as genuine same-as relations. Negative pairs were again generated by random selection subject to the constraint that the topics share no semantic link in the ontology. The example `algebraic structure` as an `skos:altLabel` of `abstract algebra` illustrates that alternative labels are not automatically equivalent.

### MeSH-Rel-4K

The biomedical set contains **4,000** relations:

- **1,000 broader**
- **1,000 narrower**
- **1,000 same-as**
- **1,000 other**

MeSH uses its own `mesh:` schema rather than pure SKOS. Topics are `mesh:TopicalDescriptor`, hierarchy is represented via `mesh:broaderDescriptor`, and equivalence-like relations are represented via `mesh:relatedConcept`. Broader and narrower were extracted from `mesh:broaderDescriptor`; same-as was extracted from `mesh:relatedConcept`; and **1,000 unrelated topic pairs** were selected for `other`.

Across the benchmark, the strongest quality-control mechanism is expert review of ambiguous equivalence candidates. The paper also stresses that preserving an imbalanced same-as class in PhySH-Rel-875 was preferable to artificially expanding it at lower quality. This suggests that construction decisions were explicitly shaped by annotation reliability rather than by uniform class balancing alone.

## 4. Experimental role and evaluation protocol

PEM-Rel-8K is the central benchmark used to evaluate whether large language models can identify semantic relations between research topics and whether such capability transfers across scientific domains [2508.20693]. Its modular structure supports three experimental scenarios:

- **domain-specific evaluation**
- **cross-domain evaluation**
- **multidisciplinary evaluation**

Each domain-specific dataset and the merged benchmark use a **train / validation / test = 7:1:2** split. The resulting sizes are:

| Dataset | Train | Validation | Test |
|---|---:|---:|---:|
| IEEE-Rel-3K | 2,240 | 320 | 640 |
| PhySH-Rel-875 | 613 | 87 | 175 |
| MeSH-Rel-4K | 2,800 | 400 | 800 |
| PEM-Rel-8K | 5,653 | 807 | 1,615 |

The study evaluates **12 decoder-only open-weight LLMs** ranging from **2.51B** to **27.2B** parameters:

- mistral-7b-instruct-v0.3
- Mistral-Nemo-Instruct-2407
- Mistral-Small-Instruct-2409
- Llama-3.2-3B-Instruct
- llama-2-7b-chat
- Meta-Llama-3.1-8B-Instruct
- gemma-2b-it
- gemma-2-9b-it
- gemma-2-27b-it
- Phi-3.5-mini-instruct
- phi-4
- zephyr-sft

All models were **quantized to 4-bit**, fine-tuned with **LoRA**, and run using **KoboldAI** for zero-shot settings and **Unsloth** for fine-tuning.

Three evaluation modes are studied:

### Standard zero-shot prompting

A fixed prompt defines the task, explains the four labels, and constrains the output format.

### Bidirectional Chain-of-Thought (bCoT)

This stronger zero-shot baseline uses a two-prompt reasoning procedure in both topic orders. The model is asked to define the topics, place them in a sentence, reflect on their relation, and then produce a final classification. The process is repeated with reversed topic order, and a **rule-based referee** resolves disagreements.

### Fine-tuning

Each model is instruction-tuned on one of four datasets:

- IEEE-Rel-3K
- MeSH-Rel-4K
- PhySH-Rel-875
- PEM-Rel-8K

The standardized fine-tuning design yields **12 LLMs × 4 training datasets = 48 fine-tuned models**. Across training/testing combinations, the modular design enables **24 train/test combinations** and **288 total runs**, consisting of **96 zero-shot experiments** and **192 fine-tuning experiments**. Evaluation uses **Precision**, **Recall**, and **F1-score**. The paper does not provide explicit equations for these metrics, and it does not report few-shot prompting.

## 5. Empirical performance and error structure

The central empirical result is that fine-tuning on PEM-Rel-8K yields much stronger performance than prompt-only approaches across all disciplines [2508.20693]. The paper reports that **Bidirectional CoT consistently outperformed standard prompting**, improving average F1 by **5.1%** over standard prompting across the four datasets, with standard deviation **\(\pm 0.3\%\)**. On the full PEM-Rel-8K benchmark, the best standard prompting result was **F1 = 0.677** using **mistral-22b**, while the best bidirectional CoT result was **F1 = 0.730** using **mistral-7b**.

Fine-tuning is substantially stronger. When trained and tested on full PEM-Rel-8K, the best result is obtained by **gemma-27b**:

- **F1 = 0.935**
- **Precision = 0.935**
- **Recall = 0.935**

Other strong full-benchmark fine-tuned results are:

- **gemma-9b**: **F1 = 0.926**
- **mistral-22b**: **F1 = 0.922**
- **phi-4**: **F1 = 0.918**
- **zephyr-7b**: **F1 = 0.915**

### Domain-specific and multidisciplinary results

Best in-domain fine-tuned models are:

- **IEEE-Rel-3K**: **gemma-27b**, **F1 = 0.989**
- **MeSH-Rel-4K**: **phi-4**, **F1 = 0.917**
- **PhySH-Rel-875**: **gemma-9b**, **F1 = 0.936**

When models are trained on full PEM-Rel-8K and tested on individual domains, the best reported results are:

- **IEEE-Rel-3K**: **gemma-27b**, **F1 = 0.973**
- **MeSH-Rel-4K**: **gemma-27b**, **F1 = 0.908**
- **PhySH-Rel-875**: **gemma-27b**, **F1 = 0.925**

The paper highlights that models trained on full PEM-Rel-8K were, on average, only **1.2% lower in F1** than the best discipline-specific models. This is one of the benchmark’s most consequential findings: a single multidisciplinary training set can nearly match separate per-domain tuning.

### Cross-domain transfer

Cross-domain transfer is a primary design goal of PEM-Rel-8K. The best cross-domain results are:

- **Test on IEEE-Rel-3K**, trained on **PhySH-Rel-875**: **phi-4**, **F1 = 0.947**
- **Test on MeSH-Rel-4K**, trained on **PhySH-Rel-875**: **gemma-27b**, **F1 = 0.877**
- **Test on PhySH-Rel-875**, trained on **MeSH-Rel-4K**: **phi-4**, **F1 = 0.864**
- **Test on full PEM-Rel-8K**, best single-domain-trained transfer result: trained on **MeSH-Rel-4K**, **phi-4**, **F1 = 0.919**

The aggregate transfer result is that the best cross-domain model for each discipline was on average only **5.1% F1** below the best in-domain model, with standard deviation **\(\pm 1.7\%\)**, and cross-domain fine-tuned models improved by **16.1 percentage points** over the best zero-shot strategy, with standard deviation **\(\pm 1.6\%\)**.

### Per-relation performance and errors

For models fine-tuned and evaluated on PEM-Rel-8K, the paper reports that **other** is the easiest class and **same-as** is the hardest. For **gemma-27b**, the per-label F1 scores are:

- **broader**: **0.935**
- **narrower**: **0.922**
- **other**: **0.966**
- **same-as**: **0.916**

Averaged across all 12 fine-tuned models on PEM-Rel-8K:

- **broader**: **0.904**
- **narrower**: **0.893**
- **other**: **0.943**
- **same-as**: **0.885**

The main reported confusion pattern is between hierarchical relations and equivalence. Two examples are emphasized:

- In MeSH, **“microsporea”** is more specific than **“microsporidians”**, but all top models predicted `same-as`.
- In IEEE, **“Nanoscale technology”** is more specific than **“Nanotechnology”**, yet all five top models treated them as `same-as`.

This suggests that high lexical overlap remains a major source of error, especially when hierarchy and near-equivalence are difficult to separate lexically.

## 6. Significance, applications, and limitations

According to the authors, PEM-Rel-8K is the **first large-scale benchmark specifically designed to study LLM-based identification of semantic relationships between research topics across multiple scientific disciplines** [2508.20693]. Its significance lies in three properties: it operationalizes a core ontology-learning subtask, spans multiple disciplines, and enables transfer-learning analysis rather than only within-domain evaluation.

The benchmark’s principal strengths are explicitly identified in the paper:

- **multidisciplinary coverage** across engineering/computer science, physics, and biomedicine;
- **modular design**, allowing per-domain, merged, and transfer experiments;
- **grounding in authoritative ontologies**, namely IEEE, PhySH, and MeSH;
- **coverage of hierarchy, equivalence, and negatives** rather than only positive hierarchical links;
- **human quality control** for ambiguous same-as candidates;
- **strong empirical usefulness**, especially under fine-tuning.

These features support several applications. The paper frames PEM-Rel-8K as useful for **semantic relation prediction**, **ontology generation and expansion**, **taxonomy alignment and integration**, and broader **research knowledge organization**. Better topic ontologies can support repositories, digital libraries, search engines, recommender systems, analytics dashboards, conversational agents, and research profiling systems. Because the benchmark includes `same-as`, `broader`, `narrower`, and `other`, it is suited to workflows where both hierarchical and equivalence structure must be inferred without forcing a semantic relation onto every pair.

The limitations are also explicit. Domain coverage is restricted to **three STEM fields**, and the paper notes that future work should extend to additional STEM areas and later to the social sciences, humanities, and linguistics. The **same-as** label is inherently difficult because equivalence is interpreted inconsistently across ontologies and thesauri. **PhySH-Rel-875** is imbalanced because reliable equivalence cases were scarce. Positive transfer results are shown only among relatively related STEM domains, so transfer to conceptually distant disciplines remains open. The benchmark addresses **pairwise relation prediction**, not end-to-end ontology induction, conflict resolution, logical consistency enforcement, or structural optimization. Finally, the paper does not fully document field-level schema details or normalization rules.

PEM-Rel-8K therefore occupies a specific methodological position. It is not a full ontology-construction system, but a benchmark for one of its central inference primitives: assigning ontology-relevant semantic relations between topic pairs. Its reported results indicate that this primitive can be learned effectively in a multidisciplinary setting, and that the learned signal transfers nontrivially across domains.

Source: https://www.emergentmind.com/topics/pem-rel-8k