---
title: 'SAMER Lexicon: Arabic Readability Resource'
url: https://www.emergentmind.com/topics/samer-lexicon
type: topic
---

# SAMER Lexicon: Arabic Readability Resource

Searching arXiv for the most relevant papers on the SAMER lexicon and closely related Arabic readability resources.
The **SAMER Lexicon** is a **lemma-based readability lexicon** for Arabic that functions as the lexical backbone of the **Simplification of Arabic Masterpieces for Extensive Reading** resource ecosystem. In the literature, it is described as a **large-scale leveled readability lexicon** for **Modern Standard Arabic**, organized into graded difficulty levels and used both for **word-level readability annotation** and for **lexically controlled simplification**. In the SAMER corpus, it anchors document- and word-level annotation and constrains simplification toward target learner levels; in later modeling work, it is used as a source of explicit lexical difficulty knowledge, contributing **average readability**, **frequency of occurrence**, and **part-of-speech tags** for matched lemmas [2404.18615][2509.22870].

## 1. Definition and Resource Scope

SAMER is presented as part of a broader Arabic pedagogical NLP infrastructure rather than as an isolated word list. The corpus paper defines simplification primarily as **lexically controlled rewriting**: replacing difficult words with simpler alternatives of equivalent meaning, while keeping the text grammatically correct and semantically faithful. A key design decision is that the corpus is anchored in a **readability-leveled lexicon** rather than in free-form simplification [2404.18615].

The lexicon is described as having been **initially developed by Al Khalil et al. (2020)** and **later extended by Jiang et al. (2020)**. In this formulation, SAMER is a readability resource centered on **lemmas**, not merely surface word forms. The corpus paper explicitly characterizes it as a **controlled readability infrastructure** used for automatic analysis and guided simplification, and also notes related tooling such as a **Google Docs add-on** by **Hazim et al. (2022)** for word-level readability visualization and simplification support, with **Arabic WordNet-based substitution suggestions** and morphological disambiguation and lemmatization tools used to map surface forms to lemmas [2404.18615].

A common misconception is to treat SAMER as a conventional dictionary. The published descriptions are narrower and more technical: it is a **graded readability lexicon** whose central function is to attach difficulty metadata to lemmas and to support downstream annotation and modeling.

## 2. Readability Representation and Lexical Schema

The lexicon contains **26K lemmas** in its original version and was later extended to **40K lemmas**. These lemmas are organized into **five readability levels**, from **L1** to **L5**, where **L1** is described as low difficulty / easy readability and **L5** as high difficulty / hard readability [2404.18615].

| Level | Characterization | Grade / age mapping |
|---|---|---|
| L1 | basic words like *house, tree, to make, but* | not specified in the mapping |
| L2 | easier than upper levels | grades **2–3**, ages **7–8** |
| L3 | intermediate readability | grades **4–5**, ages **9–10** |
| L4 | advanced school readability | grades **6–8**, ages **11–14** |
| L5 | hardest readability | grades **9+**, ages **15+** |

Within the later readability-prediction paper, SAMER is summarized as containing **lemmas associated with an average readability level across different dialects**, plus additional information such as **frequency of occurrence** and **part-of-speech tags** for each lemma/readability pair [2509.22870]. This establishes the lexicon’s operational schema in model-facing terms: lemma identity, graded readability, frequency, and POS. The same paper treats SAMER as the system’s main source of explicit lexical difficulty knowledge.

In the annotation interface used with the SAMER corpus, two extra categories are also defined: **Level 0** for proper nouns and **Level 6** for unknown words not found in the lexicon [2404.18615]. These are not presented as core readability levels of the lexicon itself, but as operational labels for annotation and tooling.

## 3. Role in the SAMER Corpus and Simplification Pipeline

The SAMER Arabic Text Simplification Corpus is described as the **first manually annotated Arabic parallel corpus for text simplification** targeting **school-aged learners**. It comprises texts of **159K words** selected from **15 publicly available Arabic fiction novels**, most of which were published between **1865 and 1955**. The corpus includes readability level annotations at both the **document** and **word** levels, as well as **two simplified parallel versions** for each text targeting learners at two different readability levels [2404.18615].

The lexicon is the mechanism that makes those annotations and simplifications systematically leveled. Each document has three parallel versions: **Original**, **Level 4 simplified version**, and **Level 3 simplified version**. The simplification process is hierarchical: if a document is **Level 5**, it is first simplified to **Level 4**, and then the Level 4 version is simplified further to **Level 3**; if the original document is already **Level 4**, it is simplified directly to **Level 3**; if it is already **Level 3**, no simplification is needed [2404.18615].

The paper states that **all Level 5 words are simplified to Level 4 or lower** in the Level 4 version, and then **all Level 4 words are simplified to Level 3 or lower** in the Level 3 version. This makes the lexicon a hard constraint on target readability, not merely a descriptive tagging resource. The target audiences are explicitly aligned with the level scheme: **Level 4** corresponds to grades **6–8**, while **Level 3** corresponds to grades **4–5** [2404.18615].

At the document level, readability is defined as **the highest readability level found among the words in the document** [2404.18615]. This convention is lexicon-driven: document difficulty is derived compositionally from the difficulty of its lexical items.

## 4. Annotation Workflow and Human-in-the-Loop Use

Word-level annotation in the corpus is performed through a workflow that combines automatic morphological processing with manual correction. The workflow is described as follows: **tokenize and disambiguate the text using CAMeL Tools**; use a **BERT unfactored morphological disambiguator** to predict **lemma** and **POS**; look up the **lemma/POS** in the readability lexicon; and assign the word a readability level [2404.18615].

The annotation was carried out by **three professional female computational linguists**, all **native Arabic speakers**, hired through a linguistic annotation firm. They used a **Google Docs add-on** developed by Hazim et al. (2022), which provided automatic word-level readability highlighting, document-level readability summaries, explicit markup options, and synonym/substitution suggestions from Arabic WordNet [2404.18615].

The markup convention permits explicit level tagging with the prefix:

\[
\#\langle \text{level} \rangle\#
\]

For example, a word may be marked as `#3#word` to indicate **Level 3** [2404.18615]. The interface also supports **Assign** and **Assign All**, allowing annotators to modify the readability level of a specific occurrence or all occurrences of that word in the document.

The annotation guidelines are lexicon-centered. Annotators first determine the original text’s readability by inspecting the automatic word-level levels and then deriving the document level from the highest word level present. They are instructed to fix anomalies caused by missing lexicon entries or morphological tagging errors, to simplify while preserving meaning and grammatical correctness, and then to **re-run the add-on** after edits to verify that the document has reached the intended readability level [2404.18615]. This establishes a quality-control loop in which lexical leveling is repeatedly checked against the resource.

## 5. Use in Arabic Readability Prediction

In later work on Arabic document readability prediction, the SAMER lexicon is integrated as an explicit knowledge source in a **heterogeneous sentence–lemma graph**. The graph is defined as

\[
\mathcal{G} = (\mathcal{V}, \mathcal{E})
\]

with multiple node types and edge types [2509.22870]. The node set includes **Sentences**, **Lemmas**, **Classes**, and **Domains**. Sentence nodes are represented by **768-dimensional contextual embeddings** from the fine-tuned Arabic transformer model `readability-arabertv2-d3tok-CE`, then augmented with linguistic features. Lemma nodes are characterized by statistical attributes such as **average readability** and **frequency**. Classes and domains are one-hot encoded [2509.22870].

The lexicon-driven step is explicit. Lemmas are extracted from sentences using **CAMeL Tools Morphology Analyzer**, preserving POS tags and recording diacritics. Each extracted lemma is then **matched against the SAMER lexicon** to enrich it with statistical attributes such as **average readability, frequency, and POS**. The paper states that this alignment ensures SAMER contributes **directly as node features**, not merely as an external lookup resource [2509.22870].

The graph includes edges `sentence → lemma` (**HAS_LEMMA**), `lemma ↔ lemma` (**OCCUR_WITH**), `sentence → class` (**IN_CLASS**), and `sentence → domain` (**IN_DOMAIN**) [2509.22870]. In this architecture, SAMER mainly affects the **lemma nodes** and, through sentence-to-lemma links, indirectly enriches **sentence nodes**.

The model equations given in the paper are generic graph-model equations rather than SAMER-specific formulas, but they define the computational setting in which the lexicon is used:

\[
h_v^{(k)} = \sigma \left( \text{AGGREGATE}_{\text{type}} \left( \left\{ h_u^{(k-1)} : u \in \mathcal{N}_\text{type}(v) \right\} \right) \right)
\]

\[
h_v^{(k)} \leftarrow \text{LayerNorm}\left(h_v^{(k)} + h_v^{(k-1)}\right)
\]

\[
y_v = \text{MLP}(h_v^{(L)})
\]

For document-level prediction, sentence-level outputs are aggregated using **max pooling**, with the document label taken as the **most difficult predicted sentence level** [2509.22870]. This mirrors the corpus convention that document readability is governed by the hardest lexical material.

The reported comparisons are between **GNN Only** and **Late Fusion** rather than with and without SAMER. At the document level, **GNN Only** yields **QWK 75.6, Acc 40.0**, while **Late Fusion** yields **QWK 76.9, Acc 42.0**. At the sentence level, **GNN Only** yields **QWK 78.5, Acc 50.0**, while **Late Fusion** yields **QWK 78.5, Acc 41.4** [2509.22870]. The paper’s interpretation is that fusion helps at the document level, whereas the GNN-only approach remains stronger for sentence-level exact classification.

## 6. Limitations, Interpretive Boundaries, and Research Significance

The published descriptions place several limits on what can be claimed about the SAMER lexicon. The readability-prediction paper does **not** define explicit **SAMER thresholds**, **SAMER-specific class labels**, any formula for converting SAMER readings into a difficulty score, or any explicit mapping from SAMER readability levels to the paper’s **19 target readability classes** [2509.22870]. It also does **not** provide a direct ablation of SAMER versus no-SAMER, so the exact marginal gain from SAMER lexicon features alone is not isolated.

The corpus paper implies a different set of limits. The resource is restricted to **fiction novels**, mostly historical literary texts, focused on **Modern Standard Arabic**, with **dialectal Arabic not included**. Simplification is restricted to **lexical simplification**; syntax is intentionally left unchanged to avoid inconsistency; and the scale is constrained by an **annotation budget** [2404.18615]. These limitations matter because the lexicon’s operational semantics are tied to the corpus and the learner-facing readability framework in which it was deployed.

Another misconception is to assume that SAMER exhaustively models Arabic lexical difficulty. The papers support a narrower interpretation: SAMER is a **strong lexical backbone** for readability-aware Arabic NLP, but its contribution is mediated by **accurate lemmatization and lexicon coverage**, and in the graph-based setting it is used in a **coarse, feature-based way** rather than through richer thresholding or rule-based calibration [2509.22870].

Its significance lies in the combination of three properties documented across the papers: it is **lemma-based**, **graded**, and **operationalized in real annotation and modeling pipelines**. In the corpus, it supports controlled simplification and document/word-level readability annotation; in downstream modeling, it supplies explicit lexical knowledge that complements transformer embeddings and graph structure. This suggests that SAMER is best understood not as a static lexical artifact, but as a readability-oriented lexical infrastructure for Arabic pedagogical and readability-sensitive NLP [2404.18615][2509.22870].

Source: https://www.emergentmind.com/topics/samer-lexicon