Papers
Topics
Authors
Recent
Search
2000 character limit reached

SAMER Lexicon: Arabic Readability Resource

Updated 13 July 2026
  • SAMER Lexicon is a lemma-based graded resource for Modern Standard Arabic, underpinning text simplification and word-level readability annotation.
  • It organizes 26K to 40K lemmas into five readability levels (L1–L5) and integrates frequency, POS tags, and average readability for precise annotation.
  • The lexicon serves as a controlled infrastructural backbone in Arabic NLP pipelines, enabling systematic lexically controlled rewriting and improved readability prediction.

Searching arXiv for the most relevant papers on the SAMER lexicon and closely related Arabic readability resources. The SAMER Lexicon is a lemma-based readability lexicon for Arabic that functions as the lexical backbone of the Simplification of Arabic Masterpieces for Extensive Reading resource ecosystem. In the literature, it is described as a large-scale leveled readability lexicon for Modern Standard Arabic, organized into graded difficulty levels and used both for word-level readability annotation and for lexically controlled simplification. In the SAMER corpus, it anchors document- and word-level annotation and constrains simplification toward target learner levels; in later modeling work, it is used as a source of explicit lexical difficulty knowledge, contributing average readability, frequency of occurrence, and part-of-speech tags for matched lemmas (Alhafni et al., 2024, Elchafei et al., 26 Sep 2025).

1. Definition and Resource Scope

SAMER is presented as part of a broader Arabic pedagogical NLP infrastructure rather than as an isolated word list. The corpus paper defines simplification primarily as lexically controlled rewriting: replacing difficult words with simpler alternatives of equivalent meaning, while keeping the text grammatically correct and semantically faithful. A key design decision is that the corpus is anchored in a readability-leveled lexicon rather than in free-form simplification (Alhafni et al., 2024).

The lexicon is described as having been initially developed by Al Khalil et al. (2020) and later extended by Jiang et al. (2020). In this formulation, SAMER is a readability resource centered on lemmas, not merely surface word forms. The corpus paper explicitly characterizes it as a controlled readability infrastructure used for automatic analysis and guided simplification, and also notes related tooling such as a Google Docs add-on by Hazim et al. (2022) for word-level readability visualization and simplification support, with Arabic WordNet-based substitution suggestions and morphological disambiguation and lemmatization tools used to map surface forms to lemmas (Alhafni et al., 2024).

A common misconception is to treat SAMER as a conventional dictionary. The published descriptions are narrower and more technical: it is a graded readability lexicon whose central function is to attach difficulty metadata to lemmas and to support downstream annotation and modeling.

2. Readability Representation and Lexical Schema

The lexicon contains 26K lemmas in its original version and was later extended to 40K lemmas. These lemmas are organized into five readability levels, from L1 to L5, where L1 is described as low difficulty / easy readability and L5 as high difficulty / hard readability (Alhafni et al., 2024).

Level Characterization Grade / age mapping
L1 basic words like house, tree, to make, but not specified in the mapping
L2 easier than upper levels grades 2–3, ages 7–8
L3 intermediate readability grades 4–5, ages 9–10
L4 advanced school readability grades 6–8, ages 11–14
L5 hardest readability grades 9+, ages 15+

Within the later readability-prediction paper, SAMER is summarized as containing lemmas associated with an average readability level across different dialects, plus additional information such as frequency of occurrence and part-of-speech tags for each lemma/readability pair (Elchafei et al., 26 Sep 2025). This establishes the lexicon’s operational schema in model-facing terms: lemma identity, graded readability, frequency, and POS. The same paper treats SAMER as the system’s main source of explicit lexical difficulty knowledge.

In the annotation interface used with the SAMER corpus, two extra categories are also defined: Level 0 for proper nouns and Level 6 for unknown words not found in the lexicon (Alhafni et al., 2024). These are not presented as core readability levels of the lexicon itself, but as operational labels for annotation and tooling.

3. Role in the SAMER Corpus and Simplification Pipeline

The SAMER Arabic Text Simplification Corpus is described as the first manually annotated Arabic parallel corpus for text simplification targeting school-aged learners. It comprises texts of 159K words selected from 15 publicly available Arabic fiction novels, most of which were published between 1865 and 1955. The corpus includes readability level annotations at both the document and word levels, as well as two simplified parallel versions for each text targeting learners at two different readability levels (Alhafni et al., 2024).

The lexicon is the mechanism that makes those annotations and simplifications systematically leveled. Each document has three parallel versions: Original, Level 4 simplified version, and Level 3 simplified version. The simplification process is hierarchical: if a document is Level 5, it is first simplified to Level 4, and then the Level 4 version is simplified further to Level 3; if the original document is already Level 4, it is simplified directly to Level 3; if it is already Level 3, no simplification is needed (Alhafni et al., 2024).

The paper states that all Level 5 words are simplified to Level 4 or lower in the Level 4 version, and then all Level 4 words are simplified to Level 3 or lower in the Level 3 version. This makes the lexicon a hard constraint on target readability, not merely a descriptive tagging resource. The target audiences are explicitly aligned with the level scheme: Level 4 corresponds to grades 6–8, while Level 3 corresponds to grades 4–5 (Alhafni et al., 2024).

At the document level, readability is defined as the highest readability level found among the words in the document (Alhafni et al., 2024). This convention is lexicon-driven: document difficulty is derived compositionally from the difficulty of its lexical items.

4. Annotation Workflow and Human-in-the-Loop Use

Word-level annotation in the corpus is performed through a workflow that combines automatic morphological processing with manual correction. The workflow is described as follows: tokenize and disambiguate the text using CAMeL Tools; use a BERT unfactored morphological disambiguator to predict lemma and POS; look up the lemma/POS in the readability lexicon; and assign the word a readability level (Alhafni et al., 2024).

The annotation was carried out by three professional female computational linguists, all native Arabic speakers, hired through a linguistic annotation firm. They used a Google Docs add-on developed by Hazim et al. (2022), which provided automatic word-level readability highlighting, document-level readability summaries, explicit markup options, and synonym/substitution suggestions from Arabic WordNet (Alhafni et al., 2024).

The markup convention permits explicit level tagging with the prefix:

#⟨level⟩#\#\langle \text{level} \rangle\#

For example, a word may be marked as #3#word to indicate Level 3 (Alhafni et al., 2024). The interface also supports Assign and Assign All, allowing annotators to modify the readability level of a specific occurrence or all occurrences of that word in the document.

The annotation guidelines are lexicon-centered. Annotators first determine the original text’s readability by inspecting the automatic word-level levels and then deriving the document level from the highest word level present. They are instructed to fix anomalies caused by missing lexicon entries or morphological tagging errors, to simplify while preserving meaning and grammatical correctness, and then to re-run the add-on after edits to verify that the document has reached the intended readability level (Alhafni et al., 2024). This establishes a quality-control loop in which lexical leveling is repeatedly checked against the resource.

5. Use in Arabic Readability Prediction

In later work on Arabic document readability prediction, the SAMER lexicon is integrated as an explicit knowledge source in a heterogeneous sentence–lemma graph. The graph is defined as

G=(V,E)\mathcal{G} = (\mathcal{V}, \mathcal{E})

with multiple node types and edge types (Elchafei et al., 26 Sep 2025). The node set includes Sentences, Lemmas, Classes, and Domains. Sentence nodes are represented by 768-dimensional contextual embeddings from the fine-tuned Arabic transformer model readability-arabertv2-d3tok-CE, then augmented with linguistic features. Lemma nodes are characterized by statistical attributes such as average readability and frequency. Classes and domains are one-hot encoded (Elchafei et al., 26 Sep 2025).

The lexicon-driven step is explicit. Lemmas are extracted from sentences using CAMeL Tools Morphology Analyzer, preserving POS tags and recording diacritics. Each extracted lemma is then matched against the SAMER lexicon to enrich it with statistical attributes such as average readability, frequency, and POS. The paper states that this alignment ensures SAMER contributes directly as node features, not merely as an external lookup resource (Elchafei et al., 26 Sep 2025).

The graph includes edges sentence → lemma (HAS_LEMMA), lemma ↔ lemma (OCCUR_WITH), sentence → class (IN_CLASS), and sentence → domain (IN_DOMAIN) (Elchafei et al., 26 Sep 2025). In this architecture, SAMER mainly affects the lemma nodes and, through sentence-to-lemma links, indirectly enriches sentence nodes.

The model equations given in the paper are generic graph-model equations rather than SAMER-specific formulas, but they define the computational setting in which the lexicon is used:

hv(k)=σ(AGGREGATEtype({hu(k−1):u∈Ntype(v)}))h_v^{(k)} = \sigma \left( \text{AGGREGATE}_{\text{type}} \left( \left\{ h_u^{(k-1)} : u \in \mathcal{N}_\text{type}(v) \right\} \right) \right)

hv(k)←LayerNorm(hv(k)+hv(k−1))h_v^{(k)} \leftarrow \text{LayerNorm}\left(h_v^{(k)} + h_v^{(k-1)}\right)

yv=MLP(hv(L))y_v = \text{MLP}(h_v^{(L)})

For document-level prediction, sentence-level outputs are aggregated using max pooling, with the document label taken as the most difficult predicted sentence level (Elchafei et al., 26 Sep 2025). This mirrors the corpus convention that document readability is governed by the hardest lexical material.

The reported comparisons are between GNN Only and Late Fusion rather than with and without SAMER. At the document level, GNN Only yields QWK 75.6, Acc 40.0, while Late Fusion yields QWK 76.9, Acc 42.0. At the sentence level, GNN Only yields QWK 78.5, Acc 50.0, while Late Fusion yields QWK 78.5, Acc 41.4 (Elchafei et al., 26 Sep 2025). The paper’s interpretation is that fusion helps at the document level, whereas the GNN-only approach remains stronger for sentence-level exact classification.

6. Limitations, Interpretive Boundaries, and Research Significance

The published descriptions place several limits on what can be claimed about the SAMER lexicon. The readability-prediction paper does not define explicit SAMER thresholds, SAMER-specific class labels, any formula for converting SAMER readings into a difficulty score, or any explicit mapping from SAMER readability levels to the paper’s 19 target readability classes (Elchafei et al., 26 Sep 2025). It also does not provide a direct ablation of SAMER versus no-SAMER, so the exact marginal gain from SAMER lexicon features alone is not isolated.

The corpus paper implies a different set of limits. The resource is restricted to fiction novels, mostly historical literary texts, focused on Modern Standard Arabic, with dialectal Arabic not included. Simplification is restricted to lexical simplification; syntax is intentionally left unchanged to avoid inconsistency; and the scale is constrained by an annotation budget (Alhafni et al., 2024). These limitations matter because the lexicon’s operational semantics are tied to the corpus and the learner-facing readability framework in which it was deployed.

Another misconception is to assume that SAMER exhaustively models Arabic lexical difficulty. The papers support a narrower interpretation: SAMER is a strong lexical backbone for readability-aware Arabic NLP, but its contribution is mediated by accurate lemmatization and lexicon coverage, and in the graph-based setting it is used in a coarse, feature-based way rather than through richer thresholding or rule-based calibration (Elchafei et al., 26 Sep 2025).

Its significance lies in the combination of three properties documented across the papers: it is lemma-based, graded, and operationalized in real annotation and modeling pipelines. In the corpus, it supports controlled simplification and document/word-level readability annotation; in downstream modeling, it supplies explicit lexical knowledge that complements transformer embeddings and graph structure. This suggests that SAMER is best understood not as a static lexical artifact, but as a readability-oriented lexical infrastructure for Arabic pedagogical and readability-sensitive NLP (Alhafni et al., 2024, Elchafei et al., 26 Sep 2025).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (2)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to SAMER Lexicon.