Papers
Topics
Authors
Recent
Search
2000 character limit reached

CLEARS Challenge: Spanish Text Adaptation

Updated 11 July 2026
  • CLEARS is a shared task that automatically adapts municipal Spanish texts into accessible Plain Language and Easy-to-Read formats.
  • It emphasizes factual preservation, controlled language rules, and readability metrics to ensure public information remains clear and reliable.
  • The challenge employs methods from prompt engineering to iterative post-editing, highlighting the need for human evaluation alongside automated metrics.

CLEARS is a shared task at IberLEF 2025 devoted to automatic Spanish text adaptation. Its purpose is to convert original public-facing Spanish documents into accessibility-oriented rewritings in either Plain Language or Easy-to-Read form, using real municipal news rather than artificial examples. In the participating-system literature, the challenge is framed not simply as stylistic simplification, but as a mechanism for widening access to public information for readers with diverse literacy, cognitive, and linguistic needs (Ayesh et al., 5 Aug 2025, Calleja et al., 15 Sep 2025).

1. Definition, scope, and accessibility rationale

The challenge is described as the “Challenge for Plain Language and Easy-to-Read Adaptation for Spanish texts” and is organized within IberLEF 2025 (Calleja et al., 15 Sep 2025). Its central object is automatic Spanish text adaptation: a system receives an original Spanish document and must produce a rewritten version that is more accessible while preserving the underlying information content.

A defining feature of CLEARS is its practical orientation. The source material is described as municipal news from Alicante, with topics including sports, culture, leisure, festivities, and local events, so the task is grounded in ordinary administrative and civic communication rather than textbook-style simplification benchmarks (Ayesh et al., 5 Aug 2025). This framing places accessibility at the center of the evaluation problem. A recurrent theme in participant reports is that the task concerns whether people can understand and use public information, not whether a system merely produces shorter or more fluent paraphrases.

This emphasis has a normative dimension. One participating paper explicitly characterizes the task as being about “access to information as a human right,” especially for texts involving municipal announcements, cultural events, and public services (Ayesh et al., 5 Aug 2025). A plausible implication is that CLEARS sits at the intersection of text simplification, controlled language, public-sector communication, and accessibility engineering.

2. Target varieties: Plain Language and Easy-to-Read

CLEARS defines two adaptation targets. The first is Plain Language (PL), described as clear, concise, general-audience Spanish aligned in one system description with ISO 24495-1:2023 principles. The second is Easy Read / Easy-to-Read (ER, E2R, Lectura Fácil), a stricter accessibility format aligned in participant reports with Inclusion Europe guidelines and Spanish standards such as UNE 153101 EX or NE 153101:2018 EX (Ayesh et al., 5 Aug 2025, Calleja et al., 15 Sep 2025).

Plain Language is intended for a broad audience, including people with limited literacy, non-native speakers, and readers who struggle with dense bureaucratic or technical prose. The target properties repeatedly cited are clear and concise phrasing, active voice, everyday vocabulary, reduced jargon, and a coherent text structure that does not fragment excessively (Ayesh et al., 5 Aug 2025). In system prompts, PL is also associated with simple words, short sentences, logical structure, and short, well-separated paragraphs (Calleja et al., 15 Sep 2025).

Easy-to-Read is presented as a more constrained format intended explicitly for people with cognitive, intellectual, or learning disabilities, including dyslexia, ADHD, and aphasia. The defining traits are very short and simple sentences, reduced lexical and syntactic complexity, strong support for clarity and redundancy, explicit explanation of difficult notions, and often visually structured text with headings or line-by-line organization (Ayesh et al., 5 Aug 2025). One participant system operationalized ER with especially strict prompt constraints, including sentences shorter than 15 words, avoidance of passive voice and relative clauses, and a final “Palabras difíciles” section explaining complex terms (Calleja et al., 15 Sep 2025).

The distinction between PL and ER is therefore not merely one of degree. PL aims at broad comprehensibility for the general public, whereas ER is a highly regulated accessibility variety with stronger cognitive-load constraints and stronger expectations about segmentation, repetition, and explicitness.

3. Linguistic constraints and controlled-language rules

A notable feature of CLEARS is the extent to which participant systems embed explicit controlled-language rules into their generation pipelines. CardiffNLP reports that its prompts incorporated guidelines extracted from Spanish and European easy-to-read standards, including sentence-level and lexical prescriptions such as separating linked ideas with a full stop rather than a comma, avoiding semicolons, preferring simple sentences, avoiding passive voice and impersonal constructions, using affirmative sentences, keeping one main idea per sentence, and preserving dates, years, and numbers exactly (Ayesh et al., 5 Aug 2025).

The same report emphasizes lexical consistency and factual preservation. The guideline “Do not paraphrase numbers” is illustrated with the example that “2000 personas” must not become “muchas personas,” and dates and years must be preserved exactly (Ayesh et al., 5 Aug 2025). This is crucial because the challenge is not licensed to trade away factual specificity for readability. Simplification that alters dates, quantities, or proper names is treated as an error rather than a stylistic variant.

Vicomtech’s prompt design shows how these norms can be operationalized in a document-level post-editing framework. Its PL and ER prompts instruct the model to minimize words per sentence and syllables per word, separate ideas with periods, identify agents through subject–verb–predicate structure, avoid technical vocabulary unless explained, prefer repetition over variation for clarity, and, in the ER setting, avoid complex punctuation and explain difficult words in a dedicated section (Calleja et al., 15 Sep 2025). This suggests that CLEARS is also a testbed for rule-grounded prompting, not only for end-to-end neural rewriting.

A common misconception is that accessibility rewriting can be evaluated purely as generic simplification. The task descriptions contradict this. CLEARS requires concurrent attention to readability, meaning preservation, terminology control, discourse organization, and audience-specific accessibility norms.

4. Corpus design and evaluation framework

One participant paper describes the shared resource as the CLEARS corpus, consisting of 3,000 municipal news articles from Alicante, each with an original version, a human-created PL version, and a human-created E2R version, split into 2,100 training articles and 900 test articles (Ayesh et al., 5 Aug 2025). Another participant paper refers to the challenge resource as the ClearSim corpus, distributed as original Spanish documents together with human PL and ER adaptations, and reports additional filtered train/dev counts for its own experimental setup after enforcing a context-length constraint (Calleja et al., 15 Sep 2025). The naming difference reflects participant descriptions rather than a change in task definition.

At the task level, CLEARS is defined as document-level adaptation: systems are given original Spanish documents and must output PL or ER documents, depending on the subtask (Calleja et al., 15 Sep 2025). However, participant methods differ in how they operationalize this requirement. CardiffNLP treats the task as sentence-level rewriting by segmenting each article into sentences, simplifying them, and concatenating the results (Ayesh et al., 5 Aug 2025). This indicates that document-level evaluation does not preclude sentence-wise generation strategies.

The official evaluation is described as multi-metric. Vicomtech reports three official metrics: the Fernández Huerta Readability Index, bag-of-words or TF-IDF cosine similarity, and embedding cosine similarity. It also reports the combined similarity measure

CosineAvg=12(simTF-IDF+simEmb),\text{CosineAvg}=\frac{1}{2}\left(\text{sim}_{\text{TF-IDF}}+\text{sim}_{\text{Emb}}\right),

and the final ranking score as an average over the scaled official metrics:

Scoreavg=1N∑i=1Nmi,\text{Score}_{\text{avg}}=\frac{1}{N}\sum_{i=1}^{N} m_i,

where the metric set consists of TF-IDF cosine multiplied by 100, embedding cosine multiplied by 100, and Fernández Huerta readability (Calleja et al., 15 Sep 2025).

CardiffNLP additionally reports internal analyses using SentenceBERT cosine similarity, BERTScore F1, and an organizer-provided cosine-based script, and notes that the organizers average TF-IDF and BERT-based similarity to obtain an official similarity score (Ayesh et al., 5 Aug 2025). The same paper stresses that ER/E2R includes properties that are difficult to capture automatically, especially segmentation and visual formatting, so current metrics are at best partial proxies for accessibility (Ayesh et al., 5 Aug 2025). This limitation is central to understanding the challenge: a system can optimize readability indices and similarity scores without necessarily maximizing actual user comprehension.

5. Methodological landscape in CLEARS-2025

CardiffNLP’s submission is an example of a purely prompt-based approach with no fine-tuning. The team initially experimented with LLaMA-3.2 and later adopted Gemma-3 for the final submission, using zero-shot, one-shot, and few-shot prompts in English and Spanish. The final configuration used Spanish few-shot prompts, sentence-level instructions, explicit preservation constraints for names, dates, and numbers, E2R/PL guidelines embedded directly in the prompt, and structured Python-dictionary-style outputs to reduce extraneous text and improve automatic parsing (Ayesh et al., 5 Aug 2025). The paper also documents recurrent failure modes during prompt engineering, including hallucinated dates, number paraphrases, copying of demonstration outputs, format non-compliance, and occasional language switching.

Vicomtech presents a more heterogeneous architecture. Its initial adaptation stage uses LLaMA 3.1 8B Instruct with zero-shot prompting, few-shot retrieval-augmented generation via BM25 or embedding similarity, supervised fine-tuning with QLoRA, and Direct Preference Optimization. These initial outputs are then refined with Automatic Post-Editing Cycles (APEC), an iterative procedure in which the model first analyzes the adaptation against PL or ER guidelines, then generates a correction, and the system keeps revisions only when readability and embedding-based meaning-preservation scores improve (Calleja et al., 15 Sep 2025). On test data, Vicomtech ensembles APEC outputs initialized from BM25-based and DPO-based generations and selects the final version with the best internal metric average.

These two systems illustrate different methodological interpretations of CLEARS. CardiffNLP emphasizes prompt control, rule injection, and output-format discipline. Vicomtech combines retrieval, fine-tuning, preference optimization, and iterative self-revision. Together they show that the challenge accommodates both prompt-only and training-based paradigms, provided that the resulting adaptations satisfy accessibility and meaning-preservation constraints.

6. Results, limitations, and research significance

The reported results show that strong performance can be obtained without a single dominant methodology. CardiffNLP reports third place in Subtask 1 and second place in Subtask 2 after moving from LLaMA-3.2 to Gemma-3 and refining few-shot Spanish prompts (Ayesh et al., 5 Aug 2025). Vicomtech reports that, when ranking systems by the average of all official metrics, its submissions achieved first place in Plain Language with a score of 79.49 and second place in Easy Read with 75.72, with especially strong Fernández Huerta readability values of 82.98 for PL and 85.44 for ER (Calleja et al., 15 Sep 2025).

These outcomes support two more general observations. First, prompt engineering and iterative post-editing are both competitive strategies for Spanish accessibility-oriented rewriting. Second, optimizing readability alone is insufficient: the most successful systems explicitly balance readability against semantic similarity and factual preservation. CardiffNLP’s error analysis highlights hallucinations, insufficient simplification, and over-compression; Vicomtech notes that improved readability can coincide with lower lexical similarity to references because legitimate paraphrases move away from reference wording (Ayesh et al., 5 Aug 2025, Calleja et al., 15 Sep 2025).

The main unresolved issue is evaluation adequacy. Both papers argue, in different terms, that automatic metrics are incomplete proxies for accessibility. ER/E2R depends on layout, segmentation, redundancy, and audience-specific comprehensibility, all of which are only partially captured by readability formulas and embedding similarity. Both works therefore point toward human evaluation, especially with target users, as a necessary complement to automatic scoring (Ayesh et al., 5 Aug 2025, Calleja et al., 15 Sep 2025).

CLEARS is consequently significant not only as a shared task in Spanish text simplification, but as a benchmark for accessibility-aware generation under normative constraints. It links controlled language, LLM prompting, retrieval-augmented generation, preference optimization, and automatic post-editing to a concrete public-information use case. A plausible implication is that future progress in this area will depend less on generic paraphrasing quality than on better accessibility metrics, stricter factuality control, and evaluation protocols that measure whether adapted texts are genuinely understandable and usable for the populations they are intended to serve.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to CLEARS Challenge.