---
title: 'DS-TF2-EN-RO-15K: Moral Fables Corpus'
url: https://www.emergentmind.com/topics/ds-tf2-en-ro-15k
type: topic
---

# DS-TF2-EN-RO-15K: Moral Fables Corpus

DS-TF2-EN-RO-15K is a compact English–Romanian parallel corpus of 15,000 moral fables introduced within the TinyFabulist Translation Framework (TF2) for literary machine translation in a low-resource setting. The corpus pairs English source fables from the earlier monolingual DS-TF1-EN-3M pool with Romanian translations generated by GPT-o3, and it is explicitly positioned as a “silver-standard” literary-resource dataset rather than a conventional news- or web-domain MT corpus. Within TF2 it has a dual function: it is the supervised resource used to adapt open models to English→Romanian literary translation, and it is also the reference benchmark against which later systems are evaluated [2509.07829].

## 1. Corpus identity and position within the TF2 pipeline

DS-TF2-EN-RO-15K is defined in the paper as a “15 thousand English→Romanian parallel set of moral fables.” Its language pair is English–Romanian, its genre is short didactic or moral fables, and its intended use is literary machine translation for stylistically sensitive narrative text that must preserve morals, tone, coherence, and cultural framing [2509.07829].

The dataset occupies Stage 2 of a four-stage TF2 pipeline. Stage 1 evaluates candidate translators. Stage 2 creates DS-TF2-EN-RO-15K. Stage 3 fine-tunes open models on it. Stage 4 uses the best fine-tuned model to translate the much larger English-only corpus into DS-TF2-EN-RO-3M. Its relation to the other TinyFabulist datasets is central. DS-TF1-EN-3M is the earlier English-only corpus of 3 million synthetic moral fables; it is monolingual and not a translation dataset. DS-TF2-EN-RO-15K is derived by sampling 15,000 stories from that English TF1 pool and translating them into Romanian. DS-TF2-EN-RO-3M is then the large-scale bilingual extension produced later by translating the full remaining English corpus using TF2 fine-tuned models.

This progression gives DS-TF2-EN-RO-15K a specific methodological status. It is not merely a held-out benchmark and not merely a training set. It is the reference-quality seed corpus from which the later bilingual ecosystem is bootstrapped. A plausible implication is that the scientific role of the dataset lies less in scale than in leverage: a relatively small parallel corpus is used to supervise open models that subsequently generate a much larger bilingual resource.

## 2. Construction methodology and reference generation

The construction procedure is described as a sampling-plus-translation pipeline. First, the authors benchmarked 13 candidate systems—proprietary LLMs, commercial MT APIs, and open instruction-tuned models—on literary translation quality. Because there were no human Romanian references at this stage, the benchmark did not initially use BLEU. Instead, evaluation relied only on an LLM-based five-dimension rubric applied to a randomly selected subset of 100 fables per model. The evaluation prompt instructed the judge: “You are a professional translation evaluator. You will be given an English fable and a Romanian translation. Evaluate the translation for accuracy, fluency, coherence, style, and cultural/pragmatic fidelity. Provide a score from 1 to 5 for each category along with a brief justification. Output your evaluation in valid JSON format with fields for each score and justification” [2509.07829].

On this Stage 1 benchmark, GPT-o3 (2025-04-16) achieved the highest reported rubric scores: Accuracy 4.86, Fluency 4.92, Coherence 4.89, Style 4.96, Cultural 4.97, with an average of 4.92 on 100 samples. It was therefore selected as the “reference translator,” and the paper states: “We therefore adopt GPT-o3 outputs as our silver-standard references for automatic comparisons elsewhere in the paper.”

Stage 2 then samples 15,000 short fables from DS-TF1-EN-3M and translates each into Romanian with GPT-o3. The paper states: “The pipeline begins by sampling 15,000 short fables from the DS-TF1-EN-3M dataset… Rather than relying on scarce human-translated sources, we use the top-performing LLM-based translator, selected via rigorous benchmarking, to generate Romanian counterparts for each English story.” No additional algorithmic selection procedure, ranking scheme, or human post-editing process is described beyond this sampling-plus-translation step. The paper also does not specify whether the 15,000 were sampled uniformly, stratified by prompt type, or filtered for length or style beyond being “short fables.”

Quality control is therefore primarily indirect. It derives from the translator-selection benchmark and from provenance retention rather than from corpus-wide human adjudication. The repeated designation “silver-standard” is consequently substantive rather than merely terminological: the Romanian side is model-generated, not human-translated.

## 3. Data structure, splits, and alignment granularity

DS-TF2-EN-RO-15K contains exactly 15,000 English–Romanian pairs. The split reported in Section 3.2 is 12,000 training pairs, 1,500 validation pairs, and 1,500 test pairs [2509.07829].

| Split | Pairs |
|---|---:|
| Train | 12,000 |
| Validation | 1,500 |
| Test | 1,500 |

Each example is stored as one JSON object in JSONL format. The fields listed for the release are `fable`, `translated_fable`, `pipeline_stage`, `source_lang`, `target_lang`, `prompt_hash`, `llm_name`, `translation_model`, and `generation_timestamp`. The `fable` field contains the original English text, described as “typically 1–3 paragraphs, always ending with an explicit moral.” The `translated_fable` field contains the Romanian translation generated by the selected model. The remaining fields preserve provenance, including the pipeline stage, language identifiers, a SHA-256 hash of the generation prompt, model identifiers, and a Unix timestamp for auditing.

The dataset is aligned at the story or document level rather than at the sentence level. The paper does not describe sentence-aligned pairs, sentence boundaries, discourse tags, paragraph-level annotations beyond the raw text, or moral labels inside the 15K release. This document-level organization matters because the evaluation criteria emphasize coherence, style, and moral framing—properties that depend on narrative context. The dataset therefore encodes complete short narratives rather than decontextualized sentence pairs.

The paper does not provide token totals for the 15K subset alone, but it does give indicative lengths for the fable format. In the benchmarked evaluation setting, each input sequence contains on average 300–400 tokens, with outputs averaging 350–450 tokens. Elsewhere, the cost section states that each fable contains roughly 300–600 input tokens and translations contain 300–600 output tokens. These figures are not isolated specifically for DS-TF2-EN-RO-15K, but they likely characterize the same fable format.

## 4. Evaluation protocol and quality model

The authors’ notion of dataset quality is tied to domain specificity, document-level narrative structure, and reference generation with the strongest translator found in their benchmark. Section 3.2 calls the corpus a “high-quality parallel dataset tailored to the narrative and stylistic demands of the target genre,” while also acknowledging that the Romanian side is synthetic. The dataset is explicitly designed to serve both as training supervision and as a benchmark reference set [2509.07829].

Evaluation combines corpus-level BLEU with a five-dimension rubric. BLEU is introduced only after DS-TF2-EN-RO-15K exists, because prior to that there were no Romanian references. The paper notes that BLEU is reported on a normalized 0–1 scale, so \(0.0926\) corresponds conventionally to 9.26 BLEU points. At the same time, BLEU is repeatedly described as limited for creative translation and is presented as complementary to rubric-based literary evaluation rather than as a standalone measure.

The rubric dimensions are defined as follows: Accuracy is “Fidelity to the full semantic content of the source.” Fluency is “Grammaticality, naturalness, and readability in the target language.” Coherence is “Logical flow and clarity at the story level.” Style is “Appropriateness of tone and consistency with the narrative voice.” Cultural/Pragmatic Adaptation is “Adequacy of cultural references and moral framing for the target audience.” Each dimension is scored from 1 to 5, and the aggregate rubric score is the mean of the five dimensions.

The primary evaluator in the main experiments is GPT-o3-mini. For rubric evaluation, the paper samples 100 test instances per model rather than scoring all 1,500 test items. It later performs a cross-family bias check with Grok-3-mini on the same 100-item subset and the same rubric. This design addresses, but does not eliminate, a core methodological concern: the same model family produces the silver-standard references and also supplies the main judge.

The paper’s qualitative discussion clarifies the error model that the benchmark is intended to penalize. Appendix B contrasts untuned Gemma outputs with TF2-12B outputs on difficult fables and highlights species mistranslations such as “skunk→Fumeg,” “cheetah→ceată,” and “hippo→iepuraș,” as well as typos and unidiomatic Romanian. Section 6 further notes that even strong systems may produce literal translations, unusual word choices, or minor syntactic issues, while weaker systems may drop details or mishandle Romanian morphology such as gender and tense.

## 5. Function in model adaptation and reported empirical results

In Stage 3, DS-TF2-EN-RO-15K is the supervised dataset for LoRA-based fine-tuning of Gemma-3 backbones at 1B, 4B, and 12B scales, producing TF2-1B, TF2-4B, and TF2-12B. Training pairs are converted into instruction–response format; the English input is prefixed with an instruction such as “Translate the following fable from English to Romanian:”; examples are tokenized with the model’s native tokenizer; and the loss is masked so that only Romanian output tokens contribute. The paper’s abstract describes a two-stage fine-tuning process for TF2-12B: “(i) instruction tuning to capture genre-specific narrative style, and (ii) adapter compression for efficient deployment.” In the body, this is operationalized as LoRA-based parameter-efficient fine-tuning followed by adapter merging and quantized deployment, including 8-bit quantized variants and W8A8 compressed artifacts [2509.07829].

Because DS-TF2-EN-RO-15K provides both supervision and references, it supports direct before/after adaptation comparisons. For TF2-12B, the best reported result is at temperature \(T=0.0\): BLEU 0.0926 and rubric average 4.83. The quantized TF2-12B model is close: BLEU 0.0746 and average 4.82 at \(T=0.0\), and average 4.82 at \(T=0.2\). The untuned Gemma-3-12B-it baseline is substantially lower, with BLEU 0.0214 and rubric 4.43. The paper summarizes this as: “TF2-12B improves from 4.43 (Gemma-12B-it) to 4.83 at \(T=0.0\) (and from BLEU 0.0214 to 0.0926).”

The same corpus underlies the 4B and 1B results. TF2-4B reaches rubric 4.74 at \(T=0.0\) and BLEU 0.1153, compared with 3.81 and 0.1005 for untuned Gemma-3-4B-it. TF2-1B reaches average rubric 3.75, while the distilled variant gets BLEU 0.2180 and rubric 3.73, versus 2.02 for untuned Gemma-3-1B-it. The paper notes that TF2-4B has higher BLEU than TF2-12B but lower rubric quality, and explains such discrepancies by paraphrastic diversity and BLEU’s sensitivity to wording.

The benchmark also supports direct open-versus-proprietary comparisons. GPT-o3 remains the best system with average rubric 4.92, but TF2-12B at 4.83 is within 0.09 points on the 1–5 scale. In the cross-family robustness study, GPT-o3 reference translations score 4.92 under both judges, TF2-12B-quant scores 4.82 under GPT-o3-mini and 4.90 under Grok-3-mini, and TF2-12B scores 4.82 and 4.85 respectively. The reported “gap to o3” is 0.10 with GPT-o3-mini and narrows to 0.02–0.07 with Grok-3-mini. These figures are the empirical basis for the paper’s claim that a fine-tuned small open model can reach near parity with much larger proprietary systems on this literary benchmark.

## 6. Availability, limitations, and research significance

DS-TF2-EN-RO-15K is publicly released on Hugging Face at `https://huggingface.co/datasets/klusai/tf2-en-ro-15k` and is stated to be under the MIT License. The paper also states that all data, code, prompts, and evaluation scripts are published under permissive licenses, and that the repository contains the LLM-evaluation prompts and example JSON outputs. The inclusion of prompt hashes, model identifiers, and timestamps is part of a broader emphasis on transparency, replicability, and auditability [2509.07829].

The principal limitation is explicit: “Our reference ‘ground truth’ Romanian texts come from GPT-o3, not from human translators.” As a result, the targets may contain subtle errors or unnatural phrasing, and systems that generate better but different translations may be penalized by BLEU or even by the LLM judge. The paper also acknowledges possible evaluator/reference family bias because GPT-o3 produced the references and GPT-o3-mini served as the primary judge, though the Grok-3-mini cross-check partially mitigates this concern.

A second limitation is domain narrowness. The corpus consists only of moral fables, whose “fairly straightforward language and repetitive structure” may make translation easier than other literary genres such as novels or poetry. The dataset is therefore highly valuable for its niche but not necessarily representative of literary translation as a whole. Additional caveats include the limited coverage of rubric evaluation—100 randomly selected test samples per model rather than the full 1,500-item test set—and the possibility that the corpus reflects biases inherited from synthetic source data in TF1 and from the style or preferences of GPT-o3 as translator. The paper discusses cultural adaptation and moral framing as central issues, but it does not report a dedicated ethics audit or a human cultural-review workflow for the Romanian references.

The resource nonetheless fills a specific gap: an open, document-level, literary English–Romanian parallel benchmark centered on short moral narratives, with provenance metadata and an accompanying evaluation protocol. The paper directly mentions or strongly implies four classes of experiments enabled by the corpus: supervised fine-tuning of open translation models for Romanian literary text, benchmark comparison of proprietary and open systems on a narrative domain, study of literary evaluation beyond BLEU using a five-dimension rubric, and bootstrapping large bilingual corpora such as DS-TF2-EN-RO-3M. Its broader significance lies in operationalizing a low-resource strategy in which a very strong model is used once to create a small, high-quality silver-standard corpus, which is then used to train cheaper open models that can scale.

The cost results reported for the later 3M-fable translation stage clarify why DS-TF2-EN-RO-15K is treated as the leverage point of the entire framework. Table 5 estimates total cost for translating the full 3M-fable corpus at \$13,500 for GPT-4.1, \$2,700 for GPT-4.1-mini, \$24,300 for GPT-o3, \$13,365 for GPT-o3-mini, \$270,000 for DeepL API Pro, and about \$350 for “TF2 fine-tuned (ours).” The paper’s conclusion is therefore not merely about a dataset in isolation. It is that DS-TF2-EN-RO-15K enables a reproducible path from a compact silver-standard literary corpus to open models that approach proprietary literary translation quality at a fraction of the cost.

Source: https://www.emergentmind.com/topics/ds-tf2-en-ro-15k