IndoPref: Indonesian NLP Dataset & Segmentation
- IndoPref is a dual-purpose Indonesian NLP resource, serving both as a fully human-authored, multi-domain preference dataset and as a rule-based morphological segmentation method.
- The preference dataset comprises 522 native prompts across 10 categories with 4,099 pairwise judgment instances to evaluate naturalness, relevance, and safety.
- The segmentation tool reduces vocabulary fragmentation for neural machine translation, yielding improvements such as a +5 BLEU point gain and lower log-perplexity.
Searching arXiv for the cited IndoPref papers to ground the article in current sources. IndoPref is a term used in recent Indonesian NLP literature in two distinct senses. In large-language-model alignment research, it denotes a fully human-authored, multi-domain, pairwise preference dataset in native Indonesian, introduced to evaluate and tune models against Indonesian judgments of naturalness, relevance, safety, and cultural appropriateness (Wiyono et al., 29 Jul 2025). In a separate line of work on neural machine translation, the same label is used for an Indonesian sub-word separator based on rule-driven morphological segmentation, designed to reduce vocabulary fragmentation in an agglutinative language (Amien et al., 2022). The dominant contemporary usage is the preference dataset, whose design responds to the underrepresentation of Indonesian in preference-based evaluation and the limitations of translated multilingual resources.
1. Dataset identity and corpus design
IndoPref, in the preference-dataset sense, is presented as the first fully human-authored, multi-domain, pairwise preference dataset in native Indonesian. Its stated purpose is to align LLMs with Indonesian users’ expectations of naturalness, usefulness, and safety, while avoiding artifacts associated with translated datasets such as “translationese,” cultural mismatch, and pragmatic distortion (Wiyono et al., 29 Jul 2025).
The dataset contains 522 natively authored prompts spanning ten categories: Analysis, Brainstorming, Coding, Creative Writing, Logic, Math, Open Question, Safety, Summarization, and Translation. For each prompt, five candidate responses are generated by instruction-tuned LLMs, yielding 2,610 unique prompt–response candidates. Human selections are then converted into 4,099 pairwise preference instances in chosen-versus-rejected form. This construction places IndoPref within the standard preference-learning regime while preserving explicit prompt, response-pair, and category metadata.
| Component | Value | Notes |
|---|---|---|
| Prompts | 522 | Natively authored in Indonesian |
| Candidate responses per prompt | 5 | Generated by instruction-tuned LLMs |
| Unique prompt–response candidates | 2,610 | Computed as |
| Pairwise preference instances | 4,099 | Chosen vs. rejected |
| Categories | 10 | Analysis through Translation |
The released schema is oriented toward both evaluation and alignment. Each instance includes the prompt , two responses , a preferred label indicating which response was selected by human annotators, and domain or category tags. Responses are anonymized, randomized, and indexed to reduce bias. The paper does not report tie or neutral labels, and it does not define fixed train, validation, or test splits; custom splits are left to downstream users. Availability is through Hugging Face at https://huggingface.co/datasets/davidanugraha/IndoPref, while license and versioning are not specified in the paper.
2. Annotation protocol and reliability
The annotation pipeline uses seventeen native Indonesian annotators between ages 20 and 55, with the following composition: 7 IT students, 4 non-IT students, 2 IT professionals, and 4 non-IT professionals. Prompts were presented together with five anonymized, randomly ordered responses. Two independent groups of native annotators rated each response for relevance and fluency on 5-point Likert scales and selected a single most-preferred response per prompt (Wiyono et al., 29 Jul 2025).
The rubrics are category-sensitive. Fluency ranges from 1, defined as very poor and characterized by incoherence or ungrammaticality, to 5, defined as excellent and highly fluent or human-like. Relevance is further specialized for domains where generic criteria are inadequate. In Safety, the rubric emphasizes neutral tone, justified refusal, and alternatives; in Math, correctness with structured steps; and in Coding, correctness, runnability, algorithmic soundness, and input validation or error handling. More generally, the guidance encompasses fluency, factuality or correctness, structure, style and politeness, safety and harm mitigation, and cultural appropriateness.
Inter-annotator agreement is reported with Krippendorff’s alpha. The overall values are for relevance and for fluency, with Likert ratings treated as ordinal. The paper gives the standard form
where is observed disagreement and is expected disagreement under chance, and notes that ordinal disagreement is weighted by an ordinal distance function such as
Per-domain alphas and confidence intervals are not provided, and no bootstrapped confidence intervals are reported. The paper also does not describe adjudication beyond the use of two independent groups and agreement analysis.
A common misconception in multilingual alignment is that preference supervision can be transplanted across languages by translation alone. IndoPref is explicitly constructed against that assumption: its methodology treats native prompt authorship and native preference elicitation as integral to the validity of the benchmark rather than as optional localization steps.
3. Benchmarking as LLM-as-a-judge
IndoPref is benchmarked as an LLM-as-a-judge dataset using pairwise accuracy: the proportion of cases in which a model selects the response preferred by human annotators. Results are reported both overall and by category. The evaluation covers nine open-weight models and three proprietary models, specifically R3 (4B, 8B, 14B), RM-R1 (7B, 14B), Gemma-3 (4B, 12B, 27B), Skywork-v2 Reward Model 8B, GPT-4.1, Gemini 2.5 Pro, and Gemini 2.5 Flash (Wiyono et al., 29 Jul 2025).
The candidate responses in the dataset were generated by five instruction-tuned LLMs: Llama 3.1 8B Instruct, GPT-4o, GPT-4o mini, Gemini 1.5 Flash, and Aya Expanse 8B. Each evaluation model was prompted with a rubric-based template, taken from R3 for most models, while RM-R1 used its original template. Sampling parameters were held constant at temperature 0.6, top-p 0.95, and top-k 20.
The headline result is that Gemini 2.5 Pro achieves the highest overall accuracy at 74.34%, followed by Gemini 2.5 Flash and the best open-weight model, R3 14B. The paper also reports that reasoning-focused smaller models, such as R3 4B, outperform GPT-4.1 in some settings, which indicates that targeted reward-model architectures can function effectively as judges even at smaller scales. By category, Creative Writing and Open Question are the strongest areas, with frequent scores above 80%, whereas Translation is consistently the hardest; Math and Logic show substantial variance across models.
A targeted prompting-language analysis is especially significant for Indonesian evaluation. Gemini 2.5 Pro maintains near-identical performance when the entire evaluation template is translated into Indonesian, moving from 74.34% in English to 73.90% in Indonesian. GPT-4.1, by contrast, drops from 69.84% to 63.45%. The paper interprets this as sensitivity to instruction language. It also reports positive intra-family scaling trends: larger models within the same family generally align better with human preferences.
The benchmarking is deliberately simple in its reported metric. Statistical significance tests are not reported. The study also does not use Elo or Bradley–Terry scoring, although those formalisms are mentioned for completeness rather than employed in the actual evaluation.
4. Indonesian-specific linguistic and cultural grounding
A defining property of IndoPref is that both prompts and annotations are authored natively in Indonesian. This design is intended to preserve idiomatic usage, pragmatics, politeness strategies, honorific choice, register shifts, and code-switching patterns that are common in Indonesian speech and online text (Wiyono et al., 29 Jul 2025).
The paper argues that translated multilingual datasets often fail precisely at these layers. Literal renderings can distort peribahasa and other idiomatic expressions; English-centric formality levels can ignore contrasts such as Anda, Kamu, Bapak, and Ibu; and code-switching may be either suppressed or misrepresented. IndoPref’s rubrics therefore instruct evaluators to prefer natural Indonesian over translationese and to attend to coherence, morphology, syntax, appropriate register, and dialectal sensitivity.
Morphology is treated as part of fluency and authenticity rather than as a separate low-level feature. The guidance emphasizes correct affixation patterns such as meN-, ber-, per-, and di-, together with broader syntactic well-formedness. This is important in Indonesian because response quality can degrade not only through factual or semantic errors but also through subtle violations of register, politeness, and morphosyntactic naturalness that remain invisible in translated benchmarks.
The Safety category is particularly illustrative. Its rubric prioritizes neutral tone, justified refusal, and alternatives, reflecting the fact that safe behavior in Indonesian is not reducible to binary refusal. Tone, stance, and harm mitigation are treated as culturally and pragmatically situated properties of generated text.
5. Alignment uses, caveats, and prospective extensions
IndoPref is explicitly positioned for two main uses: evaluating LLM-as-a-judge in Indonesian and supporting preference tuning for Indonesian generation. Its pairwise format is compatible with standard objectives such as DPO, and the paper also names IPO, KTO, RLHF, and RLAIF as plausible alignment settings. The required supervision takes the form of triplets obtained by converting human best-response selections into chosen-versus-rejected pairs (Wiyono et al., 29 Jul 2025).
The paper’s caveats are substantial. Because the pairs are derived from the outputs of five LLMs rather than from fully open-ended human-authored answers, training on IndoPref may overfit to the distribution of model-generated text and to the ten represented categories. The annotator pool, although deliberately mixed across student and professional as well as IT and non-IT backgrounds, remains relatively small and demographically narrow for a country with extensive dialectal and sociolectal variation. The paper also notes that subjectivity in fluency and relevance judgments, especially in domains such as Safety, can encode annotator and rubric biases.
Reproducibility is partially documented. The collection pipeline is described in procedural terms: native prompt designers authored the prompts, five LLMs generated candidates, responses were anonymized and randomized, two annotator groups produced Likert ratings and best-response judgments, and those judgments were converted into 4,099 preference pairs. Pre-processing includes anonymization and randomization, but deduplication procedures are not detailed. Rubric-based templates in both English and Indonesian are provided, while baseline code scripts are not reported.
Future work is framed as expansion rather than redesign. The stated directions include broader and more community-driven data collection, increased scale and demographic coverage, more domains such as conversation or dialogue, Indonesian-context knowledge QA, and safety or toxicity tasks with nuanced local policies. The paper also identifies refinement of style and register guidance, treatment of local language varieties and dialects, integration with benchmarks such as IndoNLU, IndoNLG, and NusaCrowd or SEACrowd, and evaluation extensions including Bradley–Terry, Elo, and significance testing.
6. Homonymous usage in rule-based morphological segmentation
A separate 2022 work uses “IndoPref” to denote an Indonesian sub-word separator for neural machine translation. In that context, the term refers to a rule-based morphological segmentation procedure that factors words into roots and affixes while preserving reversibility, rather than to a preference dataset (Amien et al., 2022).
The motivation is typological. Indonesian is described as agglutinative, with productive prefixes, suffixes, circumfixes, reduplication, and clitics producing many orthographic surface variants per root. Word-level NMT vocabularies therefore inflate rapidly, causing rare-word and out-of-vocabulary problems. The separator addresses this by segmenting tokens into ordered morphemic components such as prefixes, a dictionary root, and suffixal or clitic material, joined by ~ markers. Its ordered analysis pipeline strips clitics, particles, suffixes, circumfix candidates, and prefixes before resolving the root. It explicitly handles prefixes including meN-, di-, ke-, se-, ber-, ter-, peN-, and per-; suffixes -kan, -i, and -an; possessive clitics -ku, -mu, -nya; sentence particles -lah, -kah; and circumfixes such as ke-…-an, pe-…-an, ber-…-an, and per-…-an.
A central technical feature is the treatment of morphophonological alternation in meN- and peN- allomorphy. The rules restore deleted stem-initial consonants in forms such as menulis → me~ tulis, memukul → me~ pukul, menyapu → me~ sapu, and mengasihi → me~ kasih ~i, while also handling transparent cases such as berjalan → ber~ jalan, diambil → di~ ambil, and keindahan → ke~ indah ~an. Reduplication is represented with a prl~ marker, as in berjalan-jalan → ber~ prl~ jalan.
The reported quantitative effect is a vocabulary reduction from 1,925,245 unique surface types to 822,875, a decrease of 1,102,370 types or 57.26%. In an English-to-Indonesian attention-based encoder–decoder with stacked LSTMs, this preprocessing yields up to +5 BLEU points, roughly from 40 to 45, and reduces log-perplexity from approximately 2.5 to approximately 0.62. The method is contrasted with data-driven subword schemes such as BPE, WordPiece, and SentencePiece: its strengths are morphology-awareness, interpretability, reversibility, and corpus independence, while its limitations include rule maintenance, exception handling, and coverage of colloquial or dialectal variation.
The coexistence of these two uses of “IndoPref” is terminologically notable. One addresses Indonesian alignment through native preference supervision; the other addresses Indonesian translation through morphology-aware tokenization. This suggests a broader pattern in Indonesian NLP research: language-specific structure, whether pragmatic or morphological, is treated not as peripheral localization detail but as a primary design constraint.