---
title: 'BrahmicTokenizer-131K: Indic Tokenizer for 9 Scripts'
url: https://www.emergentmind.com/topics/brahmictokenizer-131k
type: topic
---

# BrahmicTokenizer-131K: Indic Tokenizer for 9 Scripts

Searching arXiv for the specified paper and closely related tokenizer work.
arxiv_search(query="BrahmicTokenizer-131K Indic-capable drop-in replacement o200k_base", max_results=5)
Searching for related Indic tokenizer papers for contextualization.
arxiv_search(query="Indic tokenizer Sarvam MUTANT-Indic byte-level BPE Brahmic", max_results=10)
BrahmicTokenizer-131K is a **131,072-vocabulary byte-level BPE tokenizer** designed as an **“Indic-capable drop-in replacement for o200k_base”** while preserving the properties that made **o200k_base** attractive for English, European languages, code, and math [2605.29379]. Its central purpose is to close what the paper terms the **“Brahmic compression gap”**: the systematic under-allocation of tokenizer capacity to Brahmic scripts in mainstream multilingual tokenizers at comparable vocabulary sizes, which can force Indic text into inefficient byte-fragment tokenization. The design is not a new tokenization algorithm; it is a retrofit of **o200k_base** that keeps the same tokenizer-side interface and operational machinery, but changes the vocabulary contents and appends Brahmic merge entries so that the tokenizer becomes usable for nine Brahmic scripts covering 11 major Indian languages without giving up inherited non-Brahmic behavior.

## 1. Definition, compatibility, and intended scope

In the paper’s narrow technical sense, **“drop-in replacement”** does not mean vocabulary identity with **o200k_base**. It means that the tokenizer preserves the same **byte-level BPE algorithm**, the same **GPT-2 ByteLevel pre-tokenizer**, the same **decoder**, the same **byte handling and byte fallback conventions**, the same **special-token format**, and the same **tokenizer JSON schema** [2605.29379]. A system that already uses **o200k_base** can therefore replace the tokenizer artifact with BrahmicTokenizer-131K’s `tokenizer.json` without changing data-loader code, pre-tokenizer code, or decoder code.

What changes is the vocabulary and some added Brahmic merge-list entries. Consequently, the model’s embedding matrix and LM head require the standard vocabulary resize from **200,019 rows to 131,072 rows**. The paper stresses that this is the same kind of mechanical resize downstream training code already supports when switching tokenizers.

The target deployment profile is explicitly limited. The tokenizer is optimized for:

- English
- three major EU languages: French, German, Spanish
- code
- math
- nine Brahmic scripts covering 11 major Indian languages

It is **not** optimized for scripts intentionally removed during the crop stage: **CJK Unified Ideographs, Hangul, Hiragana + Katakana, Arabic, Cyrillic, Greek, Thai, Hebrew, and Sinhala**. This makes it a replacement for **o200k_base** at the interface level, but not a universal replacement for every multilingual setting [2605.29379].

A common point of confusion is the label **“131K.”** In this tokenizer, it denotes a **131,072-token vocabulary**, not a long-context window. This distinction matters because the paper evaluates tokenization efficiency, not context-extension algorithms.

## 2. The Brahmic compression gap

The motivating claim is that multilingual tokenizers trained by generic frequency-based byte-level BPE on multilingual corpora can underperform on Brahmic scripts because **frequency alone underproduces tokens for underrepresented scripts** [2605.29379]. In the GPT-2 ByteLevel setup, Unicode text is first mapped through a printable-byte bijection. A Brahmic character such as a Devanagari letter becomes a multi-byte UTF-8 sequence, and if learned merges do not compose those bytes into script-specific tokens, tokenization falls back toward byte fragments.

The paper describes three regimes of Brahmic tokenization:

- **Byte-fallback regime**: no meaningful script-level merges; Brahmic characters can cost 3+ tokens each.
- **Subword regime**: characters and subword pieces are represented, leading to a few tokens per word.
- **Whole-word regime**: common words themselves become single tokens, pushing common-word fertility toward 1–2 tokens/word.

This conceptual model explains the paper’s design goal: not merely incremental improvement, but movement from byte fallback toward subword and whole-word behavior for Brahmic text while leaving English, EU-language, code, and math segmentation largely inherited from **o200k_base**.

The most extreme case is **Odia**. The paper states that **Tekken/Sarvam-m**, despite being a 131K tokenizer, contains **zero tokens from the Oriya Unicode block**, so Odia text falls almost entirely to byte fallback [2605.29379]. The resulting disparity is illustrated by a sentence-level example: a **37-character Odia sentence** becomes **99 tokens** under Tekken/Sarvam-m, **39** under **o200k_cropped**, and **21** under BrahmicTokenizer-131K. The same mechanism appears in aggregate evaluation, where the paper reports a **4.31× token-count ratio** on the **60M-word Odia corpus slice**.

The paper also gives small lexical case studies showing the transition toward whole-word anchoring. The Hindi words **भारत** and **विद्यालय** tokenize as **single tokens** under BrahmicTokenizer-131K but as **two tokens** under Tekken/Sarvam-m. These examples are used to motivate the role of high-frequency word-level surgery slots.

## 3. Two-stage retrofit from o200k_base

BrahmicTokenizer-131K is constructed through a **two-stage retrofit pipeline** beginning from **o200k_base** [2605.29379].

The first stage is the **script-prune crop**. Starting from **200,019 tokens**, the tokenizer is reduced to **131,072** by removing **38,345 tokens** associated with nine out-of-scope writing systems. The resulting cropped tokenizer, **o200k_cropped**, contains **130,716 normal tokens plus 356 special tokens**.

| Removed script or writing system | Tokens removed |
|---|---:|
| CJK Unified Ideographs | 15,872 |
| Hangul | 6,401 |
| Hiragana + Katakana | 4,937 |
| Arabic | 3,204 |
| Cyrillic | 2,856 |
| Greek | 1,722 |
| Thai | 1,561 |
| Hebrew | 1,503 |
| Sinhala | 289 |

The crop also alters structural properties. The paper reports that **o200k_base** had **266 tokens >32 bytes** and **59 cross-script tokens**, whereas after cropping both counts become **zero**. Those properties are then preserved in the second stage.

The second stage is the **surgical retrofit**. Here the author identifies **2,372 “corpus-dead vocabulary slots”** in **o200k_cropped** and repurposes them for Brahmic content. The audit corpus used for this step contains **1.045 billion tokens** assembled from public AI4Bharat Indic resources plus Sarvam-AI’s **Samvaad-Hi**. On that audit corpus, **128,700** of the **131,072** tokens in **o200k_cropped** fire at least once. The **2,372 dropped slots** consist of:

- **2,083 tokens with zero fire rate**
- **289 marginal cases** with fire rate between **1 and 1,000 per billion audit tokens**

The paper reports an approximately **197,000×** difference between median kept and median dropped token fire rates, which is presented as evidence that the removed slots are operationally expendable.

The replacement budget is divided across several categories. The paper states that the surgery includes **5 character-infrastructure slots**, **1,443 word-level slots**, **922 further per-script Brahmic content slots**, and **157 numerals/danda/artifacts**. In the final accounting table, the same budget is summarized as **2,215 single-script Brahmic additions**, **149 Indic numeral merges**, **3 shared Brahmic punctuation tokens**, and **5 residual broken-UTF-8 artifacts**, totaling **2,372** [2605.29379].

## 4. Allocation policy, merge construction, and vocabulary inventory

A defining feature of the retrofit is its **script-aware allocation policy** rather than a purely frequency-driven retraining procedure. The **2,215 single-script Brahmic additions** are distributed as follows [2605.29379]:

- Oriya (Odia): **663**
- Gurmukhi (Punjabi): **328**
- Malayalam: **225**
- Gujarati: **203**
- Kannada: **177**
- Devanagari: **173**
- Bengali: **155**
- Telugu: **153**
- Tamil: **138**

This distribution is explicitly designed to favor the most underrepresented scripts. Oriya receives the largest intervention because it was essentially absent from the baseline 131K vocabulary class in competing general-purpose tokenizers.

The allocation is formalized as a **linear-programming / discrete concave allocation problem** over the nine Brahmic scripts:

$$
\max_{x \in \mathbb{Z}_{\geq 0}^{9}} \sum_{s \in S} c_s(x_s)
\quad \text{subject to} \quad
\sum_{s \in S} x_s = 2{,}215,\quad 0 \le x_s \le K_s \;\; \forall s \in S.
$$

The paper states that each $c_s$ is non-decreasing and concave by construction, so the optimization can be solved greedily by assigning each next slot to the script with largest current marginal gain:

$$
c_s(x_s+1) - c_s(x_s).
$$

The shipped heuristic and the LP-optimal greedy solver achieved **identical total savings** at the **2,215-slot budget**, although with slightly different per-script allocations because the optimum is flat across several allocations [2605.29379].

The replacement tokens are selected by **training BPE on the audit corpus restricted to the Brahmic-script subset**, then filtering candidate merges through a **no-cross-script-merge rule**. This rule forbids any new merge that combines bytes from two disjoint writing systems. The paper states that this filtering removed **4,292 candidate merges** and that **2,156** new merge-rule entries survived and were appended to the inherited merge list. The final tokenizer ships with **301,398 merge-rule entries**. The paper explains that this count is larger than the usual “vocab minus 256” intuition because the tokenizer was converted from **tiktoken** format to **Hugging Face BPE** format; the merge list enumerates valid decomposition paths rather than a canonical sequential BPE training history. The tokenizer sets `ignore_merges=true`, so whole-token vocabulary matches take priority and path multiplicity does not alter behavior.

The difference between **2,372 surgery slots** and **2,156 new merge entries** is also accounted for. Some new vocabulary entries were already reachable via existing merge chains in **o200k_cropped**, so they required only vocabulary insertion rather than a new merge rule. The same is true for atomic numeral and punctuation additions such as the **149 Indic numerals** and **3 shared Brahmic punctuation tokens** [2605.29379].

At the final inventory level, the paper reports that BrahmicTokenizer-131K contains **15,709 Brahmic-containing tokens**, versus **5,202** for Tekken/Sarvam-m, a **3.02×** increase. Script-specific comparisons are reported as:

- Devanagari: **4,165 vs 1,569**
- Bengali: **2,292 vs 839**
- Telugu: **1,508 vs 920**
- Kannada: **1,497 vs 570**
- Malayalam: **1,908 vs 406**
- Gujarati: **1,832 vs 204**
- Tamil: **1,130 vs 539**
- Oriya: **725 vs 0**
- Gurmukhi: **652 vs 155**

This yields a **Brahmic share of vocab** of **11.99%** for BrahmicTokenizer-131K versus **3.97%** for Tekken/Sarvam-m.

## 5. Empirical performance across Indic and non-Indic benchmarks

The main real-corpus benchmark uses a **27.12-million-document Indic pretraining corpus** assembled from public AI4Bharat resources plus Sarvam-AI’s **Samvaad-Hi** and the **MMLU translation set** [2605.29379]. After filtering to Indic-language rows, it contains:

- **27 million documents**
- **2.84 billion whitespace-delimited words**
- **46.21 GB of UTF-8 text**
- **11 Brahmic-script languages**

The headline result is that on this corpus BrahmicTokenizer-131K produces **26.75% fewer tokens than Tekken/Sarvam-m at the same 131K vocabulary budget**. The totals are **6,623M tokens** for BrahmicTokenizer-131K and **9,041M tokens** for Tekken/Sarvam-m, a reduction of about **2.42 billion tokens** on the same text [2605.29379].

The per-language reductions hold for all 11 Brahmic languages:

- Hindi: **1,231M vs 1,480M**, **−16.81%**
- Bengali: **1,638M vs 1,974M**, **−17.04%**
- Tamil: **684M vs 812M**, **−15.79%**
- Marathi: **529M vs 657M**, **−19.54%**
- Telugu: **599M vs 742M**, **−19.22%**
- Malayalam: **518M vs 738M**, **−29.81%**
- Kannada: **313M vs 387M**, **−19.11%**
- Punjabi: **210M vs 258M**, **−18.30%**
- Odia: **228M vs 984M**, **−76.79%**
- Gujarati: **605M vs 910M**, **−33.44%**
- Assamese: **67M vs 100M**, **−33.08%**

On **FLORES-200** across the 11 Brahmic languages, the tokenizer ranks **4th of 11** public tokenizers by mean Brahmic fertility, with mean **2.84 tokens/word**. The top three are **Sarvam-30B (2.14)**, **Sarvam-1 (2.30)**, and **Gemma-3-1B (2.70)**. However, at the **131K vocabulary class specifically**, BrahmicTokenizer-131K is the best in the comparison: **Tekken/Sarvam-m** has mean Brahmic fertility **4.87**, versus **2.84** for BrahmicTokenizer-131K, a **41.8% relative improvement**. On Odia in FLORES, the figures are **4.13 tokens/word** versus **18.18** for Tekken/Sarvam-m [2605.29379].

The non-Indic profile is presented as preservation of **o200k-grade** behavior rather than additional optimization. On FLORES-200 English, BrahmicTokenizer-131K has **1.235 tokens/word**, compared with **1.232** for **o200k_base** and **1.267** for Tekken/Sarvam-m. Internal validation against the cropped inherited tokenizer reports **63,919** English tokens for BrahmicTokenizer-131K versus **63,941** for **o200k_cropped**, a difference of **22 tokens** across **2,009 sentences**, with **96.8%** of sentences byte-identical at the per-sentence tokenization level [2605.29379].

On European languages, the reported fertilities are:

- French: **1.464 tokens/word**
- German: **1.653**
- Spanish: **1.388**

The paper characterizes this profile as within **2.5–3.0% of best** on each EU language.

On code and math, the tokenizer achieves:

- **HumanEval**: **0.295 tokens/character**
- **MBPP**: **0.320**
- **GSM8K**: **0.301**

Against Tekken/Sarvam-m at the same vocabulary budget, this is better by **4.0%**, **5.4%**, and **14.2%**, respectively. The paper attributes the result to inheritance from **o200k_base**, especially retention of **3-digit grouping** for numbers, illustrated by the example `"1234567890"` tokenizing as `"123" | "456" | "789" | "0"`. It further reports byte-identical tokenization on **100% of HumanEval** problems and **100% of MBPP-full** documents relative to **o200k_cropped**, and byte-identical tokenization for **38 common programming identifiers** [2605.29379].

A bytes-per-token view is also reported on a mixed-domain composite corpus. For BrahmicTokenizer-131K, the paper gives:

- English: **4.91 bytes/token**
- Brahmic: **1.83**
- Code: **3.39**
- Math: **3.32**
- All-corpus mean: **3.36**

The all-corpus means for several comparators are **Sarvam-30B: 3.30**, **Sarvam-1: 2.91**, **Gemma-3-1B: 3.21**, **Tekken/Sarvam-m: 3.00**, and **o200k_base / GPT-OSS-120B: 3.28**. On this basis, the paper argues that among tokenizers in or near the 131K budget class, BrahmicTokenizer-131K is the **only one simultaneously competitive on Brahmic, English, EU languages, code, and math**.

## 6. Comparisons, ablations, limitations, and release

The paper is explicit that BrahmicTokenizer-131K is **not** claimed to be the absolute best Indic tokenizer [2605.29379]. Against **Sarvam-30B** with a **262K vocab**, BrahmicTokenizer-131K loses on Indic compression: on the 27M corpus, **Sarvam-30B produces 18.7% fewer tokens**. Against **Sarvam-1** with **68K** and an Indic-only design, it loses by **12.7%** on Indic token count. **MUTANT-Indic** is treated as an important specialist precedent, but was **not publicly available** when the paper was prepared, so it was not directly benchmarked.

The argument is instead about a **general-purpose** operating point. Relative to BrahmicTokenizer-131K, **Sarvam-1** has:

- English fertility **1.431 vs 1.235**, or **15.9% worse**
- French/German/Spanish much worse
- code/math compression **26–33% worse**

Against **Sarvam-30B**, BrahmicTokenizer-131K still wins on all three code/math corpora despite having **half the vocabulary budget**:

- HumanEval: better by **13.2%**
- MBPP: better by **21.3%**
- GSM8K: better by **16.6%**

The paper attributes these trade-offs to vocabulary allocation. **Sarvam-30B** devotes **51,474 Brahmic-containing tokens** to Indic coverage, while **Sarvam-1** allocates **73.28%** of its vocabulary to Brahmic-containing tokens, compared with **11.99%** for BrahmicTokenizer-131K.

The ablation results further emphasize allocation policy. At the **2,215-slot** single-script budget, the reported token savings on the audit corpus are:

- **Shipped heuristic**: **6,949,473**
- **LP-optimal greedy**: **6,949,473**
- **Worst-script-first**: **6,947,436** (**−0.03%**)
- **Frequency-only**: **6,769,060** (**−2.60%**)
- **Equal per script**: **6,604,464** (**−4.96%**)

This is used to support the claim that exact proportions are not especially fragile, but that **biasing capacity toward under-resourced scripts matters**.

A second analysis examines utilization of the new merges:

- **0 fires**: **569 merges** (**26.4%**)
- **1–1,000 fires**: **312**
- **1,001–100,000 fires**: **681**
- **100,001–10,000,000 fires**: **487**
- **10,000,001+**: **107**

The paper acknowledges that the **26.4% unfired merge rate** may indicate either prudent future coverage or wasted budget.

The structural diagnostics are central to the tokenizer’s self-description. Relative to Tekken/Sarvam-m at the same 131K size, BrahmicTokenizer-131K has:

- **0 tokens >32 bytes** vs **56**
- **0 cross-script tokens** vs **8**
- **725 Oriya-block tokens** vs **0**

In the full 14-tokenizer structural table, only **BrahmicTokenizer-131K** and **o200k_cropped** satisfy both **zero cross-script tokens** and **no token longer than 32 bytes**.

The paper also states several limitations. First, **88% of the final vocabulary is inherited** from **o200k_cropped** and was not re-audited from first principles. Some inherited **“garbage” tokens** remain, including **broken-UTF-8 residue, zero-width/bidi-control-containing tokens, private-use-area tokens, and HTML-entity artifacts**, although these account for only **0.035%** of the vocabulary. Second, there is possible **domain mismatch**: the audit corpus and the 27M Indic evaluation corpus come from the same public source families, making the largest benchmark in-distribution; **FLORES-200** and **IN22-Gen** serve as out-of-distribution corroboration. Third, the work remains within the constraints of **byte-level BPE** rather than replacing it with a language-aware pre-tokenizer.

The release is under **Apache 2.0** at:

- `https://huggingface.co/theschoolofai/BrahmicTokenizer-131K`
- `https://github.com/theschoolofai/BrahmicTokenizer-131K`

The released materials include the `tokenizer.json`, internal audit tests, verification scripts, allocation tables, and reproduction scripts [2605.29379].

Taken together, these results position BrahmicTokenizer-131K as a tokenizer-engineering contribution rather than a new tokenization formalism. The paper’s claim is that **o200k_base** can be **surgically retrofitted** into a more Indic-capable 131K tokenizer by pruning out-of-scope scripts and repurposing dead or irrelevant slots according to a corpus-driven, script-aware policy. A plausible implication is that the work is best understood as a vocabulary-allocation correction to a known failure mode of generic multilingual byte-level BPE: it preserves inherited behavior where the baseline is already strong, while intervening selectively where frequency-driven training had left Brahmic scripts in byte fallback.

Source: https://www.emergentmind.com/topics/brahmictokenizer-131k