BrahmicTokenizer-131K: Indic Tokenizer for 9 Scripts
- The tokenizer retrofit repurposes o200k_base by pruning non-Brahmic tokens and adding 2,372 Brahmic merge entries to address the Brahmic compression gap.
- It employs a two-stage pipeline—script-prune crop and surgical retrofit—to optimize tokenization for 9 Brahmic scripts while retaining compatibility with English, EU languages, code, and math.
- Empirical benchmarks demonstrate significant token reduction on Indic texts, with improvements up to 76.79% for languages like Odia compared to existing models.
Searching arXiv for the specified paper and closely related tokenizer work. arxiv_search(query="0", max_results=5) Searching for related Indic tokenizer papers for contextualization. arxiv_search(query="Indic tokenizer Sarvam MUTANT-Indic byte-level BPE Brahmic", max_results=10) BrahmicTokenizer-131K is a 131,072-vocabulary byte-level BPE tokenizer designed as an “Indic-capable drop-in replacement for o200k_base” while preserving the properties that made o200k_base attractive for English, European languages, code, and math (Shravan, 28 May 2026). Its central purpose is to close what the paper terms the “Brahmic compression gap”: the systematic under-allocation of tokenizer capacity to Brahmic scripts in mainstream multilingual tokenizers at comparable vocabulary sizes, which can force Indic text into inefficient byte-fragment tokenization. The design is not a new tokenization algorithm; it is a retrofit of o200k_base that keeps the same tokenizer-side interface and operational machinery, but changes the vocabulary contents and appends Brahmic merge entries so that the tokenizer becomes usable for nine Brahmic scripts covering 11 major Indian languages without giving up inherited non-Brahmic behavior.
1. Definition, compatibility, and intended scope
In the paper’s narrow technical sense, “drop-in replacement” does not mean vocabulary identity with o200k_base. It means that the tokenizer preserves the same byte-level BPE algorithm, the same GPT-2 ByteLevel pre-tokenizer, the same decoder, the same byte handling and byte fallback conventions, the same special-token format, and the same tokenizer JSON schema (Shravan, 28 May 2026). A system that already uses o200k_base can therefore replace the tokenizer artifact with BrahmicTokenizer-131K’s tokenizer.json without changing data-loader code, pre-tokenizer code, or decoder code.
What changes is the vocabulary and some added Brahmic merge-list entries. Consequently, the model’s embedding matrix and LM head require the standard vocabulary resize from 200,019 rows to 131,072 rows. The paper stresses that this is the same kind of mechanical resize downstream training code already supports when switching tokenizers.
The target deployment profile is explicitly limited. The tokenizer is optimized for:
- English
- three major EU languages: French, German, Spanish
- code
- math
- nine Brahmic scripts covering 11 major Indian languages
It is not optimized for scripts intentionally removed during the crop stage: CJK Unified Ideographs, Hangul, Hiragana + Katakana, Arabic, Cyrillic, Greek, Thai, Hebrew, and Sinhala. This makes it a replacement for o200k_base at the interface level, but not a universal replacement for every multilingual setting (Shravan, 28 May 2026).
A common point of confusion is the label “131K.” In this tokenizer, it denotes a 131,072-token vocabulary, not a long-context window. This distinction matters because the paper evaluates tokenization efficiency, not context-extension algorithms.
2. The Brahmic compression gap
The motivating claim is that multilingual tokenizers trained by generic frequency-based byte-level BPE on multilingual corpora can underperform on Brahmic scripts because frequency alone underproduces tokens for underrepresented scripts (Shravan, 28 May 2026). In the GPT-2 ByteLevel setup, Unicode text is first mapped through a printable-byte bijection. A Brahmic character such as a Devanagari letter becomes a multi-byte UTF-8 sequence, and if learned merges do not compose those bytes into script-specific tokens, tokenization falls back toward byte fragments.
The paper describes three regimes of Brahmic tokenization:
- Byte-fallback regime: no meaningful script-level merges; Brahmic characters can cost 3+ tokens each.
- Subword regime: characters and subword pieces are represented, leading to a few tokens per word.
- Whole-word regime: common words themselves become single tokens, pushing common-word fertility toward 1–2 tokens/word.
This conceptual model explains the paper’s design goal: not merely incremental improvement, but movement from byte fallback toward subword and whole-word behavior for Brahmic text while leaving English, EU-language, code, and math segmentation largely inherited from o200k_base.
The most extreme case is Odia. The paper states that Tekken/Sarvam-m, despite being a 131K tokenizer, contains zero tokens from the Oriya Unicode block, so Odia text falls almost entirely to byte fallback (Shravan, 28 May 2026). The resulting disparity is illustrated by a sentence-level example: a 37-character Odia sentence becomes 99 tokens under Tekken/Sarvam-m, 39 under o200k_cropped, and 21 under BrahmicTokenizer-131K. The same mechanism appears in aggregate evaluation, where the paper reports a 4.31× token-count ratio on the 60M-word Odia corpus slice.
The paper also gives small lexical case studies showing the transition toward whole-word anchoring. The Hindi words भारत and विद्यालय tokenize as single tokens under BrahmicTokenizer-131K but as two tokens under Tekken/Sarvam-m. These examples are used to motivate the role of high-frequency word-level surgery slots.
3. Two-stage retrofit from o200k_base
BrahmicTokenizer-131K is constructed through a two-stage retrofit pipeline beginning from o200k_base (Shravan, 28 May 2026).
The first stage is the script-prune crop. Starting from 200,019 tokens, the tokenizer is reduced to 131,072 by removing 38,345 tokens associated with nine out-of-scope writing systems. The resulting cropped tokenizer, o200k_cropped, contains 130,716 normal tokens plus 356 special tokens.
| Removed script or writing system | Tokens removed |
|---|---|
| CJK Unified Ideographs | 15,872 |
| Hangul | 6,401 |
| Hiragana + Katakana | 4,937 |
| Arabic | 3,204 |
| Cyrillic | 2,856 |
| Greek | 1,722 |
| Thai | 1,561 |
| Hebrew | 1,503 |
| Sinhala | 289 |
The crop also alters structural properties. The paper reports that o200k_base had 266 tokens >32 bytes and 59 cross-script tokens, whereas after cropping both counts become zero. Those properties are then preserved in the second stage.
The second stage is the surgical retrofit. Here the author identifies 2,372 “corpus-dead vocabulary slots” in o200k_cropped and repurposes them for Brahmic content. The audit corpus used for this step contains 1.045 billion tokens assembled from public AI4Bharat Indic resources plus Sarvam-AI’s Samvaad-Hi. On that audit corpus, 128,700 of the 131,072 tokens in o200k_cropped fire at least once. The 2,372 dropped slots consist of:
- 2,083 tokens with zero fire rate
- 289 marginal cases with fire rate between 1 and 1,000 per billion audit tokens
The paper reports an approximately 197,000× difference between median kept and median dropped token fire rates, which is presented as evidence that the removed slots are operationally expendable.
The replacement budget is divided across several categories. The paper states that the surgery includes 5 character-infrastructure slots, 1,443 word-level slots, 922 further per-script Brahmic content slots, and 157 numerals/danda/artifacts. In the final accounting table, the same budget is summarized as 2,215 single-script Brahmic additions, 149 Indic numeral merges, 3 shared Brahmic punctuation tokens, and 5 residual broken-UTF-8 artifacts, totaling 2,372 (Shravan, 28 May 2026).
4. Allocation policy, merge construction, and vocabulary inventory
A defining feature of the retrofit is its script-aware allocation policy rather than a purely frequency-driven retraining procedure. The 2,215 single-script Brahmic additions are distributed as follows (Shravan, 28 May 2026):
- Oriya (Odia): 663
- Gurmukhi (Punjabi): 328
- Malayalam: 225
- Gujarati: 203
- Kannada: 177
- Devanagari: 173
- Bengali: 155
- Telugu: 153
- Tamil: 138
This distribution is explicitly designed to favor the most underrepresented scripts. Oriya receives the largest intervention because it was essentially absent from the baseline 131K vocabulary class in competing general-purpose tokenizers.
The allocation is formalized as a linear-programming / discrete concave allocation problem over the nine Brahmic scripts:
The paper states that each is non-decreasing and concave by construction, so the optimization can be solved greedily by assigning each next slot to the script with largest current marginal gain:
The shipped heuristic and the LP-optimal greedy solver achieved identical total savings at the 2,215-slot budget, although with slightly different per-script allocations because the optimum is flat across several allocations (Shravan, 28 May 2026).
The replacement tokens are selected by training BPE on the audit corpus restricted to the Brahmic-script subset, then filtering candidate merges through a no-cross-script-merge rule. This rule forbids any new merge that combines bytes from two disjoint writing systems. The paper states that this filtering removed 4,292 candidate merges and that 2,156 new merge-rule entries survived and were appended to the inherited merge list. The final tokenizer ships with 301,398 merge-rule entries. The paper explains that this count is larger than the usual “vocab minus 256” intuition because the tokenizer was converted from tiktoken format to Hugging Face BPE format; the merge list enumerates valid decomposition paths rather than a canonical sequential BPE training history. The tokenizer sets ignore_merges=true, so whole-token vocabulary matches take priority and path multiplicity does not alter behavior.
The difference between 2,372 surgery slots and 2,156 new merge entries is also accounted for. Some new vocabulary entries were already reachable via existing merge chains in o200k_cropped, so they required only vocabulary insertion rather than a new merge rule. The same is true for atomic numeral and punctuation additions such as the 149 Indic numerals and 3 shared Brahmic punctuation tokens (Shravan, 28 May 2026).
At the final inventory level, the paper reports that BrahmicTokenizer-131K contains 15,709 Brahmic-containing tokens, versus 5,202 for Tekken/Sarvam-m, a 3.02× increase. Script-specific comparisons are reported as:
- Devanagari: 4,165 vs 1,569
- Bengali: 2,292 vs 839
- Telugu: 1,508 vs 920
- Kannada: 1,497 vs 570
- Malayalam: 1,908 vs 406
- Gujarati: 1,832 vs 204
- Tamil: 1,130 vs 539
- Oriya: 725 vs 0
- Gurmukhi: 652 vs 155
This yields a Brahmic share of vocab of 11.99% for BrahmicTokenizer-131K versus 3.97% for Tekken/Sarvam-m.
5. Empirical performance across Indic and non-Indic benchmarks
The main real-corpus benchmark uses a 27.12-million-document Indic pretraining corpus assembled from public AI4Bharat resources plus Sarvam-AI’s Samvaad-Hi and the MMLU translation set (Shravan, 28 May 2026). After filtering to Indic-language rows, it contains:
- 27 million documents
- 2.84 billion whitespace-delimited words
- 46.21 GB of UTF-8 text
- 11 Brahmic-script languages
The headline result is that on this corpus BrahmicTokenizer-131K produces 26.75% fewer tokens than Tekken/Sarvam-m at the same 131K vocabulary budget. The totals are 6,623M tokens for BrahmicTokenizer-131K and 9,041M tokens for Tekken/Sarvam-m, a reduction of about 2.42 billion tokens on the same text (Shravan, 28 May 2026).
The per-language reductions hold for all 11 Brahmic languages:
- Hindi: 1,231M vs 1,480M, −16.81%
- Bengali: 1,638M vs 1,974M, −17.04%
- Tamil: 684M vs 812M, −15.79%
- Marathi: 529M vs 657M, −19.54%
- Telugu: 599M vs 742M, −19.22%
- Malayalam: 518M vs 738M, −29.81%
- Kannada: 313M vs 387M, −19.11%
- Punjabi: 210M vs 258M, −18.30%
- Odia: 228M vs 984M, −76.79%
- Gujarati: 605M vs 910M, −33.44%
- Assamese: 67M vs 100M, −33.08%
On FLORES-200 across the 11 Brahmic languages, the tokenizer ranks 4th of 11 public tokenizers by mean Brahmic fertility, with mean 2.84 tokens/word. The top three are Sarvam-30B (2.14), Sarvam-1 (2.30), and Gemma-3-1B (2.70). However, at the 131K vocabulary class specifically, BrahmicTokenizer-131K is the best in the comparison: Tekken/Sarvam-m has mean Brahmic fertility 4.87, versus 2.84 for BrahmicTokenizer-131K, a 41.8% relative improvement. On Odia in FLORES, the figures are 4.13 tokens/word versus 18.18 for Tekken/Sarvam-m (Shravan, 28 May 2026).
The non-Indic profile is presented as preservation of o200k-grade behavior rather than additional optimization. On FLORES-200 English, BrahmicTokenizer-131K has 1.235 tokens/word, compared with 1.232 for o200k_base and 1.267 for Tekken/Sarvam-m. Internal validation against the cropped inherited tokenizer reports 63,919 English tokens for BrahmicTokenizer-131K versus 63,941 for o200k_cropped, a difference of 22 tokens across 2,009 sentences, with 96.8% of sentences byte-identical at the per-sentence tokenization level (Shravan, 28 May 2026).
On European languages, the reported fertilities are:
- French: 1.464 tokens/word
- German: 1.653
- Spanish: 1.388
The paper characterizes this profile as within 2.5–3.0% of best on each EU language.
On code and math, the tokenizer achieves:
- HumanEval: 0.295 tokens/character
- MBPP: 0.320
- GSM8K: 0.301
Against Tekken/Sarvam-m at the same vocabulary budget, this is better by 4.0%, 5.4%, and 14.2%, respectively. The paper attributes the result to inheritance from o200k_base, especially retention of 3-digit grouping for numbers, illustrated by the example "1234567890" tokenizing as "123" | "456" | "789" | "0". It further reports byte-identical tokenization on 100% of HumanEval problems and 100% of MBPP-full documents relative to o200k_cropped, and byte-identical tokenization for 38 common programming identifiers (Shravan, 28 May 2026).
A bytes-per-token view is also reported on a mixed-domain composite corpus. For BrahmicTokenizer-131K, the paper gives:
- English: 4.91 bytes/token
- Brahmic: 1.83
- Code: 3.39
- Math: 3.32
- All-corpus mean: 3.36
The all-corpus means for several comparators are Sarvam-30B: 3.30, Sarvam-1: 2.91, Gemma-3-1B: 3.21, Tekken/Sarvam-m: 3.00, and o200k_base / GPT-OSS-120B: 3.28. On this basis, the paper argues that among tokenizers in or near the 131K budget class, BrahmicTokenizer-131K is the only one simultaneously competitive on Brahmic, English, EU languages, code, and math.
6. Comparisons, ablations, limitations, and release
The paper is explicit that BrahmicTokenizer-131K is not claimed to be the absolute best Indic tokenizer (Shravan, 28 May 2026). Against Sarvam-30B with a 262K vocab, BrahmicTokenizer-131K loses on Indic compression: on the 27M corpus, Sarvam-30B produces 18.7% fewer tokens. Against Sarvam-1 with 68K and an Indic-only design, it loses by 12.7% on Indic token count. MUTANT-Indic is treated as an important specialist precedent, but was not publicly available when the paper was prepared, so it was not directly benchmarked.
The argument is instead about a general-purpose operating point. Relative to BrahmicTokenizer-131K, Sarvam-1 has:
- English fertility 1.431 vs 1.235, or 15.9% worse
- French/German/Spanish much worse
- code/math compression 26–33% worse
Against Sarvam-30B, BrahmicTokenizer-131K still wins on all three code/math corpora despite having half the vocabulary budget:
- HumanEval: better by 13.2%
- MBPP: better by 21.3%
- GSM8K: better by 16.6%
The paper attributes these trade-offs to vocabulary allocation. Sarvam-30B devotes 51,474 Brahmic-containing tokens to Indic coverage, while Sarvam-1 allocates 73.28% of its vocabulary to Brahmic-containing tokens, compared with 11.99% for BrahmicTokenizer-131K.
The ablation results further emphasize allocation policy. At the 2,215-slot single-script budget, the reported token savings on the audit corpus are:
- Shipped heuristic: 6,949,473
- LP-optimal greedy: 6,949,473
- Worst-script-first: 6,947,436 (−0.03%)
- Frequency-only: 6,769,060 (−2.60%)
- Equal per script: 6,604,464 (−4.96%)
This is used to support the claim that exact proportions are not especially fragile, but that biasing capacity toward under-resourced scripts matters.
A second analysis examines utilization of the new merges:
- 0 fires: 569 merges (26.4%)
- 1–1,000 fires: 312
- 1,001–100,000 fires: 681
- 100,001–10,000,000 fires: 487
- 10,000,001+: 107
The paper acknowledges that the 26.4% unfired merge rate may indicate either prudent future coverage or wasted budget.
The structural diagnostics are central to the tokenizer’s self-description. Relative to Tekken/Sarvam-m at the same 131K size, BrahmicTokenizer-131K has:
- 0 tokens >32 bytes vs 56
- 0 cross-script tokens vs 8
- 725 Oriya-block tokens vs 0
In the full 14-tokenizer structural table, only BrahmicTokenizer-131K and o200k_cropped satisfy both zero cross-script tokens and no token longer than 32 bytes.
The paper also states several limitations. First, 88% of the final vocabulary is inherited from o200k_cropped and was not re-audited from first principles. Some inherited “garbage” tokens remain, including broken-UTF-8 residue, zero-width/bidi-control-containing tokens, private-use-area tokens, and HTML-entity artifacts, although these account for only 0.035% of the vocabulary. Second, there is possible domain mismatch: the audit corpus and the 27M Indic evaluation corpus come from the same public source families, making the largest benchmark in-distribution; FLORES-200 and IN22-Gen serve as out-of-distribution corroboration. Third, the work remains within the constraints of byte-level BPE rather than replacing it with a language-aware pre-tokenizer.
The release is under Apache 2.0 at:
https://huggingface.co/theschoolofai/BrahmicTokenizer-131Khttps://github.com/theschoolofai/BrahmicTokenizer-131K
The released materials include the tokenizer.json, internal audit tests, verification scripts, allocation tables, and reproduction scripts (Shravan, 28 May 2026).
Taken together, these results position BrahmicTokenizer-131K as a tokenizer-engineering contribution rather than a new tokenization formalism. The paper’s claim is that o200k_base can be surgically retrofitted into a more Indic-capable 131K tokenizer by pruning out-of-scope scripts and repurposing dead or irrelevant slots according to a corpus-driven, script-aware policy. A plausible implication is that the work is best understood as a vocabulary-allocation correction to a known failure mode of generic multilingual byte-level BPE: it preserves inherited behavior where the baseline is already strong, while intervening selectively where frequency-driven training had left Brahmic scripts in byte fallback.