---
title: 'maiBERT: Maithili-Specific BERT Model'
url: https://www.emergentmind.com/topics/maibert
type: topic
---

# maiBERT: Maithili-Specific BERT Model

maiBERT is a BERT-based masked language model pre-trained specifically for Maithili, a low-resource Indo-Aryan language written in Devanagari and referred to in the paper also as Mithila. It was introduced to address the absence of large-scale pre-trained language models and the broader scarcity of computational resources for Maithili natural language understanding. The model is trained on a newly constructed Maithili corpus using Masked Language Modeling (MLM) and then fine-tuned for Maithili news classification, where it reports an accuracy of 87.02% and a macro-F1 of 86.9, outperforming the regional baselines evaluated in the study [2509.15048].

## 1. Origin, motivation, and problem setting

The motivation for maiBERT is rooted in a resource asymmetry. Maithili is described as being spoken by millions in Nepal and India, yet it has minimal computational resources and no large-scale pre-trained language model. The paper identifies this absence as a direct obstacle to Maithili’s participation in digital and AI-driven applications. It further characterizes Maithili as exhibiting dialectal variation, unique lexical items, and morphological structures that differ from Hindi and Nepali, which makes transfer from adjacent language models incomplete rather than automatic [2509.15048].

Within that framing, maiBERT is positioned as a language-specific alternative to multilingual and regional models such as mBERT, MuRIL, HindiBERT, NepBERTa, and NepaliBERT. The stated gap it fills is not merely data availability, but the ability to capture Maithili-specific syntax and orthography. The reported results on news classification are used as evidence that a Maithili-targeted encoder can outperform both multilingual models and regional Indic baselines on a downstream Maithili task. A plausible implication is that linguistic proximity to Hindi or Nepali does not eliminate the need for Maithili-specific pre-training when the target distribution is genuinely monolingual and orthographically normalized for Devanagari.

## 2. Corpus construction and text normalization

The pre-training corpus, named MithilaTextCorpus-v1, was built from diverse Maithili sources including online articles, books, and conversational data. The reported corpus statistics are 1,028,017 sentences, an average sentence length of 16.88 words, an average word length of 4.44 characters, and 1,357,081 unique tokens. The paper also reports that a Maithili-specific punctuation marker appears 1,044,037 times, indicating that punctuation behavior was treated as a salient corpus characteristic rather than incidental noise [2509.15048].

Preprocessing is described as extensive and includes normalization, lemmatization, and tokenization. The normalization pipeline is Devanagari-only, uses Unicode NFC, and filters non-Devanagari characters. An earlier preprocessing description mentions a customized `BertTokenizer` preserving Devanagari features, but the final pre-training setup uses a SentencePiece-BPE subword tokenizer with vocabulary size $|\mathcal{V}| = 30{,}000$. The coexistence of these descriptions suggests a staged pipeline in which script-preserving preprocessing preceded the final subword vocabulary induction. Code-switching handling is not described; the filtering of non-Devanagari characters suggests that non-Devanagari material was removed rather than explicitly modeled.

This corpus design is central to the paper’s interpretation of maiBERT’s gains. No formal ablation isolates corpus composition, tokenizer choice, or normalization strategy, but the Results and Discussion section qualitatively attributes the model’s advantage to the Maithili-specific corpus and Devanagari-aware preprocessing and tokenization.

## 3. Encoder design and MLM pre-training regime

maiBERT is implemented as a BERT-style Transformer encoder in TensorFlow. Positional encodings are added to token embeddings according to
$$
Z^{(0)} = E + P,
$$
where $E \in \mathbb{R}^{n \times d}$ and $P \in \mathbb{R}^{n \times d}$. The embedding dimension is reported as $d = 768$, and the maximum sequence length as $n \leq 512$. The model has 0.11 billion parameters. The paper does not specify the number of layers $L$, the number of attention heads $h$, or the intermediate feed-forward size, and it does not report architectural modifications such as whole-word masking. The attention mechanism follows the standard scaled dot-product form
$$
\mathrm{Attention}(Q, K, V) = \mathrm{softmax}\!\left(\frac{QK^\top}{\sqrt{d_k}}\right)V.
$$
Pre-training uses MLM only; Next Sentence Prediction is explicitly not used [2509.15048].

The training setup is reported with concrete hyperparameters: maximum sequence length 512, batch size 64, learning rate $\eta = 5 \times 10^{-5}$, Adam optimizer with bias correction and weight decay, dropout $p = 0.1$, weight decay $\lambda = 0.01$, early stopping with patience of 3, and total training steps $T = 1{,}000{,}000$. The masking ratio is $p = 0.15$ with the standard BERT replacement convention of 80% `[MASK]`, 10% random tokens, and 10% unchanged tokens, combined with dynamic masking. Hardware is reported as an NVIDIA RTX 4090 with 24 GB VRAM, mixed precision FP16, and TensorFlow XLA. The computational table reports memory usage of 4.6 GB, time of 0.85 hours, and an efficiency ratio of 172.11%, although the table does not clarify whether the time refers to pre-training or fine-tuning.

The paper also reports convergence statistics for MLM pre-training. Training loss and perplexity are listed as 0.45 and 1.23 at Epoch 1, 0.37 and 1.01 at Epoch 3, 0.21 and 1.00 at Epoch 6, and 0.012 and 1.00 at Epoch 10. These values are presented as evidence that the MLM objective stabilizes over training under the reported regime. Because no warmup or decay schedule is specified, the documented schedule is best understood as fixed learning rate training with early stopping rather than the more common warmup-decay configuration used in many BERT variants.

## 4. Downstream evaluation and reported empirical performance

The downstream benchmark is a Maithili news classification task with 11,959 labeled instances distributed across 10 classes: Culture, Economy, EduTech, Entertainment, Health, Interview, Literature, Opinion, Politics, and Sports. The split is reported as 70% train, 15% validation, and 15% test. Annotation quality is measured with Cohen’s Kappa of 0.82, which the paper describes as strong agreement. Fine-tuning uses `TFAutoModelForSequenceClassification`, dropout 0.1, and early stopping with patience 3. Exact downstream learning rate, batch size, and number of epochs are not specified. Evaluation uses accuracy, macro-precision, macro-recall, and macro-F1; micro metrics, statistical significance tests, confusion matrices, and detailed error analysis are not reported [2509.15048].

A data inconsistency is explicitly visible in the paper’s dataset reporting. The class-count table sums to 12,383 entries, whereas the labeled classification set used for training is stated as 11,959. The paper does not resolve this discrepancy, and therefore the benchmark should be read with that unresolved accounting issue in mind.

The main comparative results are as follows:

| Model | Accuracy | Macro-F1 |
|---|---:|---:|
| maiBERT | 87.02 | 86.9 |
| HindiBERT | 86.89 | 86.3 |
| NepBERTa | 86.4 | 86.0 |
| MuRIL | 86.3 | 86.0 |
| NepaliBERT | 82.0 | 81.3 |
| mBERT | 55.2 | 54.1 |

The full table also reports macro-precision and macro-recall: for maiBERT these are 86.8 and 87.0, respectively. The paper states an overall accuracy gain of 0.13% over the strongest baseline, HindiBERT, and claims 5–7% improvement across various classes, although per-class metrics are not tabulated. Traditional machine-learning baselines are also reported as weaker, including a linear SVM at 76.70% accuracy and a Neural Network at 72.62%. Within the experimental frame chosen by the paper, maiBERT therefore functions as a monolingual encoder that improves modestly over the strongest related-language Transformer baseline and substantially over the multilingual baseline mBERT.

## 5. Position within language-specific BERT research

maiBERT belongs to the broader class of language-specific BERT models that were surveyed comparatively in “What the [MASK]? Making Sense of Language-Specific BERT Models” [2003.02912]. That survey does not include maiBERT, which was published later, but its aggregate finding is that language-specific BERT models outperform mBERT on average across common tasks, including text classification, where the survey reports 88.96 accuracy for language-specific models versus 85.22 for mBERT. The survey also notes that gains can be especially pronounced in low-resource settings, where language-specific data curation compensates for the broad but diffuse coverage of multilingual pre-training [2003.02912].

maiBERT’s reported behavior is consistent with that survey-level pattern. In the paper’s own benchmark, the Maithili-specific model exceeds mBERT, MuRIL, and the evaluated Hindi- and Nepali-oriented baselines. The authors attribute this advantage qualitatively to Maithili-specific corpus construction and Devanagari-aware preprocessing rather than to an architectural departure from BERT. In that sense, maiBERT is best understood not as a new encoder family, but as a monolingual specialization of the BERT paradigm for a low-resource Indo-Aryan language whose orthographic and morpholexical profile is insufficiently captured by neighboring-language models.

This positioning matters methodologically. It indicates that monolingual specialization can remain beneficial even when multilingual encoders and related-language encoders already exist, particularly when the target language has distinct lexical inventory, dialectal variation, and script-normalization requirements.

## 6. Availability, applications, limitations, and name disambiguation

The model is open-sourced on Hugging Face as `rockerritesh/maiBERT_TF`, and the paper provides a pre-training corpus link at `https://huggingface.co/datasets/rockerritesh/maithiliNewsData`. Example usage is given for both masked language modeling through `TFAutoModelForMaskedLM` and downstream sequence classification through `TFAutoModelForSequenceClassification`. The paper presents news categorization as the demonstrated application and identifies sentiment analysis and Named Entity Recognition as intended downstream extensions. It also states that the Maithili-specific corpus and Devanagari-aware tokenizer can support civic information access, local media analytics, cultural preservation through digital tools, and educational, governmental, and media services in Maithili [2509.15048].

The limitations are explicit and substantial. The paper states that generative capabilities and domain adaptability are limited, especially for creative and poetic contexts. It does not address code-switching, does not analyze orthographic variation across dialects in detail, and does not evaluate generalization beyond classification. No formal ablation studies are reported for vocabulary size, corpus size, domain composition, script normalization, or tokenization choice. The paper also omits several reproducibility details: deduplication and language identification steps in corpus creation, learning-rate warmup or decay schedules, random seeds, framework versions, pre-training checkpoints, license terms, and model card specifics. A plausible implication is that the model may be most reliable on formal written Maithili close to the news-domain distribution from which much of the training and evaluation data are drawn.

Future work is described as expanding the corpus with conversational, literary, and informal texts, and exploring causal language modeling and instruction tuning. The same section suggests that broader corpora could mitigate topic and register bias introduced by the current emphasis on public Maithili news sites and formal written material.

Because of name similarity, maiBERT should also be distinguished from two unrelated models. “MaBERT” is a padding-safe interleaved Transformer–Mamba hybrid encoder for efficient extended-context masked language modeling, rather than a Maithili-specific BERT [2603.03001]. “MauBERT” is a multilingual extension of HuBERT for few-shot acoustic units discovery in speech, rather than a text encoder for Maithili NLU [2512.19612]. The similarity is nominal only; the architectures, modalities, and research problems are different.

Source: https://www.emergentmind.com/topics/maibert