SinBERT: Sinhala Error Classifier
- SinBERT is a Sinhala-specific language model that diagnoses dyslexia-inspired errors in transcribed text using a RoBERTa-based architecture.
- It leverages extensive pre-training on the sin-cc-15M corpus to capture Sinhala morphological and syntactic patterns for effective sentence-level classification.
- As an error classification component in an accessibility pipeline, SinBERT guides downstream corrections by predicting substitution, insertion, omission, or reversal errors.
SinBERT most directly denotes, in the literature surveyed here, a Sinhala-specific LLM used as the error-classification component of a low-resource speech-driven NLP pipeline for Sinhala dyslexia assistance. In that system, SinBERT is described both as an open-sourced fine-tuned BERT model trained for Sinhala and, more specifically, as a Sinhala-specific LLM built on the RoBERTa architecture, extensively pre-trained on sin-cc-15M and then fine-tuned for sentence-level classification. Its function is diagnostic rather than generative: it predicts the dominant dyslexia-inspired error type in a transcribed Sinhala sentence, and that prediction is used to guide subsequent correction by mT5 and Mistral (Perera et al., 6 Oct 2025).
1. Nomenclature and scope
The label “SinBERT” is not lexically stable across neighboring research areas. Closely named models include SenseBERT, a BERT variant that adds WordNet supersense prediction to masked language modeling (Levine et al., 2019); Sentence-BERT (SBERT), a siamese/triplet modification of BERT for sentence embeddings and cosine-based similarity search (Reimers et al., 2019); and SindBERT, a Turkish-specific, encoder-only, RoBERTa-style transformer trained from scratch (Scheible-Schmitt et al., 24 Oct 2025). This suggests that “SinBERT” often functions as an approximate or context-dependent identifier rather than a single canonical model family.
Within that ambiguous naming landscape, the clearest direct use of the exact form “SinBERT” is in the Sinhala dyslexia-assistance system, where it names the intermediate classifier responsible for error identification in transcribed Sinhala text (Perera et al., 6 Oct 2025). The term should therefore be interpreted in context: in Sinhala accessibility work it refers to a Sinhala-specific classifier; in lexical-semantics contexts it may be confused with SenseBERT; in sentence-embedding contexts with SBERT; and in Turkish NLP with SindBERT.
| Name | Referent | Distinguishing description |
|---|---|---|
| SinBERT | Sinhala classifier in dyslexia-assistance pipeline | Error-type prediction for Sinhala text |
| SenseBERT | Lexical-semantic BERT variant | Predicts masked words and WordNet supersenses |
| Sentence-BERT (SBERT) | Sentence embedding model | Siamese/triplet BERT for cosine similarity |
| SindBERT | Turkish monolingual encoder | Large-scale RoBERTa-style Turkish model |
2. Architectural characterization in Sinhala NLP
The dyslexia-assistance paper characterizes SinBERT as a Sinhala-specific LLM built on the RoBERTa architecture and previously introduced for Sinhala NLP by Dhananjaya et al. (2022) (Perera et al., 6 Oct 2025). It is said to have been extensively pre-trained on sin-cc-15M, a large monolingual Sinhala web corpus. The paper attributes to that pretraining coverage of Sinhala morphological, syntactic, and semantic patterns, and it treats this as especially relevant in a low-resource and agglutinative language setting.
The same account emphasizes two architectural properties. First, SinBERT inherits RoBERTa’s pretraining setup. Second, it benefits from subword-level representations, which the paper presents as useful for Sinhala because word forms can be morphologically complex (Perera et al., 6 Oct 2025). No new encoder architecture is introduced within the dyslexia-assistance system itself; rather, SinBERT is reused as a language-specific pretrained backbone and then fine-tuned for a downstream sentence-classification task.
This positioning is important for classification of the model’s research role. SinBERT is not presented there as a general-purpose generative model, a speech recognizer, or a correction model. It is instead treated as a language-aware encoder whose principal value lies in sentence-level diagnosis of error patterns before downstream correction (Perera et al., 6 Oct 2025).
3. Functional role and error taxonomy
In the Sinhala dyslexia-assistance pipeline, SinBERT performs error classification rather than correction itself. After speech input is transcribed by Whisper, SinBERT takes the transcribed Sinhala sentence and predicts which predefined dyslexia-like error category is most likely present. The paper defines four primary error categories: substitution, insertion, omission, and reversal (Perera et al., 6 Oct 2025).
Operationally, this makes SinBERT a 4-class sentence classifier. The paper describes its task as predicting the dominant dyslexia-related error type present in a transcribed Sinhala sentence, and it further states that this prediction is used to adaptively guide the next correction stage (Perera et al., 6 Oct 2025). If the predicted class is omission, the correction stage can be prompted to restore missing characters or particles; if substitution, to replace wrong letters; similarly for insertion and reversal. The classifier therefore serves as an error-aware conditioning mechanism between transcription and correction.
The system-level sequence is explicit: user speech input, preprocessing, Whisper for speech-to-text, SinBERT for error classification, mT5 for core correction, Mistral for fluency or stylistic refinement, and gTTS for text-to-speech playback (Perera et al., 6 Oct 2025). Within that architecture, SinBERT is the diagnostic classifier that makes the correction stage adaptive rather than generic.
4. Data construction and fine-tuning regime
The fine-tuning setup is dictated by data scarcity. The paper states that there is no large annotated dyslexic Sinhala corpus, so the authors create a synthetic dataset instead (Perera et al., 6 Oct 2025). The source corpus is the OpenSLR SLR63 Sinhala Read Speech corpus. From this, they generate a 3,000-sample parallel dataset by taking clean Sinhala sentences and applying rule-based transformations to simulate substitution, insertion, omission, and reversal errors. Each dyslexic sentence is paired with its original clean version, and audio paths are preserved for Whisper evaluation (Perera et al., 6 Oct 2025).
The dataset is split 80% train and 20% test using stratified sampling by error type, so that the four labels remain balanced (Perera et al., 6 Oct 2025). SinBERT is then fine-tuned on this synthetic dyslexic error dataset for the sentence-classification task. The paper describes it as a lightweight and efficient sentence-level classifier, but it provides only limited low-level optimization detail. It does not specify the learning rate, batch size, optimizer, number of epochs, maximum sequence length, exact classification head architecture, loss function, or hardware used for training (Perera et al., 6 Oct 2025).
The preprocessing description is also asymmetrical. The paper gives more detail for Whisper than for SinBERT: audio is resampled to 16 kHz for Whisper, and the synthetic dataset is created through rule-based error injection (Perera et al., 6 Oct 2025). A system diagram mentions tokenization, stop-word removal, lemmatization, phonetic normalization, statistical analysis, and edit distance methods, but the main text does not explicitly attribute those operations to the SinBERT classifier itself. A cautious reading is therefore that SinBERT consumes transcribed Sinhala text using its pretrained subword tokenizer, while the broader diagram reflects system-level conceptual processing rather than a fully specified SinBERT preprocessing pipeline.
5. Empirical status and reported performance
The evaluation methodology gives SinBERT a central architectural role but only limited isolated measurement. The paper states that evaluation covers speech-to-text, error classification, and correction, yet the reported quantitative results focus overwhelmingly on Whisper transcription, mT5 plus Mistral correction, and overall end-to-end system accuracy (Perera et al., 6 Oct 2025). No dedicated classification accuracy, macro-F1, precision, recall, or confusion matrix is reported for SinBERT alone.
The only explicit standalone performance figure attributable directly to SinBERT is inference latency: SinBERT classification is reported at 60 ms (Perera et al., 6 Oct 2025). Its effectiveness is otherwise assessed indirectly through the full system. The abstract reports 0.66 transcription accuracy, 0.7 correction accuracy, and 0.65 overall system accuracy (Perera et al., 6 Oct 2025). For the mT5-small plus Mistral API correction stage, the paper reports BLEU 0.359, GLEU 0.575, Accuracy 0.70, WER 0.322, and Edit Distance 1.66. For Whisper-Sinhala it reports BLEU 0.279, GLEU 0.444, Accuracy 0.659, WER 0.333, and Edit Distance 0.545 (Perera et al., 6 Oct 2025).
The baseline comparison is likewise system-level rather than classifier-level. The baselines named are rule-based correction using dictionary logic and the mT5 plus Mistral full-pipeline correction system (Perera et al., 6 Oct 2025). The paper states that the combination of error classification and hybrid correction significantly outperformed both isolated correction methods and rule-based strategies, which implies a positive contribution from SinBERT. However, there is no ablation isolating performance without SinBERT, with SinBERT, or with generic prompting versus error-aware prompting. The magnitude of SinBERT’s independent contribution therefore remains unspecified.
6. Limitations and broader significance
Several limitations directly affect interpretation of SinBERT. The largest is that the classifier is trained on synthetic dyslexic errors rather than real dyslexic writing at scale (Perera et al., 6 Oct 2025). This creates a distributional risk: the model may learn the regularities of the rule-based perturbation generator more strongly than the full variability of authentic dyslexic Sinhala. A second limitation is the absence of standalone classifier metrics, which prevents precise assessment of per-class strengths, calibration, or confusion structure. A third is label ambiguity: the paper notes that the system can misclassify ambiguous omissions, suggesting overlap among the four error categories (Perera et al., 6 Oct 2025).
A further complication is interaction with upstream ASR noise. SinBERT operates on Whisper-transcribed text, so its input may contain a mixture of dyslexia-inspired errors and transcription errors (Perera et al., 6 Oct 2025). The paper does not disentangle those sources of noise. This makes the classifier’s effective operating distribution more complex than the clean synthetic training formulation alone would suggest.
Despite those limitations, the model’s significance is clear within the paper’s application domain. SinBERT injects Sinhala-specific linguistic knowledge into a low-resource accessibility pipeline and supplies a compact control signal for downstream correction (Perera et al., 6 Oct 2025). A plausible implication is that its main contribution is not raw classification novelty but modularity: it inserts an interpretable intermediate diagnosis between speech recognition and text correction. In that sense, SinBERT is best understood as a specialized Sinhala error-identification engine embedded in a multimodal feedback loop, rather than as a general-purpose benchmark model or a broadly standardized BERT derivative.