MT-Bench-Hi: Hindi Conversational Benchmark
- MT-Bench-Hi is a Hindi conversational benchmark that adapts the English MT-Bench framework with culturally relevant prompts for accurate native evaluation.
- It employs a hybrid methodology that combines translated and human-verified items with bespoke prompts to maintain both technical validity and cultural context.
- The benchmark uses an LLM-as-a-Judge scoring system to assess multi-turn dialogue quality, reasoning, and instruction following while preserving cross-lingual comparability.
MT-Bench-Hi is a Hindi adaptation of the English MT-Bench conversational benchmark, designed to evaluate open-ended, multi-turn instruction-following and dialogue quality in Hindi. Within the benchmark suite introduced in "Benchmarking Hindi LLMs: A New Suite of Datasets and a Comparative Analysis" (Kamath et al., 27 Aug 2025), it functions as the principal test of conversational ability, response quality, and reasoning over multiple turns in Hindi, while preserving the MT-Bench / LLM-as-a-Judge evaluation paradigm established for the original English benchmark (Zheng et al., 2023). Its defining premise is that direct translation of English conversational benchmarks is insufficient for Hindi, because literal transfer can erase cultural context and test translated-English phrasing rather than native Hindi competence.
1. Origin and benchmark rationale
MT-Bench-Hi inherits its basic evaluation philosophy from MT-Bench, which was introduced as a benchmark of 80 high-quality multi-turn questions intended to measure conversational ability across multiple turns while remaining challenging enough to distinguish reasoning, math, coding, and knowledge capabilities (Zheng et al., 2023). In the original formulation, MT-Bench items are two-turn dialogues, manually authored across 8 categories, and judged under an LLM-as-a-Judge framework designed to better track human preference in open-ended chat settings than traditional closed-ended benchmarks.
MT-Bench-Hi was proposed because Hindi LLM evaluation lacked a public Hindi counterpart to MT-Bench for assessing conversational ability, response quality, and reasoning over multiple turns (Kamath et al., 27 Aug 2025). The benchmark addresses a specific validity problem: English-style MT-Bench evaluation is not sufficient for Hindi because direct translation can erase or distort cultural context, and a model may end up being tested on translated-English phrasing rather than on native Hindi competence. This concern is especially salient for conversational tasks, where subtle linguistic choices, culturally grounded prompts, and turn-level coherence matter more than literal translation.
The benchmark’s motivation is therefore twofold. First, it fills a gap in instruction-tuned Hindi evaluation by providing a Hindi benchmark for multi-turn dialogue quality. Second, it attempts to make the benchmark genuinely Hindi rather than merely translated English by adapting some categories into culturally relevant Indian contexts (Kamath et al., 27 Aug 2025). A common misconception is that MT-Bench-Hi is simply a translated copy of MT-Bench; the dataset description explicitly rejects that characterization.
2. Construction methodology
MT-Bench-Hi was built using a hybrid approach that combines translated-and-verified items with prompts created from scratch (Kamath et al., 27 Aug 2025). This hybridization is the central methodological distinction between MT-Bench-Hi and a pure translation benchmark.
| Prompt subset | Creation method | Categories |
|---|---|---|
| Universal technical categories | Translated from English into Hindi using GCP, then human-verified | STEM, math, reasoning, and coding |
| Culturally sensitive / deeply contextual categories | Authored directly by human specialists | Writing, roleplay, humanities, and extraction |
The appendix-level workflow is more specific. Annotators were shown the original English MT-Bench question and follow-up, the corresponding model response example, and the AI judge’s rating/judgment for that response (Kamath et al., 27 Aug 2025). This setup was intended to help annotators understand what capability the original prompt was probing, after which they crafted analogous Hindi prompts contextualized for India. The resulting process is neither pure translation nor pure de novo authoring; it is an explicitly mixed construction pipeline.
This design has methodological significance for benchmark validity. The translated-and-verified component preserves comparability for technical categories, while the from-scratch component is used where direct translation would be culturally flat or unnatural. This suggests that the benchmark treats cultural adaptation not as a post hoc stylistic adjustment, but as part of the measurement design itself.
3. Dataset structure and task coverage
The paper states that the original MT-Bench contains 80 multi-turn questions, and Table 1 reports MT-Bench-Hi = 200 samples (Kamath et al., 27 Aug 2025). The benchmark is repeatedly described as using “extended dialogues,” with questions and follow-ups, and the appendix specifically notes that annotators used the original English question and follow-up as references when creating analogous Hindi conversations.
Its task organization mirrors the English MT-Bench structure, but with Hindi adaptation. The categories are divided into two broad scoring regimes:
| Evaluation regime | Categories |
|---|---|
| Reference-free, subjective categories | Writing, Roleplay, Humanities, Extraction, STEM |
| Reference-answer categories | Reasoning, Math, Coding |
The paper states that the final distribution spans the same broad MT-Bench-style categories, with a special emphasis on culturally adapted instructions (Kamath et al., 27 Aug 2025). Examples are not reproduced in the text, but the description specifies Indian-themed multi-turn prompts for subjective categories, technical questions translated into Hindi for STEM/math/reasoning/coding, and prompts that preserve the original instruction type while anchoring the content in Indian settings.
Compared with the original MT-Bench taxonomy, there is a notable categorical shift. The English benchmark groups its 80 questions into writing, roleplay, extraction, reasoning, math, coding, knowledge I (STEM), and knowledge II (humanities/social science) (Zheng et al., 2023). MT-Bench-Hi preserves the broad MT-Bench-style structure but expresses it in a Hindi evaluation suite through the categories listed above. A plausible implication is that category adaptation was guided less by literal taxonomic replication than by preserving the benchmark’s functional coverage under Hindi-specific prompt design.
4. Scoring protocol and LLM-as-a-Judge framework
MT-Bench-Hi follows the MT-Bench / LLM-as-a-Judge paradigm (Kamath et al., 27 Aug 2025). The original MT-Bench paper studied several judge modes, including pairwise comparison, single-answer grading, and reference-guided grading, and found that strong LLM judges such as GPT-4 could match human preferences well in multi-turn evaluation (Zheng et al., 2023). MT-Bench-Hi retains this general framework rather than introducing a new scoring formalism.
The paper states that the original MT-Bench uses a powerful judge model such as GPT-4o to score responses on a 1–10 scale, and MT-Bench-Hi preserves this framework (Kamath et al., 27 Aug 2025). The category-specific protocol is as follows:
- For reference-free categories such as Writing, Roleplay, Humanities, Extraction, and STEM, responses are scored directly by the judge model.
- For categories with reference answers such as Reasoning, Math, and Coding, responses are evaluated by pairwise comparison against the reference answer.
The reported benchmark metric is a single MT-Bench-Hi score on a 1–10 scale (Kamath et al., 27 Aug 2025). The paper does not give a new LaTeX scoring equation for MT-Bench-Hi itself; instead, it states that the evaluation framework is aligned with the original MT-Bench methodology and retains the same category-specific scoring logic.
One especially important comparability decision is that, for categories with reference answers, the authors retain the original English reference answer during evaluation (Kamath et al., 27 Aug 2025). Factually, this is presented as a mechanism for preserving comparability to the English MT-Bench setup. This suggests a deliberate trade-off between cross-lingual comparability and fully localized reference construction.
5. Quality control and Hindi adaptation choices
MT-Bench-Hi incorporates several explicit quality-control measures (Kamath et al., 27 Aug 2025). The benchmark was created by human specialists with Hindi proficiency and cultural familiarity, and annotators worked with a specialized interface, supplementary instructions for multi-turn benchmark creation, and examples of judged model outputs. These provisions were intended to guide prompt construction toward stress-testing the intended capability rather than merely paraphrasing existing English prompts.
The paper enumerates concrete quality-control steps:
- Specialist annotators: MT-Bench-Hi was created by human specialists with Hindi proficiency and cultural familiarity.
- Guided prompt creation: annotators were given reference examples from English MT-Bench, a specialized interface, supplementary instructions for multi-turn benchmark creation, and examples of judged model outputs.
- Human review: 50% of newly created samples were reviewed weekly by a developer, and revisions were requested when necessary.
- Preserving benchmark comparability: the same general evaluation structure was kept, and the original English reference answer was retained for reference-answer categories.
- Indian contextualization: culturally sensitive categories were authored from scratch instead of translated.
These choices clarify what MT-Bench-Hi is intended to measure. It is not only a Hindi-language fluency benchmark; it is a benchmark for culturally natural, instruction-tuned, multi-turn Hindi interaction under a scoring framework comparable to MT-Bench. The benchmark’s adaptation choices therefore concern both linguistic naturalness and evaluation continuity.
A second common misconception is that cultural adaptation implies loss of comparability. The paper instead describes a mixed strategy: it preserves the general evaluation structure and certain reference-answer mechanisms while localizing culturally sensitive prompts (Kamath et al., 27 Aug 2025). The benchmark is thus positioned as both adapted and comparable, rather than choosing exclusively between the two.
6. Empirical findings and comparative results
MT-Bench-Hi is one of the key columns in the comparative evaluation table reported in the Hindi benchmark suite paper (Kamath et al., 27 Aug 2025). The benchmark is used to compare open-source LLMs supporting Hindi, with results summarized as a single score on the 1–10 MT-Bench-Hi scale.
The paper highlights the following results:
- Best SLM: Gemma-2-9b-it, with 7.37.
- Best LLM: GPT-OSS-120B (reasoning low), with 8.70.
- Other notable results include Gemma-3-27b-it at 8.31, Sarvam-M at 8.25, and Llama-3.1-405B at 7.17.
The discussion draws several broader conclusions from these scores (Kamath et al., 27 Aug 2025). No single model dominates all tasks; model size is not the sole determinant of performance; language-specific or Indic-specific tuning can matter a great deal; and conversational ability in Hindi remains uneven across models. The paper additionally notes that GPT-OSS may have an inherent advantage due to its reasoning mode, even when set to “low,” and flags the possibility of judge bias because the judge is GPT-4o, a sibling OpenAI model.
The benchmark is also used diagnostically rather than only as a leaderboard. According to the discussion, MT-Bench-Hi reveals that stronger models exhibit better multi-turn coherence, stronger instruction following in Hindi, and more reliable reasoning and answer quality in conversational settings (Kamath et al., 27 Aug 2025). At the same time, smaller models can lag substantially in dialogue quality, and models that are strong in one domain may not generalize well across all MT-Bench categories.
7. Limitations and research significance
The paper is explicit about MT-Bench-Hi’s limitations (Kamath et al., 27 Aug 2025). It does not cover every possible instruction type or conversational scenario. Its use of an LLM-as-a-Judge introduces inherent bias, especially because Hindi nuance may not always be judged perfectly by the judge model. The translation-based subsets, even when human-verified, could still be improved by full human curation. For MT-Bench-Hi specifically, the paper implies that future work could expand the set of culturally grounded prompts and improve judge robustness for Hindi.
These limitations should be understood in light of the broader MT-Bench methodology. The original MT-Bench paper documented judge-side issues such as position bias, verbosity bias, self-enhancement bias, and limited reasoning ability in math and reasoning grading, while also showing that strong judges can still achieve high agreement with human preferences under carefully designed protocols (Zheng et al., 2023). MT-Bench-Hi inherits both the benefits and the constraints of that paradigm.
Within Hindi LLM evaluation, MT-Bench-Hi’s main significance is that it enables a more realistic comparison of instruction-tuned models in a setting where cultural naturalness, turn-level coherence, and open-ended response quality matter. The paper’s broader interpretation is that state-of-the-art English-centric models are not automatically best in Hindi, and that human-centered local adaptation matters for valid evaluation (Kamath et al., 27 Aug 2025). In that sense, MT-Bench-Hi is best understood as a benchmark for Hindi conversational competence under controlled cross-benchmark comparability, rather than as a simple localization of an English dataset.