---
title: 'Samiksha: Community LLM Benchmark'
url: https://www.emergentmind.com/topics/samiksha
type: topic
---

# Samiksha: Community LLM Benchmark

Searching arXiv for recent papers on “Samiksha” and closely related community-centered LLM evaluation in healthcare.
Samiksha is a community-driven evaluation pipeline for benchmarking LLM-based healthcare chatbots in ways that are grounded in the lived realities of actual users, particularly in culturally and linguistically diverse settings in India. The term is of Sanskrit origin and means “analysis, review or thorough investigation.” In the paper that introduces it, Samiksha is defined simultaneously as an end-to-end evaluation pipeline, a methodology for benchmark construction, and a framework for community-centered LLM evaluation in which community feedback informs what to evaluate, how the benchmark is built, and how outputs are scored [2509.24506].

## 1. Definition, scope, and conceptual basis

Samiksha was proposed in response to a specific critique of contemporary LLM evaluation: general-purpose benchmarks and many domain-specific benchmarks do not adequately reflect the social, cultural, linguistic, and practical conditions in which people actually use AI systems. In healthcare chatbot settings, the paper argues that a response cannot be assessed only for medical plausibility. It must also be understandable, culturally aware, socially appropriate, sensitive to stigma, and aligned with the user’s practical context [2509.24506].

The pipeline is therefore centered on everyday community health needs rather than abstract model capability. Its motivating assumptions are explicit. Health queries are often indirect or incomplete; culturally local beliefs and practices shape what counts as a useful answer; communities may face stigma, family surveillance, low access to care, dialect variation, or mistrust; and harmful outputs may arise not only from factual mistakes but also from cultural misalignment. This position distinguishes Samiksha from translated, synthetic, or institutionally curated medical QA benchmarks that may elevate institutional priorities above those of the communities most affected [2509.24506].

A common misconception is that a multilingual or healthcare-specific benchmark is necessarily contextually grounded. Samiksha rejects that equivalence. It treats benchmark construction itself as the primary methodological problem, and not merely the downstream scoring of model outputs. This suggests that, in high-stakes domains, evaluation design is inseparable from questions of representation, inclusion, and social fit.

## 2. Stakeholders, institutional structure, and deployment context

Samiksha relies on three actor groups: civil-society organizations, community members or data workers, and researchers or partner organizations. CSOs provide domain expertise, regional familiarity, exposure to ground-level healthcare realities, insight into typical user concerns, expectations for ethical chatbot behavior, and guidance on evaluation criteria. The paper emphasizes that these organizations are not treated as generic annotators; they are described as culturally rooted and technically equipped to provide actionable evaluation inputs [2509.24506].

Community members or data workers generate benchmark queries, localize chatbot logs into target languages, and evaluate model outputs. Their lived experience is treated as central rather than incidental. Researchers and partner organizations operationalize the pipeline. The three collaborating organizations had distinct roles: Karya handled recruitment, training, task deployment, coordination, quality control, and ethical data-work infrastructure; the Collective Intelligence Project helped design the community-centered study and evaluation process; and Microsoft Research India provided methodological guidance and data-creation support [2509.24506].

The demonstration setting was the healthcare domain in India over three months, from June to August 2025. The study consulted 5 CSOs, involved 15 data workers for benchmark creation, and later involved 23 human evaluators for response evaluation, of whom 12 had also participated in query creation. The benchmark itself covered Hindi, Malayalam, and Kannada. The CSOs worked across multiple language settings, including Hindi, Marathi, Telugu, Kannada, English, and Malayalam, while the represented user populations included rural and economically disadvantaged communities, marginalized and underrepresented populations, women facing reproductive or maternal health stigma, caregivers, community health workers, multilingual Indian users, and people navigating local cultural beliefs and family authority structures [2509.24506].

The paper is careful not to claim that CSOs fully represent all users. It reports the view that CSOs cannot speak on behalf of users, but can identify practices and concerns that have worked in their settings. This is an important qualification, because it places Samiksha within a participatory but non-totalizing model of representation.

## 3. The three-phase pipeline

Samiksha is organized into three named phases: Query Curation, Query Generation, and Response Evaluation. These phases form the paper’s end-to-end workflow for grounding chatbot benchmarking in community realities [2509.24506].

In **Query Curation**, the goal is to identify realistic healthcare query themes. The authors conducted semi-structured interviews with 5 CSOs, mostly in English, with some blending English with Hindi or Malayalam. The interview protocol covered community health needs, prior digital tool use, chatbot deployment and adoption, common health queries, harms from false information, desirable features of an ethical chatbot, response “green flags” and “red flags,” and willingness for ongoing collaboration. Interviews were not recorded; detailed notes were taken; and questions were iteratively refined until the authors reached theoretical saturation. From thematic analysis, recurring points were grouped into 15 themes and then consolidated into 8 broader healthcare themes [2509.24506].

In **Query Generation**, those themes were converted into benchmark items through two workflows. One workflow asked data workers to create queries from scratch in their native language. Instructions asked them to imagine asking a healthcare expert such as a doctor, nurse, or pharmacist; write realistic queries; include context where needed; allow multi-part or chained questions; and avoid personal information. The second workflow localized de-identified chatbot logs shared by CSOs into the target language while preserving intent. Karya trained workers through learning modules on chatbot interaction, AI evaluation, and query guidelines, and tasks were run on Karya’s annotation platform [2509.24506].

In **Response Evaluation**, benchmark queries were answered by multilingual LLMs and then scored by human evaluators and LLM judges. Responses were generated under a uniform setup: the same instruction template, a designated “health expert” role, a target-language constraint, and a common word limit to mitigate response-length bias. The models used their recommended decoding hyperparameters. Evaluation then proceeded in two modes: standalone rubric-based scoring and comparative pairwise judgment [2509.24506].

The significance of the three-phase design lies in where community input enters. It does not appear only at the final annotation stage. Community participation shapes topic selection, example construction, and output judgment, thereby influencing all three questions the paper foregrounds: what to evaluate, how to build the benchmark, and how to score outputs.

## 4. Benchmark construction and content

The benchmark begins with CSO interviews rather than an inherited medical ontology. This methodological choice produces queries that are not merely health questions in Indian languages, but health questions shaped by local health-seeking behavior, social determinants, household power structures, stigma, traditional beliefs, and practical care dilemmas [2509.24506].

The paper reports that the 15 interview-derived themes were consolidated into 8 broader healthcare themes: access to community healthcare or primary healthcare; managing injuries and infectious disease; managing chronic conditions; wellness habits; reproductive health; maternal health; children’s health; and senior care. Table 3 also contains “Everything else” as a catch-all category, which introduces a small inconsistency relative to the repeated claim that the benchmark uses eight healthcare topics [2509.24506].

The benchmark is multilingual in Hindi, Kannada, and Malayalam. The paper does not state that it systematically includes code-mixed inputs, and it is therefore not possible to confirm code-mixing as a benchmark property. What is explicit is that the benchmark contains non-English, culturally specific, and locally grounded examples in a broader low-resource multilingual setting [2509.24506].

For the query-creation workflow, each participant drafted about 50 queries; workers produced 270 data points per language, totaling 810 data points across the three languages, evenly distributed across three languages and eight topics. For the localization workflow, the paper reports 260 data points per language, totaling 780 localized queries. The benchmark used for model evaluation was the 810 data points generated via query creation, although one sentence in the paper describing models answering “around half the queries in the benchmark” introduces an editorial ambiguity [2509.24506].

The benchmark’s examples are deliberately multi-layered and embedded in family and community life. They include confidentiality concerns with gynecologists, fasting during Shivratri with low blood pressure, beliefs about mango or jackfruit in pregnancy, papaya leaves for dengue, caring for an elder who refuses dietary restrictions, unofficial payments at village clinics, fertility rumors linked to wearing jeans, and conflicts between elder women’s advice and hospital breastfeeding protocol. These cases encode stigma, superstition, indigenous remedies, family control, healthcare access barriers, and tensions between traditional and biomedical knowledge [2509.24506].

## 5. Evaluation design, rubric, and empirical findings

Three multilingual LLMs were evaluated: Sarvam-M, Qwen3-235B-A22B, and Llama-3.1-405B-Instruct. Responses were assessed by 23 native-speaker human evaluators, GPT-4o as an LLM judge, and Sarvam-M as an LLM judge. The use of Sarvam-M both as an evaluated model and as a judge was intentional, in order to probe possible self-bias [2509.24506].

Standalone evaluation used four rubric dimensions derived from CSO feedback: **Clarity & Fluency**, **Helpfulness & Relevance**, **Accuracy (General Perception)**, and **Completeness & Conciseness**. Each dimension used the labels Yes, Somewhat, and No. Both humans and LLM judges produced rationales; human evaluators also recorded voice notes explaining their ratings. Pairwise evaluation compared all answer pairs for the same query:
$$
\{Q, A_1, A_2\}, \quad \{Q, A_1, A_3\}, \quad \{Q, A_2, A_3\}
$$
and annotators answered, “Which response do you think is better?” with options A, B, or Not sure [2509.24506].

The paper states that standalone ratings were aggregated into mean scores per model across items, criteria, and languages, while pairwise comparisons were aggregated into pairwise win-rates and per-model win-share summaries. However, it does not provide explicit formulas for ordinal-to-numeric conversion, exact mean-score computation, exact win-share calculation, tie handling, confidence intervals, significance tests, or p-values [2509.24506].

The main empirical findings concern judge behavior as much as model behavior. LLM judges produced compressed, near-ceiling scores with relatively little differentiation between models, whereas human judges produced lower and more dispersed scores. Coarse ranking under standalone scoring was broadly similar, with Qwen3 generally highest, Sarvam-M close behind, and Llama-3.1-405B lowest. In pairwise evaluation, Llama-3.1-405B was consistently the least preferred model, but the ordering of the two stronger models depended on the judge: humans favored Qwen3 overall, GPT-4o favored Sarvam-M, and Sarvam-M-as-judge gave the highest win rate to Qwen3 [2509.24506].

Inter-evaluator alignment was summarized with Pearson correlations over item-level mean scores:
$$
r(\text{Sarvam-M judge}, \text{GPT-4o judge}) \approx 0.40
$$
$$
r(\text{LLM judge}, \text{human}) \approx 0.13
$$
These values indicate moderate LLM–LLM agreement and weak human–LLM agreement. The paper interprets this as evidence that automated judges share a model-centric notion of quality that differs from community-grounded human judgment [2509.24506].

## 6. Limits, misconceptions, and broader implications

Samiksha is explicitly a pilot deployment. Its scale remains limited to one domain, one country context, three languages, five CSOs, and a relatively small cohort of workers and evaluators. Query generation required heavy coordination, topic-specific examples introduced bias toward analogous outputs, and localization from chatbot logs often preserved context-thin fragments rather than yielding strong standalone benchmark items. The paper also notes that some healthcare judgments may require domain experts in addition to community evaluators or LLM judges [2509.24506].

A second misconception addressed by the work is that LLM-as-judge can replace human evaluation in multilingual, culturally grounded healthcare settings. Samiksha’s evidence argues against that conclusion. The paper identifies three specific weaknesses of LLM judges in this setting: compression or ceiling effects, under-detection of completeness and nuanced factual inconsistency, and lower language-specific sensitivity than humans. Its practical recommendation is a hybrid workflow in which LLM judges are used for high-throughput ranking and triage, while targeted human annotation is retained for validation and calibration, especially in important or risky cases [2509.24506].

The work also has broader implications for benchmark design. It argues that responsible evaluation should be built from community realities rather than translated abstractions, should involve stakeholders throughout the pipeline, should combine community and expert perspectives, and should support multilingual and culturally grounded assessment. In healthcare, the implication is that passing standard benchmarks is not sufficient. Systems intended for real communities must be tested on local languages, realistic social dilemmas, culturally sensitive scenarios, and safety-critical use cases where partial correctness is inadequate [2509.24506].

Ethically, the paper reports institutional review approval, approval by Karya, use of de-identified data, compensation for data workers, and avoidance of personally identifiable information in queries. Substantively, its core caution is that healthcare chatbot evaluation must consider not only whether a model answers, but whether it answers responsibly in context. This suggests a generalizable principle: evaluation quality in high-stakes AI depends not only on metric design, but on whose realities define the benchmark in the first place.

Source: https://www.emergentmind.com/topics/samiksha