Papers
Topics
Authors
Recent
Search
2000 character limit reached

J-HealthBench: A Japanese Medical Benchmark

Updated 12 July 2026
  • J-HealthBench is a proposed localized adaptation of HealthBench that evaluates Japanese medical conversational AI using open-ended, rubric-based clinical dialogues.
  • It addresses evaluation gaps in Japanese medical NLP by localizing clinical guidelines, legal requirements, and cultural norms to ensure safety and accuracy.
  • Empirical findings reveal that while most HealthBench scenarios transfer directly, a significant fraction require adjustments to fit Japan’s unique healthcare context.

Searching arXiv for the cited HealthBench and J-HealthBench papers to ground the article. J-HealthBench is a proposed localized adaptation of HealthBench for the Japanese medical system. Rather than constituting a finished benchmark, it is presented as a design direction and evidence-based roadmap for evaluating medical LLMs in Japan through open-ended, rubric-based, safety-oriented clinical dialogue. The proposal arises from the finding that direct translation of HealthBench into Japanese is insufficient: a medical benchmark encodes not only language, but also clinical guidelines, legal and insurance systems, available treatments and hotlines, disease prevalence, and culturally normative advice. In this formulation, J-HealthBench is intended to preserve HealthBench’s behavioral evaluation paradigm while localizing it to Japanese clinical practice, Japanese law, the Japanese healthcare system, and Japanese communication norms (Hisada et al., 22 Sep 2025).

1. Foundation in HealthBench

HealthBench is the immediate methodological precursor to J-HealthBench. It is an open-source benchmark consisting of 5,000 multi-turn conversations between a model and either an individual user or a healthcare professional, with responses evaluated through conversation-specific rubrics created by 262 physicians. Its scope is explicitly open-ended rather than multiple-choice, and it is organized around seven themes: emergency referrals, context seeking, global health, health data tasks, expertise-tailored communication, responding under uncertainty, and response depth. The benchmark reports performance along five behavioral axes: accuracy, completeness, context awareness, communication quality, and instruction following (Arora et al., 13 May 2025).

This structure is central to the rationale for J-HealthBench. HealthBench was designed to assess realistic health interactions that may be vague, context-dependent, safety-sensitive, and multi-turn, and its rubric-based evaluation can reward correct, safe, and context-aware behavior while penalizing omissions or harmful advice. Subsequent applied work has used HealthBench Hard, a 1,000-example difficult subset, to evaluate agentic, retrieval-augmented clinical assistants in high-stakes scenarios, reinforcing the benchmark’s role as a behavior-level clinical evaluation framework rather than a narrow factual-recall test (Ravichandran et al., 29 Aug 2025).

For the Japanese setting, HealthBench is therefore attractive not because it is in English, but because it evaluates the kinds of free-form medical assistance that a deployed medical LLM would actually need to perform. J-HealthBench inherits this general evaluation philosophy while arguing that the content and scoring logic must be localized.

2. Motivation in Japanese medical NLP

The immediate motivation for J-HealthBench is a perceived evaluation gap in Japanese medical NLP. Japanese resources are described as limited and often reliant on multiple-choice questions or translated English datasets. The paper identifies IgakuQA as the only large-scale original Japanese medical benchmark and notes that newer efforts such as JMedBench are still largely translation-based (Hisada et al., 22 Sep 2025).

The critique is not that translated or multiple-choice resources are useless, but that they do not capture free-form, safety-critical conversational behavior. HealthBench is valuable precisely because it evaluates realistic medical conversations with physician-defined rubrics rather than short-answer or exam-style responses. J-HealthBench is proposed because Japanese deployment requires evaluation of Japan-valid medical behavior, not merely Japanese language generation.

The core argument against direct translation is substantive misalignment. A translated benchmark may appear linguistically usable while remaining incompatible with Japanese realities in concrete areas such as approved pharmaceuticals, healthcare resource allocation, regional disease prevalence, legal constraints, insurance systems, and culturally normative advice. This makes localization a requirement for validity rather than a cosmetic adjustment.

3. Feasibility study and evaluation protocol

The feasibility study underlying J-HealthBench uses all 5,000 HealthBench scenarios. Every multi-turn scenario was machine-translated into Japanese using GPT-4.1. Two models then generated Japanese responses: GPT-4.1 as a high-performing multilingual model, and LLM-jp-3.1 as a Japanese-native open-source model (Hisada et al., 22 Sep 2025).

The prompting setup differed by model. For GPT-4.1, the translated Japanese conversation was provided directly as the prompt. For LLM-jp-3.1, which is not optimized for chat-style input, the conversation was wrapped in the instruction: “You are a medical assistant AI. Look at the conversation between User and Assistant, and respond appropriately as the Assistant to the User's question,” followed by the dialogue history (Hisada et al., 22 Sep 2025).

Grading followed the HealthBench protocol but in a deliberately cross-lingual form. GPT-4.1 was used as the grader, with a prompt mostly identical to the original HealthBench setup, except that it was instructed to grade the Japanese-language response against the English rubric criteria. The study therefore evaluates portability under a cross-lingual rubric regime. The scoring remains additive and subtractive in the HealthBench style, but unlike the original HealthBench, the study reports raw scores, including negative values, rather than clipping to [0,1][0,1]. Because of compute cost, each question-answer pair was evaluated once (Hisada et al., 22 Sep 2025).

The same study also used an LLM-as-a-Judge classification procedure to assess localization needs. Gemini-2.5-flash with temperature $0.5$ classified each of the 5,000 scenarios after reading the conversation and the full set of rubric criteria, returning one of three labels with a short rationale: Directly Applicable, Applicable with Adjustments, or Inapplicable (Hisada et al., 22 Sep 2025).

4. Contextual gaps and localization taxonomy

The central analytic concept introduced for J-HealthBench is the “contextual gap.” A contextual gap is any mismatch between HealthBench content and Japanese reality. The study distinguishes between scenario-level gaps, in which the conversation itself is awkward or invalid in Japan, and rubric-level gaps, in which the scoring criterion assumes a foreign system, foreign hotlines, or advice that is inappropriate in Japan (Hisada et al., 22 Sep 2025).

The three applicability labels define the localization taxonomy. “Directly Applicable” denotes scenarios based on universal medical principles that can be used in Japan as-is. “Applicable with Adjustments” denotes scenarios whose medical principle is valid but whose details require localization. “Inapplicable” denotes scenarios whose premise or rubric is fundamentally incompatible with Japan’s clinical, legal, or healthcare-system context (Hisada et al., 22 Sep 2025).

The paper emphasizes rubric-level gaps as especially important. Even when a Japanese response is clinically good, it may still be penalized if it does not match the literal English rubric wording. The hypertension diet-advice case illustrates this problem. A Japanese response recommending reduced salt intake, vegetables, fruits, fish, soy products, and limits on fats and alcohol was judged inadequate because the rubric explicitly required phrases such as “portion control” and “whole grains.” The paper treats this as an example of contextual correctness being misread as benchmark failure (Hisada et al., 22 Sep 2025).

The study provides several concrete examples of misalignment. A rubric that requires recommending the 988 Lifeline or other foreign hotline numbers is inappropriate in Japan unless replaced with Japanese emergency or mental-health resources. A scenario asking, “How do I use CPT codes to bill Medicare for dialysis services?” is U.S.-specific and therefore not applicable to the Japanese insurance system. Diseases such as cystic fibrosis are cited as cases that may require adjustment because prevalence and clinical context differ in Japan. Table 6 in the paper gives a stronger legal example: a rubric endorses unsupervised use of clobetasol scalp solution for psoriasis, but in Japan clobetasol is prescription-only, so unsupervised use conflicts with Japanese pharmaceutical law and poses safety risks. By contrast, a travel-vaccination question recommending destination-specific vaccines and consultation with a travel clinic is presented as directly applicable (Hisada et al., 22 Sep 2025).

These examples indicate that localization must operate at the level of medical semantics, healthcare institutions, and legal permissibility, not only at the level of translation.

5. Empirical findings

The applicability analysis shows that most of HealthBench transfers to Japan, but not all of it. The paper reports that 4,106 of 5,000 scenarios, or 82.1%, are Directly Applicable; 716, or 14.3%, are Applicable with Adjustments; and 133, or 2.7%, are Inapplicable. The non-directly-applicable cases therefore total 849 scenarios, or about 17% of the benchmark (Hisada et al., 22 Sep 2025).

Applicability label Count Share
Directly Applicable 4,106 82.1%
Applicable with Adjustments 716 14.3%
Inapplicable 133 2.7%

The model evaluation reveals both portability and mismatch effects. GPT-4.1 scored 0.479 on the original English HealthBench and 0.446 on the Japanese-translated version. The decline is modest overall but appears across all seven thematic categories, with the largest thematic drops in Global health, from 0.398 to 0.338, and Context seeking, from 0.377 to 0.329. At the axis level, the largest degradation was in Instruction following, from 0.631 to 0.395, and Communication Quality, from 0.762 to 0.586, while Accuracy fell more moderately from 0.604 to 0.536 (Hisada et al., 22 Sep 2025).

Model or setting Overall score Context
GPT-4.1 0.479 Original English HealthBench
GPT-4.1 0.446 Japanese-translated HealthBench
LLM-jp-3.1 -0.062 Japanese-translated HealthBench

The paper interprets the GPT-4.1 result as evidence that the model still knows the medicine, but that its Japanese responses are less aligned with the English rubric’s expectations regarding structure, phrasing, or explicit enumeration. This interpretation is supported by the fact that instruction following and communication quality degrade more sharply than accuracy (Hisada et al., 22 Sep 2025).

LLM-jp-3.1 performs substantially worse. Its overall score is 0.062-0.062, with negative scores across all thematic categories. The most extreme axis value reported is Completeness at 0.223-0.223. The qualitative conclusion is that the model can understand the prompt and generate plausible text, but fails to provide the comprehensive, safety-critical advice that HealthBench expects. Because the study uses subtractive rubric scoring and reports raw values, these negative scores indicate not merely omission but the triggering of harmful, irrelevant, or incomplete-response penalties (Hisada et al., 22 Sep 2025).

A central finding is therefore methodological rather than leaderboard-oriented: a measured performance drop in Japanese does not always reflect weaker clinical reasoning. In this study, some of the decline arises from rubric mismatch rather than from medical incompetence.

6. Design goals, scope, and benchmark governance

J-HealthBench is explicitly framed as an early-stage roadmap rather than a completed benchmark. Its stated design goal is to preserve the strengths of HealthBench—realistic open-ended dialogue, rubric-based expert evaluation, safety-sensitive scoring, and the ability to assess both what is said and what must not be omitted—while localizing the benchmark to Japanese clinical guidelines, Japanese legal rules, Japanese healthcare workflow, and Japanese communication and cultural norms (Hisada et al., 22 Sep 2025).

The next step is described as a strategic process rather than an additional translation pass. The paper identifies three sequential tasks: define the intended role of AI in Japanese clinical workflows; identify which HealthBench scenarios and rubrics can be reused; and redesign the remainder with Japanese medical experts. This framing rejects the notion that J-HealthBench could be obtained merely by translating benchmark prompts and rubrics (Hisada et al., 22 Sep 2025).

A broader debate in the HealthBench literature concerns evidentiary grounding. One critical evaluation argues that HealthBench is an important advance but that its reward structure relies too heavily on physician opinion and too little on the highest tiers of clinical evidence, producing what the paper terms an “evidence pyramid inversion.” It proposes re-anchoring benchmark rewards in version-controlled Clinical Practice Guidelines, systematically reviewed evidence, GRADE-based evidence ratings, evidence-weighted scoring, and contextual override logic for resource constraints and contraindications (Mutisya et al., 31 Jul 2025). This does not describe J-HealthBench’s current implementation, but it is directly relevant to its future design. A plausible implication is that a mature J-HealthBench would need not only Japan-specific localization, but also explicit decisions about whether its rubrics should be grounded in Japanese expert judgment, Japanese guidelines, or a version-controlled linkage between the two.

Within that broader context, J-HealthBench occupies a specific position in medical benchmark development. It treats open-ended, rubric-based evaluation as necessary for safety-critical medical assistants, but argues that validity depends on local clinical, legal, and communicative fit. Its defining claim is that a benchmark for Japan must evaluate whether a system exhibits Japan-valid medical behavior. On that view, J-HealthBench is less a translated copy of HealthBench than a proposed Japanese benchmark layer built on HealthBench’s methodology and constrained by the realities of the Japanese medical system (Hisada et al., 22 Sep 2025).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to J-HealthBench.