---
title: 'Kizuki: Language-Aware Accessibility Tool'
url: https://www.emergentmind.com/topics/kizuki
type: topic
---

# Kizuki: Language-Aware Accessibility Tool

Searching arXiv for the specified papers and closely related context.
Kizuki is a **language-aware automated accessibility testing extension** proposed in “Not All Visitors are Bilingual: A Measurement Study of the Multilingual Web from an Accessibility Perspective” [2508.18328]. Named after the Japanese word for “awareness,” it is described both as a testing extension that “identifies such mismatches and evaluates metadata based on alignment with the surrounding linguistic context” and, more concretely, as “a Lighthouse extension that incorporates language awareness in accessibility evaluation.” Its central purpose is to expose a blind spot in mainstream automated web accessibility auditing: existing tools typically verify whether accessibility metadata exists, but not whether that metadata is written in the appropriate language for the visible page context. In the study, Kizuki does not generate accessibility text or repair pages automatically; rather, it operationalizes the finding that multilingual and non-Latin-script websites frequently contain accessibility hints that are structurally present yet linguistically misaligned with the interface being presented to users.

## 1. Conceptual origin and problem formulation

Kizuki emerges from a broader argument about multilingual web accessibility. English remains predominant on the web, but many websites increasingly combine English with regional or native languages in visible content and hidden metadata. For visually impaired users, accessibility depends not only on visible page text but also on hidden accessibility hints such as `alt` text, ARIA labels, form labels, titles, and language metadata. These are the strings that screen readers consume to pronounce, interpret, and navigate content. When they are missing, empty, generic, untranslated, or written in a different language from the surrounding interface, the assistive experience diverges from the visible one [2508.18328].

The paper situates this problem in the technical limitations of assistive technology. It cites prior work showing that JAWS and NVDA “still exhibit limited support for non-Latin scripts and often perform poorly when confronted with mixed languages,” and notes that VoiceOver does not support some languages at all. In that setting, language-inconsistent accessibility hints are not merely a translation defect. They can produce mispronunciations, unintelligible speech, broken navigation cues, or cognitively taxing context switching. The paper characterizes the resulting experience as an unintended bilingual interaction: the interface may be visually native-language, while the accessibility layer remains in English.

Kizuki is therefore best understood as a diagnostic extension for a specific failure mode: the presence of accessibility metadata whose language does not match the dominant visible language of the page. This makes its scope narrower than general accessibility auditing, but more precise with respect to multilingual accessibility defects that presence-only audits systematically overlook.

## 2. Empirical foundation in LangCrUX

The empirical basis for Kizuki is LangCrUX, introduced in the same paper as “the first large-scale dataset of 120,000 popular websites across 12 languages that primarily use non-Latin scripts” [2508.18328]. LangCrUX is constructed from Chrome User Experience Report rankings plus country-specific crawling and language verification. Its coverage spans Mandarin Chinese, Hindi, Modern Standard Arabic, Bangla, Russian, Japanese, Egyptian Arabic, Cantonese, Korean, Thai, Greek, and Hebrew. The language detection machinery relies on script-specific character ranges and additional language-specific characters in overlapping-script cases.

The study uses LangCrUX to establish two distinct patterns. First, accessibility hints are often poor even under conventional criteria: many elements are missing labels, and many existing labels are empty or uninformative. Second, and more specifically relevant to Kizuki, informative accessibility text often fails to reflect the language of visible content. The paper reports that in Bangladesh, **79% of informative accessibility texts are in English**; mixed-language labels are common in Greece, Thailand, and Hong Kong; and in India and Bangladesh, **over 40% of websites have less than 10% of their accessibility text in the native language despite predominantly native-language visible content**.

Several case examples make the mismatch concrete. The Bangladeshi education portal `teachers.gov.bd` has over **98% Bangla visible content**, yet only **one of 79 images with alt text uses Bangla**. The Hindi site `cmhelpline.mp.gov.in` presents an interface almost entirely in Hindi but retains all accessibility text in English. The Thai site `khaosod.co.th` is over **92% Thai** in visible content while accessibility labels are mostly English, and the Chinese site `kjt.shaanxi.gov.cn` is almost fully Chinese but its accessibility texts are entirely English. These examples motivate Kizuki directly: they show why a structural “metadata exists” pass condition can substantially overestimate effective accessibility on multilingual sites.

## 3. System architecture and audit scope

Architecturally, Kizuki is not a standalone browser or replacement auditing framework. It is an extension to **Google Lighthouse**, which itself relies on **axe-core** for many accessibility checks [2508.18328]. The authors first identified **language-sensitive accessibility features** by inspecting Lighthouse accessibility tests and the HTML elements and axe-core rules they target. They selected twelve rules whose outcomes depend on natural-language quality:

- `button-name`
- `document-title`
- `image-alt`
- `frame-title`
- `summary-name`
- `label`
- `input-image-alt`
- `select-name`
- `link-name`
- `input-button-name`
- `svg-img-alt`
- `object-alt`

This twelve-rule set defines Kizuki’s broader conceptual scope: accessibility elements for which human-readable text is central to the user experience. However, the implemented extension evaluated in the paper focuses specifically on **image alt text**. The paper states: “Specifically, we extend the audit for image alt text to verify whether the description is written in the same language as the page’s visible content.”

The input to Kizuki is a rendered webpage in a Chromium/Lighthouse environment. From that page, the extension requires at least two classes of signals: the **visible textual content** of the page and the **accessibility metadata** associated with the audited elements, most concretely `alt` attributes for images in the evaluated prototype. The output is a modified accessibility assessment in which a page is not rewarded solely because `alt` text exists; the text must also be linguistically aligned with the surrounding interface.

Implementation-wise, Kizuki sits in a **Puppeteer**-driven **Chromium** auditing stack with Lighthouse as the underlying framework. The appendix also states that both LangCrUX and Kizuki are open-sourced and that the repository explains “how to use it and how to extend it with custom accessibility tests.” This indicates a modular implementation strategy layered atop existing browser-based auditing rather than a new conformance framework.

## 4. Detection logic, heuristics, and operative thresholds

The paper does not provide source code, pseudocode, or a formal algorithm block for Kizuki, but its detection logic is described in enough detail to reconstruct its operative workflow [2508.18328]. The core language inference mechanism is a **Unicode-based heuristic**. Visible text is matched against script-specific character ranges, with explicit mention of Devanagari, Hangul, and Cyrillic, and with additional language-specific characters used in overlapping-script cases such as Arabic and Urdu. For dataset selection, a website is retained if at least **50%** of its visible textual content is in the target language. That dominant-visible-language assumption underlies the later consistency check.

Before language alignment is evaluated, the paper applies a rule-based filtering pipeline to remove accessibility text that is present but not meaningfully informative. The filtering rules discard:

- emoji
- too-short strings
- file names
- URLs or file paths
- generic actions such as “search” or “close”
- placeholders such as “image,” “icon,” or “button”
- developer labels such as `btn-submit` or `nav_menu`
- label-number patterns such as “image 1”
- single-word entries in non-CJK scripts unless descriptive
- mixed alphanumeric identifiers such as `img123`
- ordinal phrases such as “2 of 10”

The appendix specifies language-dependent shortness thresholds: for CJK scripts, the limit is **1 character**; for others, it is **3 characters**. Although the paper does not explicitly state that Kizuki runs this exact filtering stage before scoring, the language-aware analysis on which the extension is based uses these heuristics to distinguish informative from boilerplate accessibility text. The effective logic is therefore: extract visible text and accessibility text, discard obviously uninformative strings, infer script or language on both sides, and flag inconsistency when accessibility metadata is written in a different language from the page’s dominant visible language.

The paper is explicit about why this matters for automated auditing. In isolated test pages constructed for the appendix, the authors examined each of the twelve language-sensitive rules under three conditions: missing element, empty value, and incorrect language. The critical finding is that **incorrect language passes Lighthouse**. Kizuki’s contribution is thus the introduction of a new failure mode—language inconsistency—into an audit rule that previously checked only structural presence.

## 5. Evaluation and quantitative impact on Lighthouse outcomes

The evaluation of Kizuki is deliberately narrow but empirically informative. The authors test the extension on **10,000 websites from Bangladesh and Thailand**, chosen because “language mismatch between visible content and accessibility metadata is particularly common” there [2508.18328]. For fairness, they exclude websites that fail the original Lighthouse test due to missing alt attributes. The comparison therefore isolates the added effect of language awareness rather than conflating it with already-detected absence of accessibility metadata.

The principal observable metric is the overall Lighthouse accessibility score, with scores above **90** interpreted as “good.” Under ordinary Lighthouse, **43%** of the evaluated websites receive an accessibility score above **90**, and **5.6%** achieve a perfect score. After Kizuki’s language-aware `image-alt` check is applied, only **15.8%** remain above 90, and only **1.8%** retain a perfect score. The paper uses this shift to demonstrate that many websites currently classified as highly accessible are reclassified downward once linguistic consistency is treated as part of accessibility rather than as an external localization concern.

This result is not presented as a formal scoring theorem or as a new weighted objective. No mathematical formula for score adjustment is given. Instead, the evaluation is impact-oriented: it shows that a large fraction of pages passing the original `image-alt` audit fail to preserve high accessibility scores once mismatched-language alt text is treated as defective. In methodological terms, the experiment shows that presence-only auditing yields false positives with respect to multilingual accessibility.

The paper does not provide an ablation study, classifier benchmark, or detailed detector error analysis for Kizuki itself. Nor does it include a user study demonstrating that blind users prefer pages Kizuki scores more highly. The empirical claim is therefore specific: language-aware auditing materially changes outcomes at scale, especially on localized websites whose visible and accessibility layers have drifted apart.

## 6. Scope, limitations, and significance

Kizuki addresses **language alignment**, not **semantic adequacy** [2508.18328]. An `alt` text can be in the correct language and still be vague, misleading, overlong, or otherwise unusable. The paper repeatedly separates these issues and treats uninformative metadata through filtering rather than through semantic understanding. The current evaluated prototype also extends only the **`image-alt`** audit, even though the study identifies twelve language-sensitive accessibility elements.

The reliance on script-based heuristics introduces additional limitations. True multilingual pages, transliteration, borrowed English words in native scripts, and intentional code-switching can all complicate script-based language inference. Kizuki also assumes that the visible page language is the appropriate reference language for accessibility text, which may not always hold for bilingual audiences, branded content, proper nouns, or pedagogical interfaces intentionally designed for mixed-language use. Moreover, the tool cannot solve underlying screen-reader support limitations for unsupported languages; it can only reveal metadata mismatches that exacerbate those limitations.

Within those constraints, Kizuki is significant because it operationalizes a blind spot that mainstream auditing frameworks leave unmodeled. It complements WCAG-oriented checks, especially those connected to language and text alternatives, but it does not claim full multilingual conformance checking under WCAG 3.1.2. Its contribution is narrower and more concrete: it demonstrates that accessibility metadata should be audited as part of the page’s linguistic interface, not merely as a structural requirement.

The paper implies several natural extensions. One is coverage expansion from `image-alt` to other language-sensitive elements such as `aria-label`, `label`, `button-name`, `link-name`, `document-title`, and `frame-title`. Another is richer language identification beyond script-level heuristics for mixed-script and fine-grained code-switching cases. A further implication is that future versions would need to combine **language consistency** with **semantic quality assessment**, since localization alone does not guarantee descriptive accessibility text. In that sense, Kizuki’s enduring importance lies less in complete problem closure than in making visible a specific failure mode of the multilingual web: accessibility metadata can be present, standards-compatible at a structural level, and still be operationally misaligned with the interface that users are actually navigating.

Source: https://www.emergentmind.com/topics/kizuki