---
title: 'LangCrUX: Multilingual Accessibility Dataset'
url: https://www.emergentmind.com/topics/langcrux
type: topic
---

# LangCrUX: Multilingual Accessibility Dataset

LangCrUX is a large-scale research dataset of multilingual websites designed to study web accessibility from a language and script perspective. It is built on top of the Chrome User Experience Report (CrUX) and augments popularity-ranked sites with detailed language and accessibility annotations, with a specific focus on websites whose visible content is predominantly in non-Latin scripts. Its central purpose is to make multilingual accessibility barriers empirically measurable at web scale, especially the mismatch between visible interface language and the language of accessibility hints used by screen readers and related assistive technologies [2508.18328].

## 1. Origin, definition, and research problem

LangCrUX, short for “Language Chrome UX,” was introduced to address a gap left by existing large-scale web datasets. CrUX and Tranco provide information about popularity and performance, but they do not encode language composition or accessibility metadata. As a result, they are not suitable for systematically measuring how multilingual content is exposed to screen readers, for quantifying mismatches between visible text and accessibility hints such as `alt` or `aria-label`, or for studying accessibility for non-Latin-script languages at scale [2508.18328].

The dataset is motivated by two concurrent facts. English remains predominant on the web, while large user populations rely primarily on other languages, many of them written in non-Latin scripts. At the same time, screen readers and related assistive technologies often have limited support for such scripts and can perform poorly on mixed-language pages, producing mispronunciation, unreadable output, or missing support. LangCrUX was created to make these problems observable in a structured, quantitative form.

The dataset’s scope is explicitly accessibility-centric rather than merely multilingual. It does not only record the visible language of webpages; it also records accessibility-relevant metadata and the language composition of that metadata. This design permits direct comparison between what sighted users read and what screen-reader users are likely to hear.

## 2. Dataset scope and composition

LangCrUX covers 120,000 websites, with 10,000 websites for each of 12 language-country pairs. All selected sites have measurable Chrome traffic and are “popular” in the sense of user engagement and performance metrics. The dataset focuses on languages that primarily use non-Latin scripts and collectively account for approximately 3.19 billion speakers, or about 39.5% of the world’s population [2508.18328].

| Language-country pair | Language | Approximate speaker population |
|---|---|---:|
| China | Mandarin Chinese | ~1.2 billion |
| India | Hindi | ~609 million |
| Algeria | Modern Standard Arabic | ~335 million |
| Bangladesh | Bangla | ~284 million |
| Russia | Russian | ~253 million |
| Japan | Japanese | ~126 million |
| Egypt | Egyptian Arabic | ~119 million |
| Hong Kong | Cantonese | ~85.5 million |
| South Korea | Korean | ~82 million |
| Thailand | Thai | ~71 million |
| Greece | Greek | ~13.5 million |
| Israel | Hebrew | ~9 million |

The script coverage includes Indic scripts such as Devanagari and Bengali, Arabic-based scripts, East Asian scripts, Thai, Greek, Hebrew, and Cyrillic. Languages such as Tamil, Telugu, Sinhala, and Georgian were considered initially but excluded if they did not meet the minimum site count. This indicates that LangCrUX is not intended as a typologically exhaustive multilingual corpus; it is a popularity-weighted, script-focused dataset constrained by observable web presence.

Website selection begins with CrUX. For each language-country pair, the top 10,000 domains are extracted for the country and ranked by user-experience metrics. A language-based filter is then applied to the visible text of each page using Unicode ranges for the relevant script. A site is included only if at least 50% of its visible textual content is in the target language script. If a high-ranked site fails this threshold, it is replaced by a lower-ranked candidate until the quota is reached. For languages sharing scripts, such as Arabic and Urdu, the selection heuristics are refined with language-specific characters to reduce ambiguity.

## 3. Data collection, language detection, and accessibility annotation

LangCrUX is collected with a Puppeteer-based crawler running in a Chromium environment. For each selected domain, the crawler loads the page, extracts DOM content, and records accessibility-relevant attributes. Traffic is routed through country-specific VPN servers, specifically ProtonVPN and Hotspot Shield, so that sites are more likely to serve localized content rather than English or generic global variants. Domains that detect VPNs and return restricted content are replaced with other eligible sites [2508.18328].

The core language-detection mechanism is Unicode-based script detection over visible text. For a page \(p\) and language \(L\), the visible-language share is defined conceptually as

$$
\text{share}_{\text{visible}}(p, L) =
\frac{\# \text{ of visible characters in script of } L}
{\# \text{ of visible textual characters on } p}.
$$

A site is included if \(\text{share}_{\text{visible}}(p, L) \geq 0.5\). The same mechanism is adapted to accessibility text:

$$
\text{share}_{\text{access}}(p, L) =
\frac{\# \text{ of characters in accessibility attributes in script of } L}
{\# \text{ of characters in accessibility attributes on } p}.
$$

This enables page-level comparison between the dominant language of visible content and the dominant language of accessibility hints.

For each site, LangCrUX records visible textual content, accessibility hints and metadata, DOM structure, and network/environment metadata. The language-sensitive accessibility elements listed in the paper are `button-name`, `document-title`, `image-alt`, `frame-title`, `summary-name`, `label`, `input-image-alt`, `select-name`, `link-name`, `input-button-name`, `svg-img-alt`, and `object-alt`.

A central methodological step is the filtering of uninformative accessibility text. LangCrUX discards emoji, overly short strings, file names, URLs and paths, generic actions such as “search” or “close” when standalone, placeholders such as “image,” “icon,” or “button” and their translations, developer labels such as `btn-submit` or `nav_menu`, label-number patterns such as “image 1,” mixed alphanumeric IDs such as `img123`, ordinal phrases such as “2 of 10,” and obviously generic single-word labels in non-CJK settings. If \(T\) is the set of accessibility texts on page \(p\), and \(\mathcal{F}\) is the rule set, the informative subset is

$$
T_{\text{informative}}(p) = \{ t \in T \mid t \text{ does not match any rule in } \mathcal{F} \}.
$$

This filtering matters because raw attribute presence is not treated as sufficient evidence of useful accessibility support.

## 4. Empirical findings on multilingual accessibility

LangCrUX is primarily a measurement instrument, and its main results concern the prevalence, quality, and language consistency of accessibility metadata. Across the twelve accessibility element types, missing rates are often extremely high: `label` is 98.55% missing on average, `svg-img-alt` 96.66%, `link-name` 95.96%, `input-button-name` 93.90%, and `summary-name` 90.47%. Among image-related attributes, `image-alt` is 17.12% missing on average, but 25.39% are present and empty. The paper also reports outliers, including `image-alt` values up to 261,864 characters and 12,306 words, indicating cases where entire paragraphs or metadata were dumped into alt text rather than concise descriptions [2508.18328].

Filtering reveals that much existing accessibility text is not semantically useful. Single-word labels account for more than 33% of accessibility texts in Thailand, 22.2% in Russia, 18.03% in Greece, and 17.1% in India. Too-short texts remain common, and URLs or file paths appear as labels in Hong Kong, South Korea, and Russia. This suggests that the nominal presence of accessibility attributes substantially overstates their practical value.

After filtering, the language distribution of informative accessibility text shows strong English dominance even on sites whose visible interface is primarily in a local language. In Bangladesh, 79% of informative accessibility texts are in English. Egypt, Thailand, and Greece also show strong English dominance. Mixed-language labels are also common: 35% in Greece, 34% in Thailand, 30% in Hong Kong, and more than 20% in China, Russia, Japan, and India. Because screen readers generally do not switch voices mid-label, mixed-language labels are likely to be mispronounced or confusing.

The most consequential result is the visible-versus-accessibility language mismatch. Many sites have predominantly native-language visible content but very little native-language accessibility text. In India and Bangladesh, over 40% of websites have less than 10% of their informative accessibility text in the native language despite visible user interfaces being mostly native. In Thailand, China, and Hong Kong, more than a quarter of sites fall into similar low-native-accessibility ranges. Japan and Israel show better alignment, with fewer than 10% of sites exhibiting severe mismatch.

The paper provides concrete examples. Bangladesh’s government portal `https://teachers.gov.bd` has more than 98% of visible content in Bangla, yet only 1 of 79 images has Bangla alt text. India’s `https://cmhelpline.mp.gov.in` presents an interface almost entirely in Hindi, but all accessibility text is in English. Thai news site `https://www.khaosod.co.th` has more than 92% visible text in Thai while accessibility labels are mostly English. Chinese provincial government site `https://kjt.shaanxi.gov.cn` has visible content nearly fully in Chinese but accessibility text entirely in English.

A common misconception addressed by these findings is that accessibility can be approximated by checking whether attributes such as `alt` merely exist. LangCrUX shows that presence, emptiness, informativeness, and language consistency are distinct properties.

## 5. Kizuki and language-aware accessibility auditing

LangCrUX also serves as the empirical basis for Kizuki, a language-aware automated accessibility testing extension for Google Lighthouse. Kizuki is motivated by the observation that Lighthouse and Axe-core check the presence of accessibility attributes but do not assess whether their language matches the page’s visible content. Empty attributes and incorrect-language cases can therefore pass existing audits [2508.18328].

Kizuki currently focuses on image alt text. For each page, it computes the native-language share of visible content, filters uninformative alt texts using the same discard rules as LangCrUX, detects the language of the remaining alt content, and classifies it as native-only, English-only, mixed, or other. A page passes the Kizuki check if images appearing in native-language contexts have alt text in the native language, or at least not systematically English-only, and if the alt text is neither empty nor placeholder-like. The language-consistency condition is expressed conceptually as:

$$
\forall \text{img}_i \in p,\;
\text{if } \text{contextLang}(\text{img}_i) = L_{\text{native}},
\text{ then } \text{lang}(\text{alt}_i) \approx L_{\text{native}}.
$$

The evaluation is conducted on 10,000 websites from Bangladesh and Thailand, chosen because they exhibit high mismatch rates between visible native-language content and English accessibility text. The experiment considers only sites that already pass Lighthouse’s standard alt test. Before Kizuki, 43% of websites scored above 90 and 5.6% received a perfect score of 100. After Kizuki’s language-aware alt-text check was applied, only 15.8% scored above 90 and only 1.8% retained a perfect score. The result is not that these sites became less accessible in any intrinsic sense; rather, it shows that conventional tools overestimate accessibility by ignoring language consistency.

## 6. Interpretation, limitations, and significance

LangCrUX is not a direct screen-reader user study. The paper does not run live user tests with JAWS, NVDA, VoiceOver, or other assistive technologies. Instead, it combines large-scale web measurement with previously documented screen-reader problems such as mispronunciation, misrendering, incomplete language support, and difficulties with mixed-language output. This means that LangCrUX quantifies structural conditions that are likely to degrade assistive experience, rather than measuring comprehension or task success directly [2508.18328].

The authors note several limitations. LangCrUX depends on CrUX and therefore only includes sites with measurable Chrome traffic, excluding low-traffic or highly localized sites that may still matter for marginalized language communities. VPN-based crawling can provoke generic or restricted site responses. Puppeteer-based browsing does not fully reflect dynamic, personalized, or authenticated interfaces. Script-based language detection can miss embedded content such as images containing text and can struggle with overlapping scripts or unusual orthography. The study also omits some accessibility features, most notably video captions, because caption files are often loaded dynamically and are difficult to evaluate at scale.

Despite these limits, the dataset has broad practical significance. It establishes multilingual accessibility as a measurable web-scale phenomenon rather than an anecdotal concern. It supplies structured evidence that accessibility metadata is often missing, empty, boilerplate, or linguistically inconsistent with the visible interface. It also provides an openly available resource for tool evaluation and policy work, including the dataset at `https://anonymous.4open.science/r/LangCrux-F68F/` and an interactive exploration site at `https://anonymous.4open.science/w/LangCrux-F68F/`.

A plausible implication is that multilingual accessibility requires a shift from attribute-presence auditing to language-aware auditing. LangCrUX shows that users who rely on screen readers may encounter a systematically more bilingual interface than sighted users on the same site, especially on websites whose visible content is in Bangla, Hindi, Thai, Chinese, and other non-Latin-script languages. In that sense, LangCrUX functions both as a dataset and as a measurement framework for a specific accessibility failure mode: language-inconsistent accessibility hints on the multilingual web.

Source: https://www.emergentmind.com/topics/langcrux