ColorBlindnessEval: VLM Robustness Benchmark
- ColorBlindnessEval is a benchmark that adapts Ishihara test logic to assess adversarial color-pattern recognition in Vision-Language Models.
- It employs a Monte Carlo-based image generation process to embed numbers within complex dot patterns, testing low-level visual discrimination.
- Empirical results reveal that VLMs suffer from hallucination and response bias, especially in open-ended digit recognition under adversarial settings.
ColorBlindnessEval is a benchmark designed to evaluate the robustness of Vision-LLMs (VLMs) in visually adversarial scenarios inspired by the Ishihara color blindness test. Rather than diagnosing human color vision deficiency, it adapts the pseudo-isochromatic logic of Ishihara plates to AI evaluation: numbers from 0 to 99 are embedded in complex color-dot patterns that are easy for humans with normal color vision and difficult for models, exposing hallucination, brittle perception, and weak low-level visual discrimination in multimodal systems (Ling et al., 23 Sep 2025).
1. Definition and conceptual scope
ColorBlindnessEval formalizes a specific failure mode in multimodal perception: the answer is present in the image, but it is deliberately embedded in complex color patterns instead of conventional object layouts. In this setting, a VLM may miss the answer, default to a stereotyped response pattern, or hallucinate a number that is not present. The benchmark therefore targets visually adversarial perception rather than generic OCR or ordinary image captioning (Ling et al., 23 Sep 2025).
A recurrent clarification in the literature is that Ishihara-inspired evaluation of VLMs is not equivalent to clinical testing of human observers. One closely related study uses 25 out of 38 Ishihara numeral plates to test whether large vision-LLMs can answer as if they were under normal vision, protanopia, deuteranopia, or tritanopia; it reports that these models can explain color vision deficiencies in natural language, but they cannot simulate how people with CVDs perceive color in image based tasks (Hayashi et al., 23 May 2025). Taken together, these works distinguish three separate questions: whether a model can discuss CVD, whether it can withstand adversarial color-pattern structure, and whether it can condition its answers on altered human perceptual states.
This distinction also guards against a common misconception. ColorBlindnessEval does not claim that VLMs are clinically “color blind” in the human sense. Its operative concern is robustness under Ishihara-like figure-ground encoding. A plausible implication is that the benchmark measures a conjunction of low-level color discrimination, grouping of scattered dots into a coherent symbol, and resistance to prompt-induced response bias.
2. Dataset construction and image-generation procedure
The dataset contains 500 Ishihara-like image pairs. The embedded numbers range from 0 to 99. Generation proceeds in three stages. First, a reference image is created by rendering a number in black on a white background, using Arial by default. Second, a modified Monte Carlo-based procedure generates the plate geometry by placing many non-overlapping circles of varying size at random positions in multiple layers. Third, colors are assigned to circles according to whether each circle’s center falls in the foreground numeral region or the background region of the reference image (Ling et al., 23 Sep 2025).
The image-generation algorithm is explicit. A plate is initialized as empty. For each layer , the method sets a number of circles and a size range . Each proposed circle samples a random center within the unit circle and a random radius ; the circle is accepted only if it does not overlap existing circles. After the layout is fixed, each circle is mapped back to reference-image coordinates . If , the circle receives a random color from the foreground set ; otherwise it receives a random color from the background set (Ling et al., 23 Sep 2025).
Each item has two image types: a standard Ishihara-like image and a foreground-only clear version. For each color set the authors create 100 image pairs, giving 500 pairs across five color sets. The clear version removes the adversarial background and serves as a control condition for separating basic digit recognition from failure on the full pseudo-isochromatic dot field. The five color sets were sampled from original Ishihara numeral plates and represented as RGB values, and the paper reports that VLMs generally perform best on Color Set 1 and worst on Color Set 3, suggesting that stronger foreground-background contrast materially helps model performance (Ling et al., 23 Sep 2025).
What makes the benchmark adversarial is therefore not merely unusual coloration, but the way numerical structure is distributed across randomized dots. The number is not outlined or segmented in the manner preferred by ordinary OCR pipelines. It emerges only from color-contrast relations across the dot field, which is why the benchmark is informative about grouping and color-structured perception rather than only symbol recognition.
3. Prompting protocol, scoring, and empirical findings
ColorBlindnessEval evaluates nine VLMs under Yes/No and open-ended prompting, with the clear-image control reported separately. Accuracy is exact-match accuracy:
0
with indicator function
1
The benchmark’s task modes are as follows (Ling et al., 23 Sep 2025):
| Mode | Intended behavior |
|---|---|
| 2 | Answer “yes” to the true queried number |
| 3 | Answer “no” to a random incorrect number |
| Open | Output the number in the standard Ishihara-like image |
| Open-clear | Output the number in the foreground-only clear image |
The paper notes a documentation inconsistency: the appendix prompt box swaps the labels for the Yes/No and open-ended prompts relative to the setup section, but the experimental categories and numerical results are unambiguous. This matters because the headline result is not any single prompt score, but the pattern across prompt types.
That pattern is stark. All models are dramatically better on Open-clear than on Open. GPT-4o scores 4 on Open-clear but only 5 on Open, and GPT-4o-mini scores 6 on Open-clear but only 7 on Open. Claude3-Haiku collapses from 8 on Open-clear to 9 on Open. This shows that the adversarial background, rather than digit recognition per se, is the dominant source of failure (Ling et al., 23 Sep 2025).
Yes/No prompting also reveals strong response biases. Claude3-Haiku shows 0 and 1, indicating near-universal rejection of the queried number. Qwen2-VL-Instruct-2B behaves similarly with 2 versus 3. By contrast, Qwen2-VL-Instruct-72B shows a strong “yes” bias with 4 but 5. The paper therefore argues that balanced performance across both Yes/No directions is a better sign of competence than a single high value, and by that criterion GPT-4o-mini is the most balanced tested model at 6 on both (Ling et al., 23 Sep 2025).
Open-ended recognition remains especially weak. GPT-4o-mini is best at 7, followed by GPT-4o at 8; all other reported models are below 9, many near zero. Human comparison on a calibration subset is correspondingly revealing. Twenty participants without color blindness answered 10 Open questions on standard Ishihara-like images and 10 Open-clear questions on the corresponding clear images. Their mean performance was 0 on Open-C and 1 on OpenClear-C, indicating only a modest drop for humans where models often collapse almost completely (Ling et al., 23 Sep 2025).
The paper also reports two ablations with direct benchmark significance. First, few-shot prompting with same-color demonstrations hurt Qwen2-VL models rather than helping them, suggesting that naïve in-context examples can amplify bias instead of improving perceptual extraction. Second, replacing Arial with DejaVuSans changed absolute accuracy while preserving the general trend that clear images are easier than standard images, indicating sensitivity not only to color patterning but also to typography-induced dot layout (Ling et al., 23 Sep 2025).
4. Position within the broader evaluation landscape
ColorBlindnessEval sits within a broader turn toward color-centered evaluation of multimodal systems, but it occupies a narrower and harder niche than most adjacent benchmarks. “ColorFoil” defines model “color blindness” functionally as failure to distinguish correct from incorrect color descriptions of an image, using caption foils derived from MS COCO and Flickr30k. Its task is zero-shot image-text ranking over original and color-substituted captions, and it shows that BridgeTower and ViLT substantially outperform CLIP-family models on color grounding (Samin et al., 2024). ColorBlindnessEval differs by shifting the problem from caption-level grounding to adversarial figure-ground extraction.
“ColorBench” generalizes the scope further. It contains 1,448 instances and 5,814 image-text questions across 11 tasks spanning color perception, color reasoning, and color robustness, including a dedicated “Color Blindness” task with 157 Ishihara-style items and a robustness track built from recolored variants of seed images (Liang et al., 10 Apr 2025). This suggests that ColorBlindnessEval can be read as a stress-test specialization within a larger color-understanding taxonomy: it concentrates benchmark capacity on a single difficult regime rather than distributing it across many task families.
A third nearby line of work uses human clinical methods more directly. “Diagnosing Vision LLMs' Perception by Leveraging Human Methods for Color Vision Deficiencies” evaluates six LVLMs on 25 Ishihara plates under prompts such as “You are Protanopic” or “You are Deuteranopic,” together with Linguistic Support and Visual Support variants. That study finds that models often possess declarative knowledge of CVD yet fail at image-grounded simulation of altered perception, and it further notes that Ishihara is good for red-green CVD but poor for Tritanopia (Hayashi et al., 23 May 2025). A plausible implication is that ColorBlindnessEval and the attributed-perception benchmark are complementary: the former probes adversarial visual robustness, while the latter probes whether a model can condition its answers on a specified perceptual regime.
5. Human-centered and assistive lineages
The benchmark emerged against a longer background of assistive imaging, adaptive interfaces, and user-centered accessibility research. An early proof-of-concept paper developed freeware software with an HTML interface to simulate several forms of color blindness on a loaded color image, show a rainbow preview, and apply a very simple pseudo-correction for protanopia or protanomaly by replacing red information with grayscale/luminance. It also proposed iPad, cellphone, and Google Glass-like viewing pipelines, but provided no standardized benchmark or quantitative evaluation framework (Oliveira et al., 2015). In that lineage, evaluation was primarily demonstrative and application-driven.
A later adaptive system combined weighted Ishihara responses, fuzzy severity estimation, LMS-space simulation, and four correction methods. Its most concrete reported comparison is that Method B with histogram equalization obtained the best results for about 47% of volunteers in a web-based survey. This work is notable because it links diagnosis, simulation, and compensation, but its evaluation remains subjective and largely tied to preference for improved distinguishability rather than standardized task accuracy (Lee et al., 2017).
More recent assistive work shifts toward recognition rather than static correction. “Computational Trichromacy Reconstruction” proposes an AR smartphone interface in which users interactively rotate colors in linear sRGB about the gray axis, observe induced color shifts, and learn to resolve color-name confusions over time. Through psychophysical experiments and a 9-day longitudinal study, the paper reports threshold reductions and recall scores between 18.25 and 19.125 out of 20 during recall days, showing that dynamic, learnable assistance can be evaluated in terms of discrimination, recognition, and retention rather than only immediate visibility (Zhu et al., 2024).
There is also a parallel literature on interface-level accessibility assessment. A simulation-based study of 20 popular websites asks non-CVD observers to compare original screenshots with simulated protanomaly, deuteranomaly, and tritanomaly views, rating retained functionality and aesthetics. It reports a positive correlation between functionality and aesthetics (2) and finds that an operating-system-wide high contrast mode can reduce both. This work is significant because it shows that color accessibility evaluation can target production UIs and mitigation strategies, not just abstract images or assistive transforms (Jamil et al., 2024).
Taken together, these studies suggest that ColorBlindnessEval extends a longer accessibility lineage into the evaluation of multimodal AI. Instead of measuring whether a user can be helped by recoloring or AR interaction, it measures whether a model can recover information from the kind of color-structured visual signal that human accessibility practice has long treated as nontrivial.
6. Limitations, interpretation, and future directions
ColorBlindnessEval is intentionally synthetic. Its benchmark images are generated rather than photographed, the main dataset uses one font by default, and human evaluation covers only a 20-image calibration subset with 20 participants without color blindness. The paper itself notes a prompt-label inconsistency in the appendix and shows that a font change to DejaVuSans can materially alter absolute performance even while preserving the overall trend (Ling et al., 23 Sep 2025). These are not fatal defects, but they define the scope of inference.
A second interpretive limitation is that Ishihara-like tasks privilege red-green figure-ground confusions. Related work explicitly warns that Ishihara is not suitable for evaluating Tritanopia, and that strong verbal knowledge about protanopia or deuteranopia does not imply grounded multimodal modeling of those perceptual states (Hayashi et al., 23 May 2025). This suggests that ColorBlindnessEval should not be treated as a complete test of color understanding or of human-like color-deficiency simulation.
A third limitation is task breadth. The benchmark focuses specifically on numerical recognition in pseudo-isochromatic dot fields rather than broader semantic perception, scene understanding, or color reasoning. Broader suites already show that VLMs can fail on color extraction, color counting, color proportion, color illusion, and color robustness under recoloring, with performance gaps across models that are often modest in absolute terms (Liang et al., 10 Apr 2025). A plausible implication is that a fuller ColorBlindnessEval family would combine Ishihara-like adversarial plates with complementary tasks for color grounding, reasoning, robustness, and attributed perceptual variation.
Even with these limits, the benchmark has clear diagnostic value. It shows that strong performance on ordinary multimodal tasks does not guarantee reliable extraction of information encoded in color-pattern structure. It also reveals that apparently high Yes/No accuracy may reflect prompt-conditioned bias rather than grounded perception, and that open-ended recognition on adversarial plates remains far below human performance. In that sense, ColorBlindnessEval is best understood as a focused robustness benchmark: narrow in stimulus family, but unusually effective at exposing hallucination, suggestibility, and weak low-level visual discrimination in current VLMs (Ling et al., 23 Sep 2025).