---
title: 'Color Vision Testing: Methods & Benchmarks'
url: https://www.emergentmind.com/topics/color-vision-testing-task
type: topic
---

# Color Vision Testing: Methods & Benchmarks

Color vision testing denotes a heterogeneous class of tasks for assessing how color information is detected, discriminated, interpreted, or used for action. In contemporary research, the concept extends beyond Ishihara-style screening and hue-ordering exams to include chromatic-threshold measurement with dynamic luminance masking, pseudoisochromatic simulation across severity continua, context-aware everyday tasks such as traffic-light interpretation or judging the doneness of meat, interface accessibility audits under simulated color vision deficiency, and multimodal AI benchmarks that isolate color perception, reasoning, and robustness [2208.14211][2407.04362][2504.10514].

## 1. Task scope and taxonomy

A useful way to organize color vision testing is by the level of function under examination. Some tasks probe low-level chromatic sensitivity or confusion structure, as in the CAD test or FM-100. Others probe whether color-coded meaning can be inferred in realistic scenes, whether an interface remains usable under simulated color vision deficiency, or whether a vision-language model can recover information encoded primarily by color [2208.14211][2502.10316][2401.10357][2507.11153].

| Task family | Canonical formulation | Representative source |
|---|---|---|
| Screening and thresholding | Ishihara plates, CAD red/green and yellow/blue thresholds, FM-100 hue ordering | [1712.03329], [2208.14211], [2502.10316] |
| Contextual functional tasks | Traffic lights, meat doneness, ripe fruit, clothing, transit signs | [2407.04362] |
| Interactive augmentation tasks | 4AFC oddity under rotational color shifts; longitudinal color naming | [2408.01895] |
| AI and multimodal benchmarks | ColorBench, Ishihara-like LVLM tests, VQA skill prediction | [2504.10514], [2507.11153], [2010.03160] |
| Interface and simulation audits | Simulated UI comparison, observer-specific compensation, metamer discrimination | [2401.10357], [1510.06507], [1703.04392] |

This taxonomy matters because different task families operationalize different constructs. A plate-reading task may emphasize pseudoisochromatic segregation; the CAD test isolates chromatic sensitivity from luminance cues; a contextual AR task evaluates whether color meaning can be translated into action; and a robustness benchmark evaluates whether a model changes its answer when image color is perturbed despite the answer being color-independent [2208.14211][2407.04362][2504.10514].

A recurrent misconception in the literature is that color vision testing is exhausted by isolated hue naming or static patch discrimination. Multiple recent works argue otherwise. One line of work reframes testing around semantic interpretation in real situations, while another shows that model performance on general multimodal benchmarks does not guarantee competence on controlled color tasks [2407.04362][2507.11153].

## 2. Clinical and psychophysical paradigms

The most conventional paradigm in the surveyed literature is Ishihara-based screening. Ishihara is described as a standard, widely used screening test, regularly used to screen for congenital and acquired red green inadequacies [1712.03329]. In the adaptive interface workflow demonstrated in the LifeLine smartphone application, the user first passes through an Ishihara test; the system then determines whether the user is color blind and, if so, selects a more suitable color scheme. In a study of 100 undergraduate students aged 18 to 24, 4 out of 100 were found to be color blind: 3 were deuteranopes and 1 was protanopic; all color-blind participants reportedly found the resulting application color scheme satisfactory and visible [1712.03329]. At the same time, that paper omits the number of plates, the scoring rubric, subtype decision rules, and any digital display calibration procedure, which limits reproducibility.

Threshold-based psychophysics is represented by the CAD test. The CAD test presents colored stimuli on a calibrated monitor against a background of dynamic luminance contrast noise, with the noise masking luminance contrast signals without affecting significantly either RG or YB chromatic sensitivity [2208.14211]. In the tinted-lens study, 10 young adults aged 19–26 years were tested under three viewing conditions: no lens, a slightly tinted blue-blocking filter, and a heavily tinted orange filter. The blue-blocking filter did not significantly affect either RG or YB colour vision, whereas the orange filter caused large changes in colour discrimination, especially for YB thresholds. For YB thresholds, one-way ANOVA gave \(F(2,27)=65.24,\ p<10^{-5}\), and Tukey HSD showed no-lens versus blue-blocker \(p=0.98\), but both no-lens versus orange and blue-blocker versus orange were \(p<10^{-5}\) [2208.14211]. This establishes an important testing constraint: lens condition can be a confound in chromatic-threshold assessment.

A complementary paradigm is hue ordering and pseudoisochromatic simulation. The physiologically based red–green CVD model built on the CIE 2006 physiological observer model was validated with pseudoisochromatic plates and the Farnsworth-Munsell 100 Hue test [2502.10316]. FM-100 was treated in its standard 85-cap, four-row form, with Total Error Score used as the principal index. The same paper also used four groups from the Waggoner Computerized Color Vision Test: two vanishing plate groups, one protan classification group, and one deutan classification group. Recognition-rate curves were fit against severity, and the 50% recognition point was treated as the diagnosis threshold. The proposed model yielded vanishing-plate thresholds of 13.9 nm and 13.8 nm for protan, and 15.3 nm and 16.6 nm for deutan across the two vanishing sets; more importantly, it reproduced the expected protan/deutan selectivity on classification plates, whereas the other four compared methods did not [2502.10316].

These paradigms are not interchangeable. Ishihara functions as a screening mechanism, CAD isolates channel-specific chromatic sensitivity, and FM-100 or computerized pseudoisochromatic simulations probe confusion structure and severity. This suggests that a comprehensive human testing battery should not collapse all color-vision questions into a single task class.

## 3. Ecological and context-aware task formulations

A major recent shift is from perceptual compensation to semantic assistance. In the AR- and LLM-based support system for color vision deficiency, the central claim is that many everyday problems are not reducible to distinguishing red from green in isolation; rather, the task is to infer what color means in context [2407.04362]. The system combines an augmented reality interface with a multimodal LLM-based reasoner. Input can be explicit speech, such as “Please tell me the color of the traffic light,” or a minimal button press that asks the model to infer intent from visual context. The reasoning pipeline is structured as a four-step Chain-of-Thought prompt: analyze the current environmental situation; determine the user’s intent and the type of help required; generate concise supportive content limited to 10 words; and identify key terms for emphasis presentation [2407.04362].

The evaluation scenarios are explicitly practical: identifying traffic lights, judging the doneness of meat, selecting ripe fruits, coordinating clothing, and reading color-coded signs in public transportation [2407.04362]. In a preliminary user study with two actual color vision deficient participants, each scenario was tested in two different environments for ten tests total. The paper reports that in all ten tests the multimodal LLM correctly recognized the context and user intentions and produced accurate assistance, with an average practical effectiveness rating of 8.5 out of 10 [2407.04362]. The same paper also identifies failure modes highly relevant to task design: multi-object ambiguity, especially with multiple pieces of meat on a tray; uncertainty in LLM correctness; and the need to calibrate trust in safety-critical tasks.

A different ecological formulation appears in computational trichromacy reconstruction. That work treats color naming, rather than mere discrimination, as the central difficulty for dichromats and some anomalous trichromats [2408.01895]. The proposed augmentation is a rotation about the gray axis in linear sRGB space, implemented on a smartphone AR interface; users induce color shifts by swiping and then learn the transformation-dependent signatures of otherwise confusable colors. Psychophysically, the system was evaluated with a 4AFC oddity task in 16 CVD participants: 2 protanopes, 2 protanomalous, 3 deuteranopes, and 8 deuteranomalous participants. Threshold reductions with shifts were statistically significant for green, gray, and blue in deuteranomaly (\(p<0.05\)) and more significant for blue, green, and gray in deuteranopia (\(p<0.01\)) [2408.01895].

The same work extends testing into learning and retention. In a 9-day longitudinal study with 8 participants, the average number of attempts to reach perfect 20/20 naming performance dropped from 4.375 on Day 1 to 1.625 on Day 3, with Day 1 versus Day 3 significant at \(p<0.05\). During recall on Days 4–8, average scores ranged from 18.25 to 19.125, and a \(Z\)-test against \(\mu_0=10\) showed performance significantly above chance on each recall day (\(p<0.01\)) [2408.01895]. This suggests that some color-vision testing tasks can be formulated not only as deficit measurement but also as measurement of interactive recoverability and learnability.

## 4. Machine benchmarks and multimodal color testing

In machine vision and VQA, color is often embedded within broader task structure. A VQA-oriented analysis of 27,263 visual questions defines a four-skill taxonomy—object recognition, text recognition, color recognition, and counting—and shows that color recognition is needed in 22.06% of real-user VizWiz questions and 16.7% of VQA2.0 questions [2010.03160]. Color rarely appears in isolation; it commonly co-occurs with object recognition, especially in the object+color combination. The same study uses answer entropy,
\[
E=\sum_{i=1}^{N}-p_i \log p_i,
\]
as a human-difficulty proxy and finds that color questions are harder for humans in VizWiz than in VQA2.0, with Welch’s \(t=-3.30,\ p<.001\) [2010.03160]. For testing, this implies that color-dependent evaluation often measures grounded interpretation rather than pure hue perception.

ColorBench turns that insight into a dedicated VLM benchmark. It contains 1,448 instances and 5,814 image-text questions across 11 tasks spanning three capability areas: color perception, color reasoning, and color robustness [2504.10514]. The robustness protocol is stringent: for a seed image \(I_s\) and question \(q\), robustness is 1 only if the model is correct on the original and on all recolored variants. Across 32 VLMs, the best overall perception+reasoning scores were 57.8 for Gemini-2-flash with CoT and 56.2 for GPT-4o with CoT; average CoT gains were +3.65 for perception+reasoning and +13.3 for robustness [2504.10514]. The benchmark also shows that color can mislead models: grayscale ablations improve performance on Color Illusion and Color Mimicry for many systems.

A more plate-specific benchmark is the LVLM color vision test based on synthetic Ishihara-like images. That dataset contains 5,450 entries across Numbers, Animals, Letters or Chinese characters, Objects, and Shapes, with two prompt settings: Color Vision Test Easy, which provides category cues, and Color Vision Test Hard, which does not [2507.11153]. Zero-shot performance is strikingly low. On CVTE, the best average score is 20.86 for JanusPro-7B and 20.28 for GPT-4o; on CVTH, GPT-4o reaches 18.39 and JanusPro-7B 17.76 [2507.11153]. The paper identifies five error modes: Incorrect Category Understanding (10.14%), Unidentifiable (21.36%), Complete Recognition Error (37.53%), Partial Recognition Error (18.98%), and Stochastic Fallback on Uncertainty (11.99%). Yet targeted LoRA fine-tuning of LLaVA1.5-7B raises performance from 15.72 to 94.43 on CVTE and from 11.31 to 92.23 on CVTH [2507.11153].

An adjacent question is whether LVLMs can simulate altered human perception rather than merely solve color-encoded tasks. An Ishihara-based study of LVLMs finds that models can explain CVDs in natural language, but cannot simulate how people with CVDs perceive color in image-based tasks [2505.17461]. Under the Base prompt for normal vision, GPT-4o scores 90.5% and mPLUG-Owl3 71.4% on Ishihara numeral reading. But performance collapses for protanopia and deuteranopia, with no model exceeding 24%, and linguistic or visual support does not rescue the failure [2505.17461]. This suggests a sharp dissociation between factual knowledge about CVD and perceptually grounded simulation of CVD.

## 5. Observer models, simulation frameworks, and interface-centered testing

Some color-vision testing tasks are built from explicit models of observer geometry. In the Riemannian framework for color-weak vision, an observer’s color space is defined by local JND ellipsoids and metric tensor \(G(x)\), with local line element
\[
ds^2 = dx^T G(x)\,dx.
\]
Distances are geodesic, and compensation or simulation maps are constructed as isometries between color-normal and color-weak spaces [1510.06507]. Measurements were taken at 77 sampling points across \(L^*=30,40,50,60,70\), and the color-normal observer’s discrimination ellipsoids were on average 2.6948 times larger in volume than those of the color-weak observer; the average 2D ellipse area ratio was 1.6193 [1510.06507]. This does not define a clinical test by itself, but it provides a principled way to generate observer-specific stimuli, quantify perceptual separability, and construct compensation maps for accessibility-oriented testing.

Simulation also underlies interface-level testing. A screenshot-based protocol for UI accessibility under simulated CVD used the physiologically based Machado model via Colorspacious on 20 popular applications, 3 screenshots each, and both standard and high contrast modes, for 120 stimuli total [2401.10357]. Nineteen non-CVD participants rated functionality on a 1–5 scale and aesthetics as kept or lost. Aggregate non-high-contrast functionality was \(\mu_{(d)}=4.09\), versus \(\mu_{(e)}=3.66\) for high contrast; aggregate aesthetics-kept probability was \(s'_{(d)}=0.66\) versus \(s'_{(e)}=0.41\), both with \(p \ll 0.05\), and the Pearson correlation between functionality and aesthetics was \(r=0.74\) [2401.10357]. The proposed AAA–A scheme labels UIs AAA if functionality \(>4\) and aesthetics \(>0.75\), AA if functionality \(>3.8\) and aesthetics \(>0.5\), and A when core functionality remains but aesthetics and functionality are reduced [2401.10357]. This is a testing task aimed at production interfaces rather than perception alone.

Another strand targets spectral discrimination rather than deficiency diagnosis. A binocular filtering approach to “breaking binocular redundancy” uses different filters in each eye to reduce metamer prevalence [1703.04392]. The authors report up to about 15× metamer reduction in Monte Carlo simulations and one to two orders of magnitude decrease in the more abstract volume calculation [1703.04392]. For testing, the salient task becomes filtered versus unfiltered discrimination of spectrally distinct but color-matched stimuli, especially in the blue-violet region.

For dot-pattern stimuli, exhaustive target recovery can itself become the core testing problem. A color-complexity-based method for exhaustive color-dot identification quantizes RGB or HSV space, exploits the low fraction of occupied color cubes, and then builds connected target-color dots and multiscale MST-based spatial tests [2007.14485]. The paper reports color-complexity examples of \(\frac{393014}{256^3}\approx 0.023\), \(\frac{28126}{256^3}\approx 0.002\), and \(\frac{880}{26^3}\approx 0.05\) [2007.14485]. A plausible implication is that automated scoring of pseudoisochromatic or colored-dot tasks can be decomposed into two distinct stages: chromatic recovery of target dots, and spatial verification that the recovered dots form a nonuniform symbol rather than random clutter.

## 6. Limitations, confounds, and directions of development

The literature repeatedly warns against overgeneralization. The AR-LLM support study is preliminary, with only two actual color vision deficient participants and ten scenario-environment combinations [2407.04362]. The screenshot-based UI protocol is explicitly limited to simulated CVD screenshots rated by non-CVD observers, not lived perception by CVD users [2401.10357]. The Ishihara-based LVLM simulation study shows that language-level explanation of CVD does not imply image-level simulation of CVD [2505.17461]. The synthetic LVLM plate benchmark shows that benchmark-specific fine-tuning can produce dramatic gains, but the paper itself implies that those gains may not demonstrate broad real-world generalization [2507.11153].

Another recurring issue is under-specification. The Ishihara-driven adaptive interface does not specify plate sequence, scoring rubric, subtype thresholds, or digital calibration [1712.03329]. The AR-LLM contextual assistant does not provide a full prompt template, decoding settings, latency benchmarks, fallback rules, or spatial grounding algorithms [2407.04362]. The CAD tinted-lens study gives strong statistics for YB thresholds, but remains a preliminary study with 10 young adults and only two lenses [2208.14211]. These omissions matter because test validity in color tasks is highly sensitive to display, optics, illumination, and response protocol.

A final limitation is conceptual. Several works imply that color testing is multidimensional: discrimination, naming, grounded reasoning, and robustness can dissociate. VLMs may answer “what color is this?” while failing on color-conditioned counting or pseudoisochromatic segregation [2504.10514]. Human-support systems may correctly state a traffic light’s color yet remain unsafe if trust calibration is poor or if multi-object ambiguity is unresolved [2407.04362]. This suggests that a mature color vision testing task should span at least four layers: low-level chromatic sensitivity, contextual semantic interpretation, robustness to irrelevant color perturbation, and calibration under uncertainty.

Source: https://www.emergentmind.com/topics/color-vision-testing-task