Papers
Topics
Authors
Recent
Search
2000 character limit reached

Color Vision Testing: Methods & Benchmarks

Updated 6 July 2026
  • Color Vision Testing is a heterogeneous set of assessments that measure chromatic sensitivity, semantic interpretation, and real-world color applications.
  • Methods range from traditional Ishihara plates and CAD tests to advanced FM-100 paradigms and AI-driven benchmarks, each isolating distinct aspects of color perception.
  • Recent studies emphasize integrated approaches using AR interfaces and observer models to quantify low-level discrimination, contextual reasoning, and robustness under color perturbations.

Color vision testing denotes a heterogeneous class of tasks for assessing how color information is detected, discriminated, interpreted, or used for action. In contemporary research, the concept extends beyond Ishihara-style screening and hue-ordering exams to include chromatic-threshold measurement with dynamic luminance masking, pseudoisochromatic simulation across severity continua, context-aware everyday tasks such as traffic-light interpretation or judging the doneness of meat, interface accessibility audits under simulated color vision deficiency, and multimodal AI benchmarks that isolate color perception, reasoning, and robustness (Natali et al., 2022, Morita et al., 2024, Liang et al., 10 Apr 2025).

1. Task scope and taxonomy

A useful way to organize color vision testing is by the level of function under examination. Some tasks probe low-level chromatic sensitivity or confusion structure, as in the CAD test or FM-100. Others probe whether color-coded meaning can be inferred in realistic scenes, whether an interface remains usable under simulated color vision deficiency, or whether a vision-LLM can recover information encoded primarily by color (Natali et al., 2022, Sun et al., 14 Feb 2025, Jamil et al., 2024, Ye et al., 15 Jul 2025).

Task family Canonical formulation Representative source
Screening and thresholding Ishihara plates, CAD red/green and yellow/blue thresholds, FM-100 hue ordering (Qaiser et al., 2017, Natali et al., 2022, Sun et al., 14 Feb 2025)
Contextual functional tasks Traffic lights, meat doneness, ripe fruit, clothing, transit signs (Morita et al., 2024)
Interactive augmentation tasks 4AFC oddity under rotational color shifts; longitudinal color naming (Zhu et al., 2024)
AI and multimodal benchmarks ColorBench, Ishihara-like LVLM tests, VQA skill prediction (Liang et al., 10 Apr 2025, Ye et al., 15 Jul 2025, Zeng et al., 2020)
Interface and simulation audits Simulated UI comparison, observer-specific compensation, metamer discrimination (Jamil et al., 2024, Oshima et al., 2015, Gundlach et al., 2017)

This taxonomy matters because different task families operationalize different constructs. A plate-reading task may emphasize pseudoisochromatic segregation; the CAD test isolates chromatic sensitivity from luminance cues; a contextual AR task evaluates whether color meaning can be translated into action; and a robustness benchmark evaluates whether a model changes its answer when image color is perturbed despite the answer being color-independent (Natali et al., 2022, Morita et al., 2024, Liang et al., 10 Apr 2025).

A recurrent misconception in the literature is that color vision testing is exhausted by isolated hue naming or static patch discrimination. Multiple recent works argue otherwise. One line of work reframes testing around semantic interpretation in real situations, while another shows that model performance on general multimodal benchmarks does not guarantee competence on controlled color tasks (Morita et al., 2024, Ye et al., 15 Jul 2025).

2. Clinical and psychophysical paradigms

The most conventional paradigm in the surveyed literature is Ishihara-based screening. Ishihara is described as a standard, widely used screening test, regularly used to screen for congenital and acquired red green inadequacies (Qaiser et al., 2017). In the adaptive interface workflow demonstrated in the LifeLine smartphone application, the user first passes through an Ishihara test; the system then determines whether the user is color blind and, if so, selects a more suitable color scheme. In a study of 100 undergraduate students aged 18 to 24, 4 out of 100 were found to be color blind: 3 were deuteranopes and 1 was protanopic; all color-blind participants reportedly found the resulting application color scheme satisfactory and visible (Qaiser et al., 2017). At the same time, that paper omits the number of plates, the scoring rubric, subtype decision rules, and any digital display calibration procedure, which limits reproducibility.

Threshold-based psychophysics is represented by the CAD test. The CAD test presents colored stimuli on a calibrated monitor against a background of dynamic luminance contrast noise, with the noise masking luminance contrast signals without affecting significantly either RG or YB chromatic sensitivity (Natali et al., 2022). In the tinted-lens study, 10 young adults aged 19–26 years were tested under three viewing conditions: no lens, a slightly tinted blue-blocking filter, and a heavily tinted orange filter. The blue-blocking filter did not significantly affect either RG or YB colour vision, whereas the orange filter caused large changes in colour discrimination, especially for YB thresholds. For YB thresholds, one-way ANOVA gave F(2,27)=65.24, p<105F(2,27)=65.24,\ p<10^{-5}, and Tukey HSD showed no-lens versus blue-blocker p=0.98p=0.98, but both no-lens versus orange and blue-blocker versus orange were p<105p<10^{-5} (Natali et al., 2022). This establishes an important testing constraint: lens condition can be a confound in chromatic-threshold assessment.

A complementary paradigm is hue ordering and pseudoisochromatic simulation. The physiologically based red–green CVD model built on the CIE 2006 physiological observer model was validated with pseudoisochromatic plates and the Farnsworth-Munsell 100 Hue test (Sun et al., 14 Feb 2025). FM-100 was treated in its standard 85-cap, four-row form, with Total Error Score used as the principal index. The same paper also used four groups from the Waggoner Computerized Color Vision Test: two vanishing plate groups, one protan classification group, and one deutan classification group. Recognition-rate curves were fit against severity, and the 50% recognition point was treated as the diagnosis threshold. The proposed model yielded vanishing-plate thresholds of 13.9 nm and 13.8 nm for protan, and 15.3 nm and 16.6 nm for deutan across the two vanishing sets; more importantly, it reproduced the expected protan/deutan selectivity on classification plates, whereas the other four compared methods did not (Sun et al., 14 Feb 2025).

These paradigms are not interchangeable. Ishihara functions as a screening mechanism, CAD isolates channel-specific chromatic sensitivity, and FM-100 or computerized pseudoisochromatic simulations probe confusion structure and severity. This suggests that a comprehensive human testing battery should not collapse all color-vision questions into a single task class.

3. Ecological and context-aware task formulations

A major recent shift is from perceptual compensation to semantic assistance. In the AR- and LLM-based support system for color vision deficiency, the central claim is that many everyday problems are not reducible to distinguishing red from green in isolation; rather, the task is to infer what color means in context (Morita et al., 2024). The system combines an augmented reality interface with a multimodal LLM-based reasoner. Input can be explicit speech, such as “Please tell me the color of the traffic light,” or a minimal button press that asks the model to infer intent from visual context. The reasoning pipeline is structured as a four-step Chain-of-Thought prompt: analyze the current environmental situation; determine the user’s intent and the type of help required; generate concise supportive content limited to 10 words; and identify key terms for emphasis presentation (Morita et al., 2024).

The evaluation scenarios are explicitly practical: identifying traffic lights, judging the doneness of meat, selecting ripe fruits, coordinating clothing, and reading color-coded signs in public transportation (Morita et al., 2024). In a preliminary user study with two actual color vision deficient participants, each scenario was tested in two different environments for ten tests total. The paper reports that in all ten tests the multimodal LLM correctly recognized the context and user intentions and produced accurate assistance, with an average practical effectiveness rating of 8.5 out of 10 (Morita et al., 2024). The same paper also identifies failure modes highly relevant to task design: multi-object ambiguity, especially with multiple pieces of meat on a tray; uncertainty in LLM correctness; and the need to calibrate trust in safety-critical tasks.

A different ecological formulation appears in computational trichromacy reconstruction. That work treats color naming, rather than mere discrimination, as the central difficulty for dichromats and some anomalous trichromats (Zhu et al., 2024). The proposed augmentation is a rotation about the gray axis in linear sRGB space, implemented on a smartphone AR interface; users induce color shifts by swiping and then learn the transformation-dependent signatures of otherwise confusable colors. Psychophysically, the system was evaluated with a 4AFC oddity task in 16 CVD participants: 2 protanopes, 2 protanomalous, 3 deuteranopes, and 8 deuteranomalous participants. Threshold reductions with shifts were statistically significant for green, gray, and blue in deuteranomaly (p<0.05p<0.05) and more significant for blue, green, and gray in deuteranopia (p<0.01p<0.01) (Zhu et al., 2024).

The same work extends testing into learning and retention. In a 9-day longitudinal study with 8 participants, the average number of attempts to reach perfect 20/20 naming performance dropped from 4.375 on Day 1 to 1.625 on Day 3, with Day 1 versus Day 3 significant at p<0.05p<0.05. During recall on Days 4–8, average scores ranged from 18.25 to 19.125, and a ZZ-test against μ0=10\mu_0=10 showed performance significantly above chance on each recall day (p<0.01p<0.01) (Zhu et al., 2024). This suggests that some color-vision testing tasks can be formulated not only as deficit measurement but also as measurement of interactive recoverability and learnability.

4. Machine benchmarks and multimodal color testing

In machine vision and VQA, color is often embedded within broader task structure. A VQA-oriented analysis of 27,263 visual questions defines a four-skill taxonomy—object recognition, text recognition, color recognition, and counting—and shows that color recognition is needed in 22.06% of real-user VizWiz questions and 16.7% of VQA2.0 questions (Zeng et al., 2020). Color rarely appears in isolation; it commonly co-occurs with object recognition, especially in the object+color combination. The same study uses answer entropy,

E=i=1Npilogpi,E=\sum_{i=1}^{N}-p_i \log p_i,

as a human-difficulty proxy and finds that color questions are harder for humans in VizWiz than in VQA2.0, with Welch’s p=0.98p=0.980 (Zeng et al., 2020). For testing, this implies that color-dependent evaluation often measures grounded interpretation rather than pure hue perception.

ColorBench turns that insight into a dedicated VLM benchmark. It contains 1,448 instances and 5,814 image-text questions across 11 tasks spanning three capability areas: color perception, color reasoning, and color robustness (Liang et al., 10 Apr 2025). The robustness protocol is stringent: for a seed image p=0.98p=0.981 and question p=0.98p=0.982, robustness is 1 only if the model is correct on the original and on all recolored variants. Across 32 VLMs, the best overall perception+reasoning scores were 57.8 for Gemini-2-flash with CoT and 56.2 for GPT-4o with CoT; average CoT gains were +3.65 for perception+reasoning and +13.3 for robustness (Liang et al., 10 Apr 2025). The benchmark also shows that color can mislead models: grayscale ablations improve performance on Color Illusion and Color Mimicry for many systems.

A more plate-specific benchmark is the LVLM color vision test based on synthetic Ishihara-like images. That dataset contains 5,450 entries across Numbers, Animals, Letters or Chinese characters, Objects, and Shapes, with two prompt settings: Color Vision Test Easy, which provides category cues, and Color Vision Test Hard, which does not (Ye et al., 15 Jul 2025). Zero-shot performance is strikingly low. On CVTE, the best average score is 20.86 for JanusPro-7B and 20.28 for GPT-4o; on CVTH, GPT-4o reaches 18.39 and JanusPro-7B 17.76 (Ye et al., 15 Jul 2025). The paper identifies five error modes: Incorrect Category Understanding (10.14%), Unidentifiable (21.36%), Complete Recognition Error (37.53%), Partial Recognition Error (18.98%), and Stochastic Fallback on Uncertainty (11.99%). Yet targeted LoRA fine-tuning of LLaVA1.5-7B raises performance from 15.72 to 94.43 on CVTE and from 11.31 to 92.23 on CVTH (Ye et al., 15 Jul 2025).

An adjacent question is whether LVLMs can simulate altered human perception rather than merely solve color-encoded tasks. An Ishihara-based study of LVLMs finds that models can explain CVDs in natural language, but cannot simulate how people with CVDs perceive color in image-based tasks (Hayashi et al., 23 May 2025). Under the Base prompt for normal vision, GPT-4o scores 90.5% and mPLUG-Owl3 71.4% on Ishihara numeral reading. But performance collapses for protanopia and deuteranopia, with no model exceeding 24%, and linguistic or visual support does not rescue the failure (Hayashi et al., 23 May 2025). This suggests a sharp dissociation between factual knowledge about CVD and perceptually grounded simulation of CVD.

5. Observer models, simulation frameworks, and interface-centered testing

Some color-vision testing tasks are built from explicit models of observer geometry. In the Riemannian framework for color-weak vision, an observer’s color space is defined by local JND ellipsoids and metric tensor p=0.98p=0.983, with local line element

p=0.98p=0.984

Distances are geodesic, and compensation or simulation maps are constructed as isometries between color-normal and color-weak spaces (Oshima et al., 2015). Measurements were taken at 77 sampling points across p=0.98p=0.985, and the color-normal observer’s discrimination ellipsoids were on average 2.6948 times larger in volume than those of the color-weak observer; the average 2D ellipse area ratio was 1.6193 (Oshima et al., 2015). This does not define a clinical test by itself, but it provides a principled way to generate observer-specific stimuli, quantify perceptual separability, and construct compensation maps for accessibility-oriented testing.

Simulation also underlies interface-level testing. A screenshot-based protocol for UI accessibility under simulated CVD used the physiologically based Machado model via Colorspacious on 20 popular applications, 3 screenshots each, and both standard and high contrast modes, for 120 stimuli total (Jamil et al., 2024). Nineteen non-CVD participants rated functionality on a 1–5 scale and aesthetics as kept or lost. Aggregate non-high-contrast functionality was p=0.98p=0.986, versus p=0.98p=0.987 for high contrast; aggregate aesthetics-kept probability was p=0.98p=0.988 versus p=0.98p=0.989, both with p<105p<10^{-5}0, and the Pearson correlation between functionality and aesthetics was p<105p<10^{-5}1 (Jamil et al., 2024). The proposed AAA–A scheme labels UIs AAA if functionality p<105p<10^{-5}2 and aesthetics p<105p<10^{-5}3, AA if functionality p<105p<10^{-5}4 and aesthetics p<105p<10^{-5}5, and A when core functionality remains but aesthetics and functionality are reduced (Jamil et al., 2024). This is a testing task aimed at production interfaces rather than perception alone.

Another strand targets spectral discrimination rather than deficiency diagnosis. A binocular filtering approach to “breaking binocular redundancy” uses different filters in each eye to reduce metamer prevalence (Gundlach et al., 2017). The authors report up to about 15× metamer reduction in Monte Carlo simulations and one to two orders of magnitude decrease in the more abstract volume calculation (Gundlach et al., 2017). For testing, the salient task becomes filtered versus unfiltered discrimination of spectrally distinct but color-matched stimuli, especially in the blue-violet region.

For dot-pattern stimuli, exhaustive target recovery can itself become the core testing problem. A color-complexity-based method for exhaustive color-dot identification quantizes RGB or HSV space, exploits the low fraction of occupied color cubes, and then builds connected target-color dots and multiscale MST-based spatial tests (Liao et al., 2020). The paper reports color-complexity examples of p<105p<10^{-5}6, p<105p<10^{-5}7, and p<105p<10^{-5}8 (Liao et al., 2020). A plausible implication is that automated scoring of pseudoisochromatic or colored-dot tasks can be decomposed into two distinct stages: chromatic recovery of target dots, and spatial verification that the recovered dots form a nonuniform symbol rather than random clutter.

6. Limitations, confounds, and directions of development

The literature repeatedly warns against overgeneralization. The AR-LLM support study is preliminary, with only two actual color vision deficient participants and ten scenario-environment combinations (Morita et al., 2024). The screenshot-based UI protocol is explicitly limited to simulated CVD screenshots rated by non-CVD observers, not lived perception by CVD users (Jamil et al., 2024). The Ishihara-based LVLM simulation study shows that language-level explanation of CVD does not imply image-level simulation of CVD (Hayashi et al., 23 May 2025). The synthetic LVLM plate benchmark shows that benchmark-specific fine-tuning can produce dramatic gains, but the paper itself implies that those gains may not demonstrate broad real-world generalization (Ye et al., 15 Jul 2025).

Another recurring issue is under-specification. The Ishihara-driven adaptive interface does not specify plate sequence, scoring rubric, subtype thresholds, or digital calibration (Qaiser et al., 2017). The AR-LLM contextual assistant does not provide a full prompt template, decoding settings, latency benchmarks, fallback rules, or spatial grounding algorithms (Morita et al., 2024). The CAD tinted-lens study gives strong statistics for YB thresholds, but remains a preliminary study with 10 young adults and only two lenses (Natali et al., 2022). These omissions matter because test validity in color tasks is highly sensitive to display, optics, illumination, and response protocol.

A final limitation is conceptual. Several works imply that color testing is multidimensional: discrimination, naming, grounded reasoning, and robustness can dissociate. VLMs may answer “what color is this?” while failing on color-conditioned counting or pseudoisochromatic segregation (Liang et al., 10 Apr 2025). Human-support systems may correctly state a traffic light’s color yet remain unsafe if trust calibration is poor or if multi-object ambiguity is unresolved (Morita et al., 2024). This suggests that a mature color vision testing task should span at least four layers: low-level chromatic sensitivity, contextual semantic interpretation, robustness to irrelevant color perturbation, and calibration under uncertainty.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Color Vision Testing Task.