Attribute keyboard-detection differences across assessment configurations

Determine how much of the difference in recovery of keyboard-related accessibility findings between interactive worker agents and noninteractive vision-language models is attributable to browser interaction, instructions, tools, and the evidence extracted by each procedure.

Background

The evaluation reports that worker agents recover more keyboard-related reference cases than the noninteractive vision-LLM configurations, but the tested configurations differ in several respects at once. In particular, they vary in interaction capabilities, instructions, available tools, and the evidence supplied to the model.

Because these factors are confounded in the archival comparison, the observed performance difference cannot be assigned specifically to browser interaction. The authors explicitly leave unresolved how much each factor contributes, making controlled experimentation with separately varied components necessary.

References

The results indicate a particular weakness on the keyboard cases in this archive, while leaving open how much of the difference is attributable to interaction, instructions, tools, and the evidence each procedure extracts.

Agentic Web Accessibility Auditing: Authoring and Evaluating Per-Criterion Worker Agents for WCAG  (2609.09379 - Mishra et al., 8 Sep 2026) in Section 5.2, subsection “Criterion-level variation and abstention”