Identify the mechanism underlying stable gaze–AOI divergence under combined guidance

Determine whether stable gaze–AOI divergence during combined Co-Annotator guidance occurs because the ontology-bounded VLM text satisfies residents’ information needs or because residents rely less on the gaze-aligned AOI overlay when both modalities are presented.

Background

In US3, gaze–AOI divergence remained stable across the pre-guidance, combined-guidance, and post-guidance blocks, unlike the increase observed during and after AOI-only guidance in US2. The authors propose two competing explanations: textual biomarker guidance may reduce the need for spatial reorientation, or residents may primarily engage with the VLM-generated text and use the AOI overlay less.

Survey results and qualitative reports favor the first interpretation, but they do not resolve the issue. The authors explicitly state that direct measurement of engagement with each modality is required to distinguish the explanations.

References

Two interpretations are consistent with this pattern: the VLM text may have satisfied residents' information needs, reducing pressure to spatially re-orient gaze; or residents may have relied less on the AOI overlay in the combined condition, engaging primarily with the textual draft. Survey data showing moderate perceived reliance (3.75/7) and qualitative reports of a corroboration loop between modalities favor the first interpretation, though direct measurement of per-modality engagement is needed to distinguish the two.

Co-Annotator: Expert-Distilled ViT and VLM for Visual and Documentation Guidance in Age-Related Macular Degeneration  (2608.30352 - Li et al., 31 Aug 2026) in Section 3, User Study 3, subsection “Diagnostic Accuracy and Efficiency”