Diagnosing Failures on Logo and Printed-Text Matching

Determine whether VHop-Router's poor performance on logo- and printed-text-based retrieval is primarily caused by difficulty reading visual marks, applying a matching rule absent from training, or controlling the resulting multi-step search.

Background

The VHop benchmark extends its object-identity matching task to logo matching (L6) and printed-text matching (L7). VHop-Router is trained only on L4 identity links, so these evaluations test transfer to matching rules that were not present during training. Performance on both extensions is very low, particularly for printed-text matching.

The authors explicitly leave unresolved which component is responsible for these failures: visual reading of the marks, transfer to a new matching rule, or search control after the new clue is identified. They note that broader retrieval does not resolve the issue and that targeted training together with separate perception checks would be needed to distinguish the alternatives.

References

The results do not establish whether the main cause is reading the marks, applying the new matching rule, or controlling the resulting search. Wider retrieval improves some L6/L7 agent results in Table~\ref{tab:topk-full}, but it does not resolve these questions.

— Learning to Route in Visual Space via Multi-Step Embedding Retrieval  (2609.38743 - Chen et al., 30 Sep 2026) in Appendix, Section 'Generalization and Limitations', subsection 'Extensions to Logo and Text Matching'