Evaluate FATA across broader tasks

Investigate the behavior and effectiveness of Feature-Aware Token Attack (FATA) across a broader range of vision-language tasks beyond TextVQA, VQAv2, and ScienceQA benchmarks.

Background

The reported experiments use four visually dependent benchmark subsets: TextVQA-Open, VQAv2-Open, ScienceQA-MC, and VQAv2-MC. The authors explicitly state that broader task coverage remains unresolved, leaving open whether the observed trade-off between full-token preservation and compression-triggered failure generalizes to other task types and evaluation settings.

References

Black-box transfer, broader tasks, and stronger online detectors remain open.

— Feature-Aware Token Attack for Compression-Triggered Stealthy Failures in Large Vision-Language Models  (2609.39134 - Yan et al., 30 Sep 2026) in Section 6, Limitations