Robust identification of rare safety-critical differences
Improve the robustness of object-centric set-difference-captioning methods so that they can reliably identify rare, safety-critical differences under very low concentration and low purity settings.
References
As reported in Tab.~\ref{tab:zero_shot_pointing}, all models show low spatial precision, confirming that multi-image spatial reasoning remains a fundamental open challenge.
— Spot-the-shift: Evaluating Grounded Image Difference Captioning of Long-term Changes
(2609.10356 - Liberatori et al., 9 Sep 2026) in Section 5.1, subsection “Benchmarking MLLMs on Spot-the-shift,” paragraph “Difference pointing”
Accuracy, however, drops sharply for very low concentration and low purity settings, so reliably identifying rare, safety-critical differences remains an open challenge. Future work should therefore focus on improving robustness to extreme sparsity.
— Understanding Autonomous Driving Datasets by Describing Differences between Image Subsets in Natural Language
(2609.03677 - Truetsch et al., 3 Sep 2026) in Section Summary and Conclusions