Quantitative identification of causes of recognition inaccuracies

Identify quantitatively the causes of recognition inaccuracies in NavSight’s real-world object-recognition and augmentation pipeline by analyzing appropriate real-world navigation data while preserving participant privacy.

Background

NavSight exhibited recognition errors whose occurrence appeared related to environmental conditions such as rain, wet surfaces, lighting and shadows, and nonstandard road markings and textures. The study reports participants’ observations of these errors and evaluates the fine-tuned recognition model on a curated test set, but it does not quantitatively link errors observed during deployment to specific camera views, environmental conditions, or recognition outputs.

The authors explicitly state that privacy-preserving data collection prevented quantitative analysis of these causes. They identify collection of real-world egocentric navigation data and examination of recognition-error causes as future work, leaving the quantitative causal analysis unresolved.

References

Finally, we were unable to quantitatively analyze the causes of recognition inaccuracies because we did not log users' camera feeds or recognition results to preserve privacy. Future work could deploy NavSight across more platforms and with larger, more diverse samples of PLV over longer periods, incorporate objective task-performance measures, and collect real-world egocentric navigation data to examine the causes of recognition errors and how PLV's experiences vary across platforms, individuals, and environments.

NavSight in the Wild: Understanding Real-World Use of a Mobile Augmented Reality Application for People with Low Vision in Outdoor Navigation  (2608.12759 - Wu et al., 13 Aug 2026) in Section 6.4, Limitations and Future Directions

We conjecture that this issue may stem from the limited alignment between the MLLM's scale estimation and real-world spatial scale, while fine-grained object recognition in complex scenes remains challenging.

LookStep: Efficient Vision-Language Navigation with Linguistic Foresight and Event Driven Memory  (2609.02350 - Yu et al., 2 Sep 2026) in Section Failure Case Analysis