- The paper presents the first in-the-wild diary study of mobile AR visual augmentation for outdoor navigation, following 12 people with low vision for seven days and documenting real-world use, adaptation, and safety concerns.
- NavSight averaged 25.6 frames per second with 70 ms latency, while domain-trained recognition substantially outperformed the baseline model, although crosswalks, benches, and other ambiguous surfaces remained difficult to detect.
- Participants valued simplified walkable paths, hazard cues, and earlier route planning, but divided attention, hand competition, glare, battery drain, and consequence-sensitive AI errors shaped trust and limited use in challenging environments.
Overview
NavSight in the Wild addresses a persistent gap in assistive augmented reality (AR) research for people with low vision (PLV): nearly all prior AR navigation systems—wayfinding guidance, obstacle outlining, depth recoloring, and field-of-view remapping—were evaluated in controlled indoor lab settings over short sessions with predefined tasks. The paper presents NavSight, an iOS mobile AR application that recognizes 21 categories of outdoor objects and renders real-time visual augmentations, and reports the first in-the-wild diary study of AR visual augmentation for PLV outdoor navigation. Through a seven-day deployment with 12 participants, the authors characterize usage patterns, evolving configuration strategies, user interpretation of AI errors, and environmental and social challenges that lab studies cannot capture.
System design and technical evaluation
NavSight is built on Unity with on-device inference via Apple Core ML. Its recognition model is a fine-tuned YOLO11l-seg trained on a refined Mapillary Vistas dataset split 80/10/10 into training, validation, and test sets. The fine-tuned model substantially outperforms the MS-COCO-pretrained baseline across all three instance-segmentation metrics: mAP50 of 0.588 versus 0.256, mAP75 of 0.305 versus 0.129, and mAP of 0.318 versus 0.136. Eleven navigation-critical classes (e.g., curb, sidewalk, crosswalk) score zero under the baseline because they are absent from MS-COCO, underscoring the necessity of domain fine-tuning.
Per-class analysis reveals a consequential asymmetry: cars (FNR = 0.170) and pedestrians (FNR = 0.251) are recognized reliably, while ambiguous surfaces perform poorly—crosswalks have FNR = 0.791 and FDR = 0.284, benches FNR = 0.681, sewer drains FNR = 0.697. This quantitative pattern directly foreshadows participants' later reports that walkable-path recognition degraded in real-world conditions while vehicle and pedestrian recognition remained consistent.
The application offers six augmentations inspired by prior work: four foreground effects (Contour Enhancement, Solid Overlay, Flashing, Brightness Adjustment) and two background effects (Background Darkening, Color Removal). Customization operates at three levels: object selection from 21 categories, assignment to two Augmentation Groups, and per-group augmentation selection with adjustable parameters. On an iPhone 13, end-to-end latency is 70 ms per frame at roughly 25.6 FPS; however, continuous use drains substantial battery relative to a camera-only baseline (~5% over 30 minutes), which itself became a usability concern during the study.
Study method
The seven-day diary study recruited 12 PLV participants aged 25–88 (M=59.08), five of whom were legally blind, covering central vision loss, peripheral vision loss, and severe low acuity. Participants used NavSight for 5–12 days (M=7.92), completing daily surveys rating helpfulness, safety, accuracy, distraction, and public comfort, capturing screenshots via a privacy-conscious shutter-sound feature, and finishing with semi-structured exit interviews probed by their own logs and screenshots. Analysis combined thematic coding with descriptive statistics of daily ratings.
Usage patterns and perceived benefits
All participants used NavSight regularly, averaging 2.1 sessions per day of about 9.3 minutes each, with a negative correlation between session frequency and duration (Pearson r=−0.69). Usage diverged notably: P8 and P9 ran frequent short sessions, while P4 kept the app running continuously (one 23.5-minute session), and P3 never used it while walking—instead launching it 28 times within one 36-minute walk to verify scenes while stationary, deeming walking-with-app unsafe.
The dominant benefit was scene simplification: all 12 participants used NavSight to locate and remain on walkable paths, treating sidewalk contours as navigational rails ("I just stayed in the middle of the green lines, I didn't fall over"). Ten participants detected tripping hazards through segmentation cues, including an emergent mechanism the authors highlight: hazards NavSight could not recognize appeared as "holes" in the sidewalk augmentation, which three participants exploited as detection cues. Seven participants monitored moving objects, some panning the phone toward driveways to scan for turning vehicles—one participant caught a car she would otherwise not have seen. Nine reported extended visual reach enabling earlier route planning, and five described increased confidence exploring unfamiliar routes. Two participants repurposed NavSight beyond navigation entirely—for estimating queue length at a zoo and tracking players in a soccer game—suggesting broader scene-understanding utility.
Ratings were positive but heterogeneous: median helpfulness ranged from 2.00 to 5.00 across participants, as did median safety. Notably, P6 rated it only somewhat helpful every day and P3 rated it unhelpful or below on every day, because objects were outlined only after he had already seen them—visual reach extension is valuable only when it exceeds the user's own viewing range. No participant reported a fall or injury.
Attention competition
The most significant cost was divided attention. Eight participants reported splitting attention between screen and surroundings ("when I'm watching the phone, I'm not watching where my feet are going"), which reduced awareness of unaugmented hazards. Distraction was activity-dependent: P12 rated NavSight not at all distracting when stationary but very distracting when crossing streets, and consequently preferred watching traffic himself. Five participants also reported hand competition—the phone occupied a hand needed for canes, handrails, or balance.
Importantly, adaptation occurred within the week: five participants showed declining distraction ratings, two reaching floor ratings by day two, suggesting a learning curve that short lab evaluations would miss. Six participants proposed wearable displays to resolve both attention and hand competition simultaneously.
AI errors in the wild
Participants identified four error types: segmentation inaccuracy ("squiggly" boundaries), temporal instability (flickering between frames), false positives on out-of-list objects, and missed objects at uncommon orientations. They attributed path-recognition failures to three environmental factors—rain altering surface appearance, tree shadows creating spurious bright patches, and nonstandard crosswalk markings (a red-brick crosswalk partially recognized; a yellow-striped one missed entirely).
Attitudes toward errors were asymmetric by consequence rather than type: false positives that added information were tolerated or even useful (P1 found indoor mall aisles augmented as sidewalks genuinely helpful), whereas segmentation errors extending onto steps concealed height changes and were considered dangerous. Five participants constructed mental models of error causes (phone steadiness, number of selected objects, augmentation choice) and adjusted behavior accordingly; two who could not form coherent explanations of inconsistent recognition reported reduced trust—"It lessens the faith you have... in the act of helping or protecting you."
Evolving configuration preferences
Object selections converged quickly to stable sets (common paths, moving objects, low tripping hazards), with scenario-specific additions thereafter. Grouping strategies varied: seven grouped by risk level (by motion or by height), while three preferred a single group, finding multi-group configuration cognitively burdensome. Augmentation choices evolved continuously around a visibility–occlusion tradeoff: four participants who began with solid overlays switched to contour enhancement once they learned to customize thickness, since overlays occluded the very objects being marked. Ten adapted augmentations to lighting conditions—changing colors against bright skies, darkening backgrounds at midday, brightening objects at sunset—and two rotated colors during long sessions to reduce eye strain.
Social acceptability
Ten participants felt comfortable using NavSight publicly, and eight explicitly valued its discretion compared to a white cane, which signals visual impairment ("if I have a cane for walking, everyone looks at you"). Yet this advantage carried a countervailing concern: six worried that holding up a phone would be read as filming strangers, and one participant was questioned by a passerby. One participant rated public comfort at the minimum on all seven days for this reason and restricted her use to familiar, crowded areas. A further tension emerged: one participant wished the app could signal his visual condition on demand after bumping into a pedestrian, since the phone concealed the impairment his cane would have explained.
Usability issues unique to field deployment
Three issues surfaced that lab studies would plausibly miss. First, two participants fully darkened the background, saw a black screen when nothing was augmented, and mistook their own configuration for a malfunction—one for three days, another until contacting researchers. Second, adverse weather impaired not only recognition but augmentation visibility: ten participants struggled with screen glare in sunlight, and three stopped using the app entirely in bright conditions. Third, battery drain and conflicts with other assistive apps (e.g., a separate pedestrian-signal app) constrained sustained use.
Design implications and limitations
The authors derive implications along several axes: context-adaptive augmentation to reduce manual reconfiguration burden; intelligent configuration debugging support (e.g., an agent answering "why is the screen black?"); offering options across platforms given the phone-versus-glasses tradeoff between social acceptability and hands-free operation; expanding recognition beyond discrete objects to surface attributes such as cracks and slopes; prioritizing the visual challenges shared by PLV and vision models, since NavSight was least reliable precisely when users' own vision failed them; supporting real-time trust calibration, noting that unlike query-based applications, streaming AR errors cannot be re-examined; and matching evaluation metrics to error consequences, since mAP weights all errors equally while users judge errors asymmetrically.
Stated limitations include iOS-only deployment, a sample limited to iPhone owners, self-reported outcome measures, a one-week window too short to reveal abandonment patterns, and the absence of camera-feed logging (for privacy), which precluded quantitative root-cause analysis of recognition errors.
Conclusion
This paper provides the first ecological evaluation of AR visual augmentation for PLV outdoor navigation, pairing a deployable, customizable mobile system with a week-long diary study. Its central contributions are empirical: documentation of how PLV actually integrate such tools alongside existing aids, evidence of within-week adaptation to divided attention, characterization of consequence-based error tolerance, and identification of environmental conditions—rain, shadows, nonstandard markings—that degrade recognition exactly when residual vision is least sufficient. The findings collectively argue that future AI-powered AR aids should be evaluated in the wild, adapt automatically to context, and treat trust calibration and consequence-aware metrics as first-class design requirements.