Papers
Topics
Authors
Recent
Search
2000 character limit reached

NaviSense: Integrated Navigation & Sensing

Updated 12 July 2026
  • NaviSense is an integrated design pattern that fuses spatial sensing with navigation by coupling a sensing stack, localization, and closed-loop feedback.
  • Implementations range from RF passive radar on RFSoC to ultrasound probe guidance, head-mounted assistive wearables, and mobile AR object retrieval.
  • Each system achieves real-time processing and actionable output by integrating multi-channel sensor fusion with domain-specific registration and feedback loops.

NaviSense is a label used in the literature for several integrated navigation-and-sensing systems rather than for a single standardized platform. In the available arXiv record, it denotes or is used to describe: an RFSoC-based NavIC L5 passive radar receiver for integrated navigation and remote sensing; an optical-tracking and feedback framework for blinded 4D contrast-enhanced ultrasound; a head-mounted multimodal assistive system for blind users; and a mobile application for open-world object retrieval by blind and low-vision users. Across these implementations, the recurring technical motif is the coupling of a sensing stack, a localization or registration mechanism, and a closed-loop output channel that supports either target detection, image alignment, or human guidance (Sachdeva et al., 9 Feb 2026, Kaffas et al., 2020, Bobba et al., 2023, Sridhar et al., 23 Sep 2025).

1. Scope, nomenclature, and recurring system pattern

The literature uses NaviSense for distinct systems operating in RF remote sensing, medical ultrasound, wearable assistive perception, and mobile accessibility. The term therefore does not identify a single hardware lineage or software stack. This suggests a broader design pattern: integrated navigation plus sensing, with task-specific feedback and decision support.

Variant Core sensing stack Primary task
RFSoC/NavIC receiver NavIC L5, dual synchronized RF channels, ARM PS + FPGA PL Delay-Doppler mapping of ground-based targets
4D DCE-US framework IR tracking, matrix transducer, MevisLab workstation Probe guidance, motion correction, live TIC feedback
Head-mounted assistive system RGB camera, ultrasonic array, microphones, IMU Object recognition, text reading, obstacle avoidance
Mobile assistive application Rear camera, LiDAR, ARKit, conversational AI, audio-haptics Open-world object retrieval

A common source of ambiguity is that these systems share a name while differing substantially in modality, inference stack, and evaluation protocol. In RF work, navigation refers to satellite timing and code-phase reference; in ultrasound, it refers to tracked probe positioning; in assistive systems, it refers to human wayfinding or object-directed reaching. The unifying element is not modality but the integration of spatial inference with actionable output.

2. RFSoC-based integrated navigation and sensing using NavIC

In "RFSoC-Based Integrated Navigation and Sensing Using NavIC" (Sachdeva et al., 9 Feb 2026), NaviSense denotes a passive radar receiver system mounted on an AMD Zynq RFSoC 4×2 platform. The Processing System is a quad-core ARM Cortex-A53 running PetaLinux and PYNQ for device control, NavIC satellite coarse acquisition, and delay-Doppler map generation. The Programmable Logic contains four parallel data-streaming engines implemented with AXI-Stream and DMA, multi-stage interpolators and decimators, block RAMs for sample buffering, and DSP slices for on-chip FIR and NCO operations. Two DAC tiles and two ADC tiles are clocked up to 2.45 GHz, with Multi-Tile Synchronization across all active DAC and ADC tiles via a common 245.76 MHz reference derived from the external LMX2594 PLL. Zero-phase alignment is used so that the direct-signal and ground-reflected-signal channels see identical clocks, eliminating inter-channel timing and Doppler skew.

The receiver architecture is explicitly dual-channel. The DS channel uses an RHCP patch antenna pointed toward the NavIC sky-visible hemisphere and connected to ADC#0. The GRS channel uses a nadir-pointing or quasi-hemispherical LHCP antenna connected to ADC#1. Both channels follow parallel RF chains with low-noise amplification, bandpass filtering around 1.176 GHz, ADC sampling at 2.45 GHz, PL downconversion and decimation, and transfer to the PS.

The NavIC L5 signal parameters are specified as center frequency fc=1.176 GHzf_c = 1.176\ \text{GHz}, PRN chip rate fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}, code length of 1 ms1\ \text{ms} with $1023$ chips and 100%100\% duty cycle, and bandwidth of approximately 2 MHz2\ \text{MHz}. The signal-processing chain downconverts to baseband via NCO, decimates by $40$ to 61.44 MHz61.44\ \text{MHz}, then soft-decimates by 232^3 to a final PS rate of 7.68 MHz7.68\ \text{MHz}. In loopback mode, FIR interpolation by fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}0 is implemented via three cascaded fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}1 interpolators with fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}2, fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}3, and fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}4 taps, with FIR decimation by fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}5 in reverse order on receive.

The matched-filter formulation is given by the DS and GRS baseband models

fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}6

and

fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}7

with cross-ambiguity

fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}8

In discrete form,

fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}9

The implementation uses 1 ms1\ \text{ms}0 samples/s 1 ms1\ \text{ms}1 samples per PRN period, with Doppler search over 1 ms1\ \text{ms}2 in 1 ms1\ \text{ms}3 bins of 1 ms1\ \text{ms}4.

Validation is performed in two configurations. In Configuration A, the PS generates a resampled PRN-2 code or 1 ms NavIC packet with built-in AWGN, path loss, and integrated Doppler and time delay; the PL interpolates by 1 ms1\ \text{ms}5, the DAC output is looped back to the ADC inputs, and the PL decimates by 1 ms1\ \text{ms}6 before C/A and DDM processing. In Configuration B, a Keysight M8910A AWG provides two synchronized RF channels loaded with 12-bit, 2.45 GHz DS and GRS NavIC waveforms. AWGN is added to realize SNR values of 1 ms1\ \text{ms}7, 1 ms1\ \text{ms}8, and 1 ms1\ \text{ms}9 relative to code power.

The reported performance includes native range resolution

$1023$0

with code ambiguity giving approximately $1023$1, and oversampled $1023$2 processing yielding approximately $1023$3 bin precision. Doppler resolution is $1023$4, with a $1023$5 span. The DS channel peak is approximately $1023$6 above the noise floor for the correct PRN; the GRS channel is approximately $1023$7-$1023$8 lower post-processing gain due to reflection loss. Range-error RMSE is reported as $1023$9 at 100%100\%0, 100%100\%1 at 100%100\%2, and 100%100\%3 at 100%100\%4; Doppler-error RMSE is approximately 100%100\%5, 100%100\%6, and 100%100\%7, respectively. Both systems successfully detect up to 100%100\%8 SNR.

In this usage, NaviSense is an integrated passive radar architecture in which satellite navigation signals provide both timing reference and illuminator waveform. The significance lies in consolidating RF front end, synchronized data conversion, FPGA filtering, and delay-Doppler processing on a single RFSoC.

3. Optical-tracking navigation and feedback in 4D contrast-enhanced ultrasound

In "Clinical Evaluation of Real-Time Optical-Tracking Navigation and Live Time-Intensity Curves to Provide Feedback During Blinded 4D Contrast-Enhanced Ultrasound Imaging" (Kaffas et al., 2020), the NaviSense-style framework consists of an infrared tracking camera, a 3D-printed tracking target attached to a Philips X6-1 matrix transducer, a Philips EPIQ7 system, and an interventional workstation running MevisLab. The stereo pair of Polaris IR cameras is mounted over the patient bed and oriented before each session with an integrated laser pointer to cover the lower-mid abdomen. The target holds four uniquely arranged spherical retroreflective markers, enabling full 6-DOF tracking, with calibration uncertainty less than 100%100\%9 using automated hand-eye calibration and self-consistency rejection metrics. Ultrasound frames are streamed over a 1 Gb/s Digital Navigation Link to a modular workstation for tracking-to-image linking, virtual probe rendering, live time-intensity curves, and data capture.

The coordinate model comprises world coordinates, tracker coordinates, and image coordinates. Rigid transforms 2 MHz2\ \text{MHz}0 and 2 MHz2\ \text{MHz}1 are reconstructed via hand-eye calibration, and the ultrasound image-plane to probe-head transform is obtained automatically via intramodality image registration over overlapping volumes. At time 2 MHz2\ \text{MHz}2, the IR system provides marker centroids, from which probe translation 2 MHz2\ \text{MHz}3 and rotation are deduced. The virtual probe display renders a live grey probe and a static red reference probe. Once the operator identifies an anatomical landmark in volumetric B-mode and clicks capture, the system stores the world-frame pose 2 MHz2\ \text{MHz}4 and displays subsequent live-reference alignment error as the offset between the grey and red icons.

The framework also supports live TIC acquisition during contrast infusion. The reported protocol uses Definity microbubbles 2 MHz2\ \text{MHz}5 diluted in 2 MHz2\ \text{MHz}6 saline and infused at 2 MHz2\ \text{MHz}7 for up to 2 MHz2\ \text{MHz}8. Disruption-replenishment pulses consist of three high-MI 2 MHz2\ \text{MHz}9 flash frames embedded in low-MI $40$0 volumetric contrast-mode imaging. B-mode is used only initially to pick the lock position; contrast-mode remains blind of B-mode side-by-side. An in-house MevisLab module computes and plots TICs from a user-defined VOI in real time.

Displacement is defined for the center voxel of the locked volume as $40$1 in mm, with magnitude

$40$2

For motion correction, each acquired 3D contrast volume $40$3 is associated with pose $40$4, and re-alignment to the reference frame applies the inverse rigid transform $40$5. The registration is purely rigid and does not require voxel-based image similarity.

The experimental protocol includes an operator performance study with $40$6 ultrasound operators having at least $40$7 years of experience and trained for $40$8 minutes on the tracking system. In Task A, operators identify an abdominal liver landmark in B-mode, remove the probe, reposition, and record positional error and time. In Task B, they maintain the locked pose for $40$9 minutes under three feedback modes: B-mode only, virtual transducer display with red reference probe, and blind. The clinical cohort comprises 61.44 MHz61.44\ \text{MHz}0 adult patients and 61.44 MHz61.44\ \text{MHz}1 scans with liver metastases greater than 61.44 MHz61.44\ \text{MHz}2, using two flash-replenishment cycles spaced 61.44 MHz61.44\ \text{MHz}3 minutes apart.

For repositioning in Task A, mean error is reported as approximately 61.44 MHz61.44\ \text{MHz}4 for B-mode, approximately 61.44 MHz61.44\ \text{MHz}5 for tracking-assist, and approximately 61.44 MHz61.44\ \text{MHz}6 for blind, with corresponding mean times of approximately 61.44 MHz61.44\ \text{MHz}7, 61.44 MHz61.44\ \text{MHz}8, and 61.44 MHz61.44\ \text{MHz}9. For maintenance displacement in Task B, mean 232^30 is 232^31 for B-mode, 232^32 for blind, and 232^33 for tracking; the corresponding SD values are 232^34, 232^35, and 232^36. Perfusion repeatability improves after rigid re-alignment: rBF ICC increases from 232^37 to 232^38, and rBV ICC from 232^39 to 7.68 MHz7.68\ \text{MHz}0. Time to steady state varies widely across lesion, parenchyma, and portal vein, with a range of approximately 7.68 MHz7.68\ \text{MHz}1-7.68 MHz7.68\ \text{MHz}2.

Here, NaviSense-style operation is centered on tracked probe pose, virtualized reference guidance, and quantitative stabilization of longitudinal contrast measurements. The framework is therefore a navigation system in the interventional-imaging sense, not in the RF or pedestrian sense.

4. Head-mounted multimodal assistive NaviSense for blind users

In "Newvision: application for helping blind people using deep learning" (Bobba et al., 2023), NaviSense is a lightweight, modular visor-style headset, with a target final form factor of slim spectacles, intended to help visually impaired users navigate surroundings, identify objects and people, read text, and avoid obstacles. The hardware described in the technical overview includes a forward-facing 7.68 MHz7.68\ \text{MHz}3 RGB camera at 7.68 MHz7.68\ \text{MHz}4 fps with 7.68 MHz7.68\ \text{MHz}5 horizontal field of view; an array of four ultrasonic transceivers with 7.68 MHz7.68\ \text{MHz}6 center frequency; a stereo microphone pair at 7.68 MHz7.68\ \text{MHz}7 and 7.68 MHz7.68\ \text{MHz}8 bit; an onboard compute module with a quad-core ARM Cortex-A53 CPU, an integrated 7.68 MHz7.68\ \text{MHz}9-core neural-network GPU fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}00, fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}01 LPDDR4 memory, and fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}02 flash; a fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}03, fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}04 Li-Po battery; and a small IMU. Total device mass is approximately fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}05. The ultrasonic sensors are angled slightly downward by fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}06 to capture typical walking-path obstacles.

The run-time loop operates at approximately fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}07. Sensor acquisition includes camera capture, ultrasonic time-of-flight sampled at fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}08 with a fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}09 averaging window, microphone streaming to speech-to-text, and IMU orientation at fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}10. Image preprocessing resizes frames to fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}11 and normalizes with mean and standard deviation from COCO. Ultrasonic TOF samples are converted to distance through

fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}12

where fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}13. Audio buffers are windowed and processed with voice activity detection.

The vision-language core follows a BLIP-inspired architecture. The image encoder is a fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}14-layer ViT-Base/384 with fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}15-dimensional embeddings and fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}16 attention heads; the text encoder is BERT-Base with fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}17 layers, fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}18 hidden units, and fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}19 self-attention heads. Cross-modal fusion is implemented through an image-grounded text encoder with cross-attention layers and an image-grounded text decoder initialized from a GPT-style transformer. The pre-training objectives are Image-Text Contrastive, Image-Text Matching, and Language Modeling. The summary specifies pre-training on approximately fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}20 million images and fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}21 million captions from COCO, Flickr30k, NLVR, and NoCaps, and fine-tuning on COCO detection splits and ICDAR text datasets, with AdamW, initial learning rate fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}22, weight decay fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}23, batch size fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}24, and fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}25 epochs.

Inference uses the on-board GPU in FP16 quantized form for less than fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}26 per frame. An object-detection head yields bounding boxes and class scores, while a text-recognition head uses a lightweight ResNet plus CTC decoder. Detected object centroids fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}27 and ultrasonic distances are fused into a 2D occupancy grid via a Kalman filter with state transition fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}28 and measurement fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}29, where fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}30 under a constant-velocity model. A voice-command parser invokes modules such as “What is that?” or “Navigate to X,” and navigation uses A* on a fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}31 occupancy grid with fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}32 resolution. End-to-end sensing to voice output completes within fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}33, with optional haptic buzzer feedback for imminent collisions at distance less than fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}34.

Reported benchmark metrics include COCO Image Captioning CIDEr fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}35, VQA v2 accuracy fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}36, NLVR accuracy fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}37, and in-device object detection mAP of approximately fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}38 on COCO val2017. The speech stack uses an on-device RNN-Transducer with real-time latency below fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}39, a context-free grammar with approximately fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}40 production rules, and Tacotron-2 plus WaveRNN at fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}41 with overall TTS latency below fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}42.

The pilot user study includes fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}43 visually impaired volunteers aged fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}44-fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}45, performing corridor navigation, cluttered room exploration, sign reading, and object identification. Reported outcomes are object identification accuracy of fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}46, mean ultrasonic distance estimation error of fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}47, navigation success rate of fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}48, average completion time of fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}49 the sighted baseline, end-to-end system latency of fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}50, and overall user satisfaction of fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}51 on a fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}52-fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}53 Likert scale.

In this form, NaviSense is a wearable multimodal perception-and-guidance system. Its defining technical feature is the joint use of deep vision-language inference, ultrasonic ranging, speech interaction, and local path planning on head-mounted hardware.

In "NaviSense: A Multimodal Assistive Mobile application for Object Retrieval by Persons with Visual Impairment" (Sridhar et al., 23 Sep 2025), NaviSense is a mobile assistive system that combines conversational AI, a vision-LLM, AR, LiDAR, and synchronized audio-haptic feedback for open-world object detection with real-time guidance. The app runs on an iPhone 16 Pro with A18 chipset, fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}54 unified memory, iOS 17, a fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}55 MP wide camera, and a LiDAR depth scanner. The software stack uses ARKit for pose tracking and depth fusion, Core Haptics for vibration patterns, Apple Speech for on-device ASR and TTS, Moondream 2B for cloud VLM inference, and GPT-4o-mini for intent parsing and confirmation. An Apple M2 MacBook Pro hosts containerized services for GPT-4o-mini and Moondream 2B via REST.

The interaction model is controlled by a finite state machine with transitions Idle fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}56 Listening fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}57 Processing fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}58 Speaking fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}59 Scanning fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}60 Guiding, and shaking the phone cancels the current task. A user request is transcribed on-device and sent via HTTPS to GPT-4o-mini. Approximately once per second, an image frame is sent to Moondream 2B together with a prompt such as “find X,” and the model returns 2D bounding boxes of candidate objects. Once the target is detected, ARKit maps the pixel centroid fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}61 to a 3D point fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}62 using LiDAR depth, and guidance enters a closed-loop phase.

The spatial formulation includes Euclidean distance

fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}63

a guidance vector

fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}64

and angular deviation

fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}65

The implementation description also defines the unit guidance vector fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}66. Audio feedback updates at approximately fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}67 with messages such as “turn left slightly” or “move forward,” while haptic pulse rate increases as distance decreases and angular alignment improves. Example piecewise feedback values are approximately fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}68 for fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}69, approximately fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}70 for fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}71, and approximately fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}72 for fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}73. When fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}74 exceeds a threshold of approximately fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}75, lateral pulses or stereo-panned audio indicate turning.

The system is explicitly open-world: it requires no pre-scanning and no category restriction. The example interaction includes clarification by the LLM when multiple candidate objects are present, such as distinguishing a blue ceramic coffee mug from a metal travel mug. Feedback ceases when fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}76 or upon pickup gesture, and the system returns to Listening.

Latency figures are given separately for components and the full interaction loop. ASR and TTS run locally in less than fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}77; cloud VLM round-trip averages approximately fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}78 with fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}79; the guidance loop is tuned to fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}80. The evaluation uses a within-subject comparison with fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}81 blind or low-vision participants, of whom fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}82 were blind and fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}83 low vision, aged fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}84-fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}85 with mean fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}86. Comparison systems are NaviSense, Be My AI, and Ray-Ban Meta Glasses. Three target objects were placed on a four-tier shelf, with fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}87 trials per object and system, yielding fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}88 trials per participant.

The quantitative results are reported as follows. Search time is fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}89 for Be My AI, fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}90 for Meta Glasses, and fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}91 for NaviSense. Guidance time is fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}92, fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}93, and fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}94, respectively. Total time is fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}95, fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}96, and fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}97. Undesired touches are fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}98, fchip=1.023 MHzf_{\text{chip}} = 1.023\ \text{MHz}99, and 1 ms1\ \text{ms}00. Accuracy is 1 ms1\ \text{ms}01, 1 ms1\ \text{ms}02, and 1 ms1\ \text{ms}03. Repeated-measures ANOVA yields 1 ms1\ \text{ms}04 for search time and 1 ms1\ \text{ms}05 for total time, while guidance time is non-significant. The Friedman test on errors gives 1 ms1\ \text{ms}06. Overall preference is 1 ms1\ \text{ms}07 for NaviSense, 1 ms1\ \text{ms}08 for Be My AI, and 1 ms1\ \text{ms}09 for Meta, with Friedman 1 ms1\ \text{ms}10. A supplementary evaluation on 1 ms1\ \text{ms}11 common household items and 1 ms1\ \text{ms}12 frames reports 1 ms1\ \text{ms}13 detection accuracy with 1 ms1\ \text{ms}14 CI 1 ms1\ \text{ms}15.

This implementation represents the most explicit use of NaviSense as a named assistive product. Its distinctive contribution is the combination of open-world language-conditioned detection with AR/LiDAR localization and last-metre audio-haptic guidance on commodity mobile hardware.

6. Cross-domain technical characteristics and significance

Across the four uses of NaviSense, several architectural regularities recur. First, each system is multimodal: the RFSoC receiver couples direct and ground-reflected NavIC channels; the ultrasound framework couples IR pose tracking with volumetric imaging; the headset combines RGB, ultrasonics, microphones, and IMU; and the mobile application combines camera, LiDAR, ARKit pose, ASR/TTS, and cloud models. Second, each system uses an explicit spatial transform or registration layer: RF delay-Doppler cross-ambiguity, rigid transforms 1 ms1\ \text{ms}16 and 1 ms1\ \text{ms}17, Kalman-filtered occupancy-grid state, or ARKit pose and depth projection.

Third, all implementations close the loop from sensing to action. In the RF system, the output is a delay-Doppler map for passive radar extraction. In ultrasound, it is live probe-position feedback and re-aligned TIC quantification. In the headset, it is verbal scene description, path planning, and obstacle alerts. In the mobile app, it is synchronized audio-haptic object retrieval guidance. This suggests that NaviSense functions less as a modality-specific technology than as an integration pattern in which localization and sensing are operationally inseparable.

The evaluation criteria are likewise domain-specific but structurally parallel. The RF platform reports range-error RMSE, Doppler-error RMSE, and SNR robustness. The ultrasound system reports repositioning error, displacement statistics, and ICC for rBF and rBV. The headset reports object-identification accuracy, navigation success, and latency. The mobile app reports search time, total time, undesired touches, retrieval accuracy, and subjective ratings. A plausible implication is that NaviSense systems are best understood through task-level closed-loop performance rather than through isolated component benchmarks.

The available literature therefore presents NaviSense not as a single canonical platform, but as a family of systems in which sensing, navigation, and feedback are co-designed for specific operational settings. In RF sensing, this yields compact passive radar on RFSoC; in interventional ultrasound, tracked alignment and repeatable perfusion quantification; and in accessibility, open-world perception with actionable guidance for blind and low-vision users.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to NaviSense.