NaviSense: Integrated Navigation & Sensing
- NaviSense is an integrated design pattern that fuses spatial sensing with navigation by coupling a sensing stack, localization, and closed-loop feedback.
- Implementations range from RF passive radar on RFSoC to ultrasound probe guidance, head-mounted assistive wearables, and mobile AR object retrieval.
- Each system achieves real-time processing and actionable output by integrating multi-channel sensor fusion with domain-specific registration and feedback loops.
NaviSense is a label used in the literature for several integrated navigation-and-sensing systems rather than for a single standardized platform. In the available arXiv record, it denotes or is used to describe: an RFSoC-based NavIC L5 passive radar receiver for integrated navigation and remote sensing; an optical-tracking and feedback framework for blinded 4D contrast-enhanced ultrasound; a head-mounted multimodal assistive system for blind users; and a mobile application for open-world object retrieval by blind and low-vision users. Across these implementations, the recurring technical motif is the coupling of a sensing stack, a localization or registration mechanism, and a closed-loop output channel that supports either target detection, image alignment, or human guidance (Sachdeva et al., 9 Feb 2026, Kaffas et al., 2020, Bobba et al., 2023, Sridhar et al., 23 Sep 2025).
1. Scope, nomenclature, and recurring system pattern
The literature uses NaviSense for distinct systems operating in RF remote sensing, medical ultrasound, wearable assistive perception, and mobile accessibility. The term therefore does not identify a single hardware lineage or software stack. This suggests a broader design pattern: integrated navigation plus sensing, with task-specific feedback and decision support.
| Variant | Core sensing stack | Primary task |
|---|---|---|
| RFSoC/NavIC receiver | NavIC L5, dual synchronized RF channels, ARM PS + FPGA PL | Delay-Doppler mapping of ground-based targets |
| 4D DCE-US framework | IR tracking, matrix transducer, MevisLab workstation | Probe guidance, motion correction, live TIC feedback |
| Head-mounted assistive system | RGB camera, ultrasonic array, microphones, IMU | Object recognition, text reading, obstacle avoidance |
| Mobile assistive application | Rear camera, LiDAR, ARKit, conversational AI, audio-haptics | Open-world object retrieval |
A common source of ambiguity is that these systems share a name while differing substantially in modality, inference stack, and evaluation protocol. In RF work, navigation refers to satellite timing and code-phase reference; in ultrasound, it refers to tracked probe positioning; in assistive systems, it refers to human wayfinding or object-directed reaching. The unifying element is not modality but the integration of spatial inference with actionable output.
2. RFSoC-based integrated navigation and sensing using NavIC
In "RFSoC-Based Integrated Navigation and Sensing Using NavIC" (Sachdeva et al., 9 Feb 2026), NaviSense denotes a passive radar receiver system mounted on an AMD Zynq RFSoC 4×2 platform. The Processing System is a quad-core ARM Cortex-A53 running PetaLinux and PYNQ for device control, NavIC satellite coarse acquisition, and delay-Doppler map generation. The Programmable Logic contains four parallel data-streaming engines implemented with AXI-Stream and DMA, multi-stage interpolators and decimators, block RAMs for sample buffering, and DSP slices for on-chip FIR and NCO operations. Two DAC tiles and two ADC tiles are clocked up to 2.45 GHz, with Multi-Tile Synchronization across all active DAC and ADC tiles via a common 245.76 MHz reference derived from the external LMX2594 PLL. Zero-phase alignment is used so that the direct-signal and ground-reflected-signal channels see identical clocks, eliminating inter-channel timing and Doppler skew.
The receiver architecture is explicitly dual-channel. The DS channel uses an RHCP patch antenna pointed toward the NavIC sky-visible hemisphere and connected to ADC#0. The GRS channel uses a nadir-pointing or quasi-hemispherical LHCP antenna connected to ADC#1. Both channels follow parallel RF chains with low-noise amplification, bandpass filtering around 1.176 GHz, ADC sampling at 2.45 GHz, PL downconversion and decimation, and transfer to the PS.
The NavIC L5 signal parameters are specified as center frequency , PRN chip rate , code length of with $1023$ chips and duty cycle, and bandwidth of approximately . The signal-processing chain downconverts to baseband via NCO, decimates by $40$ to , then soft-decimates by to a final PS rate of . In loopback mode, FIR interpolation by 0 is implemented via three cascaded 1 interpolators with 2, 3, and 4 taps, with FIR decimation by 5 in reverse order on receive.
The matched-filter formulation is given by the DS and GRS baseband models
6
and
7
with cross-ambiguity
8
In discrete form,
9
The implementation uses 0 samples/s 1 samples per PRN period, with Doppler search over 2 in 3 bins of 4.
Validation is performed in two configurations. In Configuration A, the PS generates a resampled PRN-2 code or 1 ms NavIC packet with built-in AWGN, path loss, and integrated Doppler and time delay; the PL interpolates by 5, the DAC output is looped back to the ADC inputs, and the PL decimates by 6 before C/A and DDM processing. In Configuration B, a Keysight M8910A AWG provides two synchronized RF channels loaded with 12-bit, 2.45 GHz DS and GRS NavIC waveforms. AWGN is added to realize SNR values of 7, 8, and 9 relative to code power.
The reported performance includes native range resolution
$1023$0
with code ambiguity giving approximately $1023$1, and oversampled $1023$2 processing yielding approximately $1023$3 bin precision. Doppler resolution is $1023$4, with a $1023$5 span. The DS channel peak is approximately $1023$6 above the noise floor for the correct PRN; the GRS channel is approximately $1023$7-$1023$8 lower post-processing gain due to reflection loss. Range-error RMSE is reported as $1023$9 at 0, 1 at 2, and 3 at 4; Doppler-error RMSE is approximately 5, 6, and 7, respectively. Both systems successfully detect up to 8 SNR.
In this usage, NaviSense is an integrated passive radar architecture in which satellite navigation signals provide both timing reference and illuminator waveform. The significance lies in consolidating RF front end, synchronized data conversion, FPGA filtering, and delay-Doppler processing on a single RFSoC.
3. Optical-tracking navigation and feedback in 4D contrast-enhanced ultrasound
In "Clinical Evaluation of Real-Time Optical-Tracking Navigation and Live Time-Intensity Curves to Provide Feedback During Blinded 4D Contrast-Enhanced Ultrasound Imaging" (Kaffas et al., 2020), the NaviSense-style framework consists of an infrared tracking camera, a 3D-printed tracking target attached to a Philips X6-1 matrix transducer, a Philips EPIQ7 system, and an interventional workstation running MevisLab. The stereo pair of Polaris IR cameras is mounted over the patient bed and oriented before each session with an integrated laser pointer to cover the lower-mid abdomen. The target holds four uniquely arranged spherical retroreflective markers, enabling full 6-DOF tracking, with calibration uncertainty less than 9 using automated hand-eye calibration and self-consistency rejection metrics. Ultrasound frames are streamed over a 1 Gb/s Digital Navigation Link to a modular workstation for tracking-to-image linking, virtual probe rendering, live time-intensity curves, and data capture.
The coordinate model comprises world coordinates, tracker coordinates, and image coordinates. Rigid transforms 0 and 1 are reconstructed via hand-eye calibration, and the ultrasound image-plane to probe-head transform is obtained automatically via intramodality image registration over overlapping volumes. At time 2, the IR system provides marker centroids, from which probe translation 3 and rotation are deduced. The virtual probe display renders a live grey probe and a static red reference probe. Once the operator identifies an anatomical landmark in volumetric B-mode and clicks capture, the system stores the world-frame pose 4 and displays subsequent live-reference alignment error as the offset between the grey and red icons.
The framework also supports live TIC acquisition during contrast infusion. The reported protocol uses Definity microbubbles 5 diluted in 6 saline and infused at 7 for up to 8. Disruption-replenishment pulses consist of three high-MI 9 flash frames embedded in low-MI $40$0 volumetric contrast-mode imaging. B-mode is used only initially to pick the lock position; contrast-mode remains blind of B-mode side-by-side. An in-house MevisLab module computes and plots TICs from a user-defined VOI in real time.
Displacement is defined for the center voxel of the locked volume as $40$1 in mm, with magnitude
$40$2
For motion correction, each acquired 3D contrast volume $40$3 is associated with pose $40$4, and re-alignment to the reference frame applies the inverse rigid transform $40$5. The registration is purely rigid and does not require voxel-based image similarity.
The experimental protocol includes an operator performance study with $40$6 ultrasound operators having at least $40$7 years of experience and trained for $40$8 minutes on the tracking system. In Task A, operators identify an abdominal liver landmark in B-mode, remove the probe, reposition, and record positional error and time. In Task B, they maintain the locked pose for $40$9 minutes under three feedback modes: B-mode only, virtual transducer display with red reference probe, and blind. The clinical cohort comprises 0 adult patients and 1 scans with liver metastases greater than 2, using two flash-replenishment cycles spaced 3 minutes apart.
For repositioning in Task A, mean error is reported as approximately 4 for B-mode, approximately 5 for tracking-assist, and approximately 6 for blind, with corresponding mean times of approximately 7, 8, and 9. For maintenance displacement in Task B, mean 0 is 1 for B-mode, 2 for blind, and 3 for tracking; the corresponding SD values are 4, 5, and 6. Perfusion repeatability improves after rigid re-alignment: rBF ICC increases from 7 to 8, and rBV ICC from 9 to 0. Time to steady state varies widely across lesion, parenchyma, and portal vein, with a range of approximately 1-2.
Here, NaviSense-style operation is centered on tracked probe pose, virtualized reference guidance, and quantitative stabilization of longitudinal contrast measurements. The framework is therefore a navigation system in the interventional-imaging sense, not in the RF or pedestrian sense.
4. Head-mounted multimodal assistive NaviSense for blind users
In "Newvision: application for helping blind people using deep learning" (Bobba et al., 2023), NaviSense is a lightweight, modular visor-style headset, with a target final form factor of slim spectacles, intended to help visually impaired users navigate surroundings, identify objects and people, read text, and avoid obstacles. The hardware described in the technical overview includes a forward-facing 3 RGB camera at 4 fps with 5 horizontal field of view; an array of four ultrasonic transceivers with 6 center frequency; a stereo microphone pair at 7 and 8 bit; an onboard compute module with a quad-core ARM Cortex-A53 CPU, an integrated 9-core neural-network GPU 00, 01 LPDDR4 memory, and 02 flash; a 03, 04 Li-Po battery; and a small IMU. Total device mass is approximately 05. The ultrasonic sensors are angled slightly downward by 06 to capture typical walking-path obstacles.
The run-time loop operates at approximately 07. Sensor acquisition includes camera capture, ultrasonic time-of-flight sampled at 08 with a 09 averaging window, microphone streaming to speech-to-text, and IMU orientation at 10. Image preprocessing resizes frames to 11 and normalizes with mean and standard deviation from COCO. Ultrasonic TOF samples are converted to distance through
12
where 13. Audio buffers are windowed and processed with voice activity detection.
The vision-language core follows a BLIP-inspired architecture. The image encoder is a 14-layer ViT-Base/384 with 15-dimensional embeddings and 16 attention heads; the text encoder is BERT-Base with 17 layers, 18 hidden units, and 19 self-attention heads. Cross-modal fusion is implemented through an image-grounded text encoder with cross-attention layers and an image-grounded text decoder initialized from a GPT-style transformer. The pre-training objectives are Image-Text Contrastive, Image-Text Matching, and Language Modeling. The summary specifies pre-training on approximately 20 million images and 21 million captions from COCO, Flickr30k, NLVR, and NoCaps, and fine-tuning on COCO detection splits and ICDAR text datasets, with AdamW, initial learning rate 22, weight decay 23, batch size 24, and 25 epochs.
Inference uses the on-board GPU in FP16 quantized form for less than 26 per frame. An object-detection head yields bounding boxes and class scores, while a text-recognition head uses a lightweight ResNet plus CTC decoder. Detected object centroids 27 and ultrasonic distances are fused into a 2D occupancy grid via a Kalman filter with state transition 28 and measurement 29, where 30 under a constant-velocity model. A voice-command parser invokes modules such as “What is that?” or “Navigate to X,” and navigation uses A* on a 31 occupancy grid with 32 resolution. End-to-end sensing to voice output completes within 33, with optional haptic buzzer feedback for imminent collisions at distance less than 34.
Reported benchmark metrics include COCO Image Captioning CIDEr 35, VQA v2 accuracy 36, NLVR accuracy 37, and in-device object detection mAP of approximately 38 on COCO val2017. The speech stack uses an on-device RNN-Transducer with real-time latency below 39, a context-free grammar with approximately 40 production rules, and Tacotron-2 plus WaveRNN at 41 with overall TTS latency below 42.
The pilot user study includes 43 visually impaired volunteers aged 44-45, performing corridor navigation, cluttered room exploration, sign reading, and object identification. Reported outcomes are object identification accuracy of 46, mean ultrasonic distance estimation error of 47, navigation success rate of 48, average completion time of 49 the sighted baseline, end-to-end system latency of 50, and overall user satisfaction of 51 on a 52-53 Likert scale.
In this form, NaviSense is a wearable multimodal perception-and-guidance system. Its defining technical feature is the joint use of deep vision-language inference, ultrasonic ranging, speech interaction, and local path planning on head-mounted hardware.
5. NaviSense as a mobile application for object retrieval by blind and low-vision users
In "NaviSense: A Multimodal Assistive Mobile application for Object Retrieval by Persons with Visual Impairment" (Sridhar et al., 23 Sep 2025), NaviSense is a mobile assistive system that combines conversational AI, a vision-LLM, AR, LiDAR, and synchronized audio-haptic feedback for open-world object detection with real-time guidance. The app runs on an iPhone 16 Pro with A18 chipset, 54 unified memory, iOS 17, a 55 MP wide camera, and a LiDAR depth scanner. The software stack uses ARKit for pose tracking and depth fusion, Core Haptics for vibration patterns, Apple Speech for on-device ASR and TTS, Moondream 2B for cloud VLM inference, and GPT-4o-mini for intent parsing and confirmation. An Apple M2 MacBook Pro hosts containerized services for GPT-4o-mini and Moondream 2B via REST.
The interaction model is controlled by a finite state machine with transitions Idle 56 Listening 57 Processing 58 Speaking 59 Scanning 60 Guiding, and shaking the phone cancels the current task. A user request is transcribed on-device and sent via HTTPS to GPT-4o-mini. Approximately once per second, an image frame is sent to Moondream 2B together with a prompt such as “find X,” and the model returns 2D bounding boxes of candidate objects. Once the target is detected, ARKit maps the pixel centroid 61 to a 3D point 62 using LiDAR depth, and guidance enters a closed-loop phase.
The spatial formulation includes Euclidean distance
63
a guidance vector
64
and angular deviation
65
The implementation description also defines the unit guidance vector 66. Audio feedback updates at approximately 67 with messages such as “turn left slightly” or “move forward,” while haptic pulse rate increases as distance decreases and angular alignment improves. Example piecewise feedback values are approximately 68 for 69, approximately 70 for 71, and approximately 72 for 73. When 74 exceeds a threshold of approximately 75, lateral pulses or stereo-panned audio indicate turning.
The system is explicitly open-world: it requires no pre-scanning and no category restriction. The example interaction includes clarification by the LLM when multiple candidate objects are present, such as distinguishing a blue ceramic coffee mug from a metal travel mug. Feedback ceases when 76 or upon pickup gesture, and the system returns to Listening.
Latency figures are given separately for components and the full interaction loop. ASR and TTS run locally in less than 77; cloud VLM round-trip averages approximately 78 with 79; the guidance loop is tuned to 80. The evaluation uses a within-subject comparison with 81 blind or low-vision participants, of whom 82 were blind and 83 low vision, aged 84-85 with mean 86. Comparison systems are NaviSense, Be My AI, and Ray-Ban Meta Glasses. Three target objects were placed on a four-tier shelf, with 87 trials per object and system, yielding 88 trials per participant.
The quantitative results are reported as follows. Search time is 89 for Be My AI, 90 for Meta Glasses, and 91 for NaviSense. Guidance time is 92, 93, and 94, respectively. Total time is 95, 96, and 97. Undesired touches are 98, 99, and 00. Accuracy is 01, 02, and 03. Repeated-measures ANOVA yields 04 for search time and 05 for total time, while guidance time is non-significant. The Friedman test on errors gives 06. Overall preference is 07 for NaviSense, 08 for Be My AI, and 09 for Meta, with Friedman 10. A supplementary evaluation on 11 common household items and 12 frames reports 13 detection accuracy with 14 CI 15.
This implementation represents the most explicit use of NaviSense as a named assistive product. Its distinctive contribution is the combination of open-world language-conditioned detection with AR/LiDAR localization and last-metre audio-haptic guidance on commodity mobile hardware.
6. Cross-domain technical characteristics and significance
Across the four uses of NaviSense, several architectural regularities recur. First, each system is multimodal: the RFSoC receiver couples direct and ground-reflected NavIC channels; the ultrasound framework couples IR pose tracking with volumetric imaging; the headset combines RGB, ultrasonics, microphones, and IMU; and the mobile application combines camera, LiDAR, ARKit pose, ASR/TTS, and cloud models. Second, each system uses an explicit spatial transform or registration layer: RF delay-Doppler cross-ambiguity, rigid transforms 16 and 17, Kalman-filtered occupancy-grid state, or ARKit pose and depth projection.
Third, all implementations close the loop from sensing to action. In the RF system, the output is a delay-Doppler map for passive radar extraction. In ultrasound, it is live probe-position feedback and re-aligned TIC quantification. In the headset, it is verbal scene description, path planning, and obstacle alerts. In the mobile app, it is synchronized audio-haptic object retrieval guidance. This suggests that NaviSense functions less as a modality-specific technology than as an integration pattern in which localization and sensing are operationally inseparable.
The evaluation criteria are likewise domain-specific but structurally parallel. The RF platform reports range-error RMSE, Doppler-error RMSE, and SNR robustness. The ultrasound system reports repositioning error, displacement statistics, and ICC for rBF and rBV. The headset reports object-identification accuracy, navigation success, and latency. The mobile app reports search time, total time, undesired touches, retrieval accuracy, and subjective ratings. A plausible implication is that NaviSense systems are best understood through task-level closed-loop performance rather than through isolated component benchmarks.
The available literature therefore presents NaviSense not as a single canonical platform, but as a family of systems in which sensing, navigation, and feedback are co-designed for specific operational settings. In RF sensing, this yields compact passive radar on RFSoC; in interventional ultrasound, tracked alignment and repeatable perfusion quantification; and in accessibility, open-world perception with actionable guidance for blind and low-vision users.