---
title: ISTASTrack Hybrid RGB-Event Tracker
url: https://www.emergentmind.com/topics/istastrack
type: topic
---

# ISTASTrack Hybrid RGB-Event Tracker

Searching arXiv for the primary paper and closely related context.
arxiv_search.query{"search_query":"all:ISTASTrack OR ti:\"ISTA Adapter\" OR ti:\"RGB-Event Tracking\"","start":0,"max_results":10}
arxiv_search.query{"search_query":"ti:\"Tracking with A Common Tracking Software\" OR ti:\"Deep Learning-Based Extended Target Tracking in ISAC Systems\" OR ti:\"Temporal Graph Neural Network for ISAC Target Detection and Tracking\"","start":0,"max_results":10}
ISTASTrack is a transformer-based ANN-SNN hybrid tracker for RGB-Event tracking introduced in “ISTASTrack: Bridging ANN and SNN via ISTA Adapter for RGB-Event Tracking” [2509.09977]. It is designed for visual object tracking from paired RGB frames and event streams, combining a vision transformer for dense spatial context with a spiking transformer for sparse spatio-temporal dynamics. Its defining component is a model-based ISTA adapter for bidirectional feature interaction between the ANN and SNN branches, derived from sparse representation theory by unfolding the iterative shrinkage thresholding algorithm, together with a temporal downsampling attention module for aligning multi-step SNN features with single-step ANN features in the latent space. The reported evaluation on FE240hz, VisEvent, COESOT, and FELT states that the method achieves state-of-the-art performance while maintaining high energy efficiency [2509.09977].

## 1. Definition, nomenclature, and scope

ISTASTrack denotes the RGB-Event visual tracker introduced in 2025, not a generic tracking framework and not an ISAC extended-target tracker [2509.09977]. In particular, the query term “ISTASTrack” does not appear in the paper “Deep Learning-Based Extended Target Tracking in ISAC Systems”; that work explicitly names its method **ISACTrackNet**, and the available description states that *ISTASTrack* is almost certainly a misnaming or typo for **ISACTrack** or **ISACTrackNet** rather than a distinct algorithm [2504.00576].

The scope of ISTASTrack is single-object visual tracking in the standard template-search formulation. Given an initial bounding box in the first frame, the tracker estimates the target bounding box in every subsequent frame. The method operates on two heterogeneous sensing streams: conventional RGB frames, which provide spatial semantics, and event data, which provide asynchronous motion-sensitive signals with microsecond temporal resolution, very high dynamic range, low latency, and low power [2509.09977].

This positioning distinguishes ISTASTrack from experiment-independent tracking software in other fields. A plausible implication is that the name can invite confusion because unrelated tracking literature uses “tracking framework” in a much broader software-engineering sense, as in Acts for charged-particle tracking in high-energy physics [2007.01239]. In the specific published usage considered here, however, ISTASTrack is a multimodal neuromorphic vision tracker.

## 2. Problem formulation and motivation

The method is motivated by the observation that RGB-event fusion is attractive because RGB supplies high-quality spatial context under normal conditions, whereas events supply robust motion cues under fast motion, low light, high dynamic range, strong camera shake, and small-target scenarios where motion is crucial [2509.09977].

The paper frames conventional ANN-based trackers as poorly matched to event streams. Standard CNN- or transformer-based trackers are built for dense frame inputs sampled at fixed intervals, so events are typically accumulated into event frames or voxel grids aligned to the RGB frame rate. The stated limitations are temporal information loss, mismatch with sparse data, and inefficiency, since full-precision MAC-heavy networks do not exploit the binary and sparse nature of spikes [2509.09977].

The work also identifies limitations in existing RGB-Event trackers and earlier ANN-SNN hybrids. ANN-only RGB-event trackers either use early fusion or dual-branch ANN backbones with cross-attention, but are described as struggling with temporal misalignment, heterogeneous modality statistics, heavy attention-based fusion, and limited interpretability. Prior hybrid approaches are described as relying on ad-hoc attention, often using unidirectional fusion, shallow SNN backbones, and unprincipled ANN-SNN interfaces [2509.09977].

ISTASTrack is presented as a direct response to these issues. Its design objective is not merely to place an SNN beside an ANN, but to construct a bidirectional, sparse, layer-wise interface between them and to make temporal alignment explicit in the latent space. This suggests that the central claim of the method is architectural: the gain is attributed less to modality concatenation than to the way heterogeneous representations are coupled.

## 3. Hybrid architecture and tracking pipeline

The tracker uses a dual-branch ANN-SNN design. The RGB branch is a vision transformer inherited from OSTrack and processes the RGB template and RGB search image, while the event branch is a modified SpikingFormer that processes event template and event search sequences over \(T\) time steps [2509.09977].

On the RGB side, the input template \( \mathbf{Z}_{I_0} \in \mathbb{R}^{3 \times H_1 \times W_1} \) and search image \( \mathbf{X}_{I_i} \in \mathbb{R}^{3 \times H_2 \times W_2} \) are patch-embedded, augmented with positional encodings, and concatenated into transformer tokens \( \mathbf{x}^I \in \mathbb{R}^{N \times D} \). The backbone consists of 12 standard ViT blocks with MSA, MLP, GELU, layer norm, and residual connections. On the event side, the input event tensors are represented as \( \mathbf{E}_i \in \mathbb{R}^{T \times 3 \times H \times W} \), passed through a spiking tokenizer and then through 8 spiking transformer blocks, producing multi-step features \( \mathbf{x}^E \in \mathbb{R}^{T \times N \times D} \) [2509.09977].

The two branches are aligned layer-wise over the first 8 ANN layers. Each aligned layer contains four adapters: ANN \(\rightarrow\) SNN and SNN \(\rightarrow\) ANN at both the attention and MLP stages. The update equations are given as
\[
\begin{aligned}
\mathbf{x}_k^{E^1} &=
\mathbf{x}_k^{E}
+ \mathrm{SpikeMSA}(\mathbf{x}_k^{E})
+ \mathcal{A}^{I \rightarrow E}_{k1}(\mathbf{x}_k^{I}) \\
\mathbf{x}_k^{E^2} &=
\mathbf{x}_k^{E^1}
+ \mathrm{SpikeMLP}(\mathbf{x}_k^{E^1})
+ \mathcal{A}^{I \rightarrow E}_{k2}(\mathbf{x}_k^{I^1})
\end{aligned}
\]
and
\[
\begin{aligned}
\mathbf{x}_k^{I^1} &=
\mathbf{x}_k^{I}
+ \mathrm{MSA}(\mathbf{x}_k^{I})
+ \mathcal{A}^{E \rightarrow I}_{k1}(\mathbf{x}_k^{E}) \\
\mathbf{x}_k^{I^2} &=
\mathbf{x}_k^{I^1}
+ \mathrm{MLP}(\mathbf{x}_k^{I^1})
+ \mathcal{A}^{E \rightarrow I}_{k2}(\mathbf{x}_k^{E^1}) .
\end{aligned}
\]
These equations formalize bidirectional interaction rather than one-way conditioning [2509.09977].

The tracker is embedded in a Siamese template-search formulation,
\[
B_i = \mathcal{T}(\mathbf{Z}_{I_0}, \mathbf{Z}_{E_0}, \mathbf{X}_{I_i}, \mathbf{X}_{E_i}, B_0),
\]
with an OSTrack-style one-stage head for classification and bounding-box regression. The paper states that inference uses only the initial template, with no explicit online template update [2509.09977].

## 4. ISTA adapter and temporal downsampling attention

The defining mechanism of ISTASTrack is the ISTA adapter, which recasts cross-modal fusion as sparse coding in a shared latent space [2509.09977]. For a signal \( \mathbf{x} \in \mathbb{R}^{M \times N} \) and dictionary \( \mathbf{W} \in \mathbb{R}^{M \times D} \), the sparse-coding objective is written as
\[
\min_{\mathbf{z}} \; \frac{1}{2}\|\mathbf{x} - \mathbf{W}\mathbf{z}\|_2^2 + \lambda \|\mathbf{z}\|_1 .
\]
Within the tracker, the RGB feature \( \mathbf{x}^I \in \mathbb{R}^{M \times N} \) and the event feature \( \mathbf{x}^E \in \mathbb{R}^{T \times M \times N} \) are modeled using modality-specific dictionaries and shared sparse codes for cross-modal adaptation.

For RGB \(\rightarrow\) event, the formulation is described as
\[
\mathbf{x}^I = \mathbf{D}_I \mathbf{a}^{I},\quad
\mathbf{x}^{E'} = \mathbf{D}_E' \mathbf{a}^{I \rightarrow E},\quad
\mathbf{a}^I = \mathbf{a}^{I \rightarrow E},
\]
with an analogous construction for event \(\rightarrow\) RGB. The sparse codes are then obtained by unfolding ISTA iterations into learnable layers:
\[
\mathbf{a}^{k}
= h_{\boldsymbol{\theta}_k} \left(\mathbf{a}^{k-1} + \mathbf{P}_k (\mathbf{x} - \mathbf{D}_k \mathbf{a}^{k-1})\right),
\]
with initialization \( \mathbf{a}^0 = \mathbf{P}_0 \mathbf{x}_0 \), synthesis \( \mathbf{x}' = \mathbf{D}_K' \mathbf{a}^K \), and a soft-thresholding function with skip connection
\[
h_{\boldsymbol{\theta}_k}(\mathbf{u})
= \frac{1}{2}\left(
\text{sign}(\mathbf{u}) \cdot (\lvert \mathbf{u} \rvert - \boldsymbol{\theta}_k)_+
+ \mathbf{u}
\right).
\]
This is the paper’s main theoretical bridge between ANN and SNN representations [2509.09977].

A separate issue is temporal mismatch: the ANN branch is single-step, while the SNN branch produces \(T\)-step features. ISTASTrack addresses this with a Temporal Downsampling Attention module embedded inside the event \(\rightarrow\) RGB adapter. After iterative sparse coding, the multi-step latent codes
\[
\mathbf{a}_k \in \mathbb{R}^{T \times D \times N}
\]
are converted into a single-step code through learned temporal weights,
\[
\mathbf{a}_k^{(d)} = \sum_{t=1}^{T} \alpha_t \, \mathbf{a}_k^{(t)},
\]
where \( \boldsymbol{\alpha} \in \mathbb{R}^{1 \times T} \) is produced from pooled temporal statistics and a shared linear layer [2509.09977]. The paper reports that removing TDA degrades performance, which supports the view that temporal alignment is not a secondary implementation detail but part of the model’s core contribution.

## 5. Training protocol, benchmarks, and efficiency

Training uses the loss
\[
L = \lambda_1 L_{\text{focal}} + \lambda_2 L_1 + \lambda_3 L_{\text{GIoU}},
\]
with \( \lambda_1 = 2 \), \( \lambda_2 = 5 \), and \( \lambda_3 = 1 \), matching OSTrack [2509.09977]. The RGB branch is initialized from OSTrack pretrained weights, the SNN branch is partially initialized from SpikingFormer pretrained on classification, the optimizer is Adam, the learning rates are \(1 \times 10^{-5}\) for pretrained backbone parameters and \(1 \times 10^{-4}\) for other parameters, weight decay is \(1 \times 10^{-4}\), training runs for 60 epochs with learning-rate decay by \(\times 0.1\) after 48 epochs, batch size is 32, and the default SNN time-step count is \(T=3\) [2509.09977].

The evaluation uses four RGB-Event tracking benchmarks: FE240hz, VisEvent, COESOT, and FELT. The reported metrics are SR, OP50, OP75, PR, and NPR [2509.09977].

| Dataset | ISTASTrack result | Comparison noted in the data |
|---|---|---|
| FE240hz | SR 64.7, PR 92.2, OP75 43.3 | BAT: SR 64.3, PR 91.9 |
| VisEvent | SR 67.3, PR 84.6, OP50 81.2 | BAT: SR 67.1, PR 84.8 |
| COESOT | SR 75.7, PR 87.1, OP50 86.8 | BAT: SR 75.6, PR 85.7 |
| FELT | SR 55.2, PR 65.8 | BAT: SR 52.9, PR 62.6 |

The paper also gives an explicit energy model,
\[
E = E_{\text{MAC}} \sum_i OP_{\text{MAC}}^i + E_{\text{AC}} \sum_i (OP_{\text{AC}}^i \cdot fr^i),
\]
with \( E_{\text{MAC}} \approx 4.6 \) pJ and \( E_{\text{AC}} \approx 0.9 \) pJ in 45 nm CMOS [2509.09977]. The reported comparison is that a dual-ANN baseline uses 178.1 M parameters, 56.4 G MACs, and about 259.3 mJ per inference; ANN-SNN with \(T=1\) uses 158.0 M parameters, 31.0 G MACs, 28.5 G ACs, and about 146.0 mJ; ANN-SNN with \(T=3\) uses 34.7 G MACs, 85.5 G ACs, and about 169.1 mJ. The ISTA adapters contribute 0.32 M parameters, 0.16 G MACs, and 0.72 mJ, while TDA adds 72 parameters and about 0.0006 mJ [2509.09977]. This suggests that the adapter-based fusion is not the dominant source of compute or energy.

Ablation results further specify that bidirectional fusion is better than one-way fusion, that early insertion is better than late insertion, and that \(N=4\) adapters placed in layers 1–4 is the best setting across datasets. The default choice \(T=3\) is described as a good tradeoff between accuracy, complexity, and SNN energy [2509.09977].

## 6. Limitations, interpretation, and research significance

The main limitation stated for ISTASTrack is that SNN states are reset after each tracking step because of memory constraints, so the SNN does not carry long-term temporal context across frames [2509.09977]. The paper identifies this as a barrier to exploiting the full potential of spiking dynamics in long sequences. Future directions explicitly include stateful SNN tracking across long sequences, applying ISTA adapters and TDA to detection, segmentation, and action recognition, extending the approach to other modality pairs such as RGB-Depth, RGB-Thermal, and radar-camera, and investigating deployment on neuromorphic hardware [2509.09977].

From a methodological perspective, the paper positions ISTASTrack against two prevailing fusion paradigms. First, relative to ANN-only RGB-event trackers, it argues that events should not be treated merely as another frame-like modality. Second, relative to earlier ANN-SNN hybrids, it replaces heuristic fusion with an unfolded optimization module based on sparse coding. This suggests that the novelty of the method is as much representational as architectural: the shared sparse code is intended to function as a compact and interpretable latent interface rather than as a generic cross-attention block [2509.09977].

The broader significance of ISTASTrack is therefore twofold. At the application level, it reports state-of-the-art RGB-Event tracking results on FE240hz, VisEvent, COESOT, and FELT while retaining the energy-efficiency advantages associated with spiking computation. At the systems level, it provides a concrete design pattern for hybrid multimodal tracking: a dense semantic ANN branch, a spike-driven temporal SNN branch, a model-based bidirectional adapter derived from ISTA, and an explicit temporal alignment module in latent space [2509.09977]. A plausible implication is that this pattern is portable beyond tracking, particularly to multimodal perception problems where one stream is dense and frame-based while another is sparse and asynchronous.

Source: https://www.emergentmind.com/topics/istastrack