Papers
Topics
Authors
Recent
Search
2000 character limit reached

UniBench300: Unified MMVOT Benchmark

Updated 3 July 2026
  • UniBench300 is an evaluation benchmark that unifies multi-modal visual object tracking by integrating RGB, thermal, depth, and event data into a balanced testbed.
  • It resolves training-testing inconsistencies by replacing multiple modality-specific passes with a single, efficient one-pass inference, cutting evaluation time by roughly 27%.
  • The benchmark supports continual-learning strategies to counteract knowledge forgetting, enabling unified models to recover and even surpass separate-model performance.

UniBench300 is an evaluation benchmark designed to unify the assessment of multi-modal visual object tracking (MMVOT) models by incorporating RGB, thermal-infrared (T), depth (D), and event (E) data within a single, balanced testbed. Addressing a longstanding inconsistency between the parallel training of unified models and their evaluation on separate, modality-specific benchmarks, UniBench300 enables one-pass inference, reduces evaluation time by 27%, and underpins rigorous analysis of continual-learning–based unification strategies for MMVOT tasks (Tang et al., 14 Aug 2025).

1. Motivation and Objectives

Existing MMVOT pipelines typically mix modality sources—visible RGB, thermal-infrared, depth, and event streams—during training, with the aim of joint optimization across modalities. However, evaluation has historically remained fragmented, relying on independent RGBT, RGBD, and RGBE benchmarks. This workflow produces a reproducible performance drop at test time, attributed to the mismatch between joint training and separate testing objectives. UniBench300 is introduced to rectify this discrepancy by providing a single evaluation suite where all three tasks coexist. Its key objectives are:

  • Eliminating the train/test inconsistency in unified-model MMVOT evaluation,
  • Providing a balanced and challenging mixture of sequences from RGBT, RGBD, and RGBE modalities,
  • Enabling systematic investigation of continual learning (CL) strategies within the unification process.

2. Dataset Composition and Modalities

UniBench300 comprises exactly 300 video sequences totaling 368,100 frames, subdivided equally among the three major MMVOT task modalities:

Modality Count Source Benchmark(s) Selection Protocol
RGBT 100 LasHeR 100 hardest by avg. IoU of strong RGBT trackers
RGBE 100 VisEvent 100 hardest by mean IoU of TENet, SeqTrackV2
RGBD 100 DepthTrack + RGBD1K Union of 50 DepthTrack + 50 RGBD1K test sequences

Each frame in UniBench300 carries an axis-aligned bounding box annotation consistent with the original datasets’ protocols. Modalities are present in the following ensemble:

  • RGB: provides rich appearance
  • T (Thermal Infrared): robust against illumination changes
  • D (Depth): encodes 3D geometry
  • E (Event): captures high-frequency motion information

3. Evaluation Protocol and Metrics

UniBench300 employs canonical MMVOT metrics tailored to the constituent modalities:

  • For RGBT and RGBE:
    • Precision Rate (PR): The fraction of frames where the center error, did_i, is below a threshold, tt.
    • Normalized Precision Rate (NPR): PR with did_i scaled by the frame-diagonal.
    • Success Rate (SR): Area under SR-curve as a function of IoU threshold τ∈[0,1]\tau\in[0,1], with

    $\mathrm{SR}(\tau) = \frac{1}{N}\sum_{i=1}^{N} \mathbf{1}[\IoU_i \ge \tau],$

    where IoU is

    $\IoU_i = \frac{|B_{\text{pred},i} \cap B_{\text{gt},i}|}{|B_{\text{pred},i} \cup B_{\text{gt},i}|}.$

  • For RGBD:

    • Precision (Pr): Precision=TPTP+FP\mathrm{Precision} = \frac{\mathrm{TP}}{\mathrm{TP}+\mathrm{FP}}
    • Recall (Re): Recall=TPTP+FN\mathrm{Recall} = \frac{\mathrm{TP}}{\mathrm{TP}+\mathrm{FN}}
    • F-score: 2Precision×RecallPrecision+Recall2\frac{\mathrm{Precision}\times\mathrm{Recall}}{\mathrm{Precision}+\mathrm{Recall}}

All scores are reported as single area-under-curve (AUC) values (for PR, SR) or scalar F-scores, with no additional mAP computation required.

4. Unified Deployment and Computational Efficiency

Traditional unified MMVOT models have required three separate inference passes—one for each evaluation benchmark (LasHeR for RGBT, DepthTrack for RGBD, VisEvent for RGBE). UniBench300 merges these into a single contiguous suite of 300 sequences. Empirically, this reduces total evaluation time:

  • SymTrack: from 148 min (3 passes) to 108 min (1 pass), a 27% reduction,
  • ViPT*: from 127 min to 93 min, a 26.8% reduction.

This single-pass evaluation paradigm not only improves efficiency but also ensures metric consistency across modalities.

5. Baseline Performance and Continual Learning Strategies

Two state-of-the-art baselines—ViPT* and a symmetric variant SymTrack—are evaluated on UniBench300 using both parallel ("mixed") unification and serial unification with replay-based continual learning (CL):

Training ViPT* PR ViPT* SR SymTrack PR SymTrack SR
mixed 0.743 0.579 0.763 0.597
+CL 0.758 0.592 0.771 0.607
Δ\Delta +1.5% +1.3% +0.8% +1.0%

Serial unification augmented with CL enables unified models to match or surpass the performance of separately trained models, particularly in overall PR and SR. This demonstrates that continual learning can effectively mitigate performance degradation (i.e., "knowledge forgetting") as new modalities are integrated.

6. Knowledge Forgetting, Network Capacity, and Modality Discrepancy

Under mixed (parallel) unification, baseline trackers display reproducible degradation attributed to knowledge forgetting:

  • ViPT*: –3.30% (PR AUC) for RGBT, –2.57% for RGBD, –1.20% for RGBE
  • SymTrack: –2.20% (PR AUC) for RGBT, –1.13% for RGBD, –0.80% for RGBE

Replaying previously encountered tasks during training—i.e., implementing continual learning—substantially recoups this lost performance.

Network capacity negatively correlates with performance degradation: reducing the depth of modality-specific branches from 12 to 2 layers increases the SR-drop on LasHeR from 2.20% to 3.00%. This suggests network expressiveness is an essential parameter in mitigating knowledge interference between modalities.

Modalities exhibit asymmetric forgetting: RGBT suffers the largest performance loss, RGBD intermediate, and RGBE the smallest (RGBT > RGBD > RGBE). Latent embedding visualizations and cross-validation experiments indicate greater distributional distance between T and D than between T and E, and higher inter-modality distance induces more interference.

7. Analysis and Recommendations for Future Research

The unified, balanced nature of UniBench300 supports in-depth analysis of cross-modal interactions and knowledge transfer. Several recommended research directions emerge:

  • Adoption of distribution alignment methodologies, such as adversarial or optimal transport techniques, to reduce inter-modality gap prior to feature fusion,
  • Development of capacity-aware unification architectures that allocate parameters proportional to observed modality heterogeneity,
  • Exploration of advanced continual learning strategies, including replay and regularization-based techniques (e.g., Elastic Weight Consolidation, Learning without Forgetting), to further curb catastrophic forgetting as new modalities are integrated.

In summary, UniBench300 establishes a rigorous, efficient, and consistent evaluation platform for unified MMVOT models, supporting both benchmarking and methodological innovation in serial continual-learning–based unification (Tang et al., 14 Aug 2025).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to UniBench300.