Papers
Topics
Authors
Recent
Search
2000 character limit reached

DSST: Efficient Scale-Adaptive Object Tracking

Updated 14 February 2026
  • The paper introduces DSST, which decouples translation and scale estimation using separate correlation filters to enhance tracking efficiency and robustness.
  • DSST employs a dedicated 1D scale filter updated in the Fourier domain, significantly reducing computational complexity compared to joint filtering methods.
  • Empirical evaluations on benchmarks like OTB-50 and VOT 2014 demonstrate DSST’s superior accuracy, real-time speed, and effective handling of scale variations.

The Discriminative Scale Space Tracker (DSST) is a scale-adaptive visual object tracking method within the tracking-by-detection paradigm. DSST achieves robust and accurate scale estimation by learning separate discriminative correlation filters for translation and scale, providing a computationally efficient solution compared to exhaustive joint search or full 3D filtering. The DSST framework explicitly models scale change using a dedicated 1D correlation filter, updating both translation and scale filters online via exponential averaging in the Fourier domain. DSST attains state-of-the-art tracking performance and real-time speeds, with superior robustness across standardized benchmarks (Danelljan et al., 2016).

1. Translation Filter: Formulation and Solution

The spatial translation filter in DSST is formulated as a multi-channel ridge regression/correlation filter. Given a set of feature patches xi∈RM×N×dx_i \in \mathbb{R}^{M \times N \times d} (e.g., HOG and gray) centered around the target in frame ii, and desired outputs yi∈RM×Ny_i \in \mathbb{R}^{M \times N} (typically Gaussian-shaped), the filter h=(h1,…,hd)h = (h^1,\ldots,h^d) is found as:

min⁡h∑i=1N∥h⋆xi−yi∥22+λ∑l=1d∥hl∥22\min_h \sum_{i=1}^N \lVert h \star x_i - y_i \rVert_2^2 + \lambda \sum_{l=1}^d \lVert h^l \rVert_2^2

where ⋆\star denotes circular correlation and λ>0\lambda > 0 is a regularization parameter. By exploiting the circulant structure and Parseval’s theorem, the optimization decouples in the Fourier domain into M ⁣× ⁣NM \! \times \! N independent d×dd \times d systems. For online tracking, exponential-windowed averaging of auto- and cross-spectra is employed:

Atl=(1−η)At−1l+η (G‾⊙Xtl)A_t^l = (1-\eta) A_{t-1}^l + \eta \, (\overline{G} \odot X_t^l)

ii0

where ii1, ii2, and ii3 denotes pointwise multiplication. The filter is updated as:

ii4

Translation in a new frame is determined by extracting a patch at the previous target position, computing its FFT, and correlating with the learned filter. The new target location is the argmax of the inverse FFT of the response (Danelljan et al., 2016).

2. Scale Filter: One-Dimensional Learning and Estimation

For scale estimation, DSST learns a separate one-dimensional filter ii5. At each frame ii6, ii7 patches ii8 of varying sizes ii9 are extracted (where yi∈RM×Ny_i \in \mathbb{R}^{M \times N}0 and yi∈RM×Ny_i \in \mathbb{R}^{M \times N}1). Each patch is resized to a fixed template and represented by a yi∈RM×Ny_i \in \mathbb{R}^{M \times N}2-dimensional feature vector yi∈RM×Ny_i \in \mathbb{R}^{M \times N}3. The scale filter minimizes:

yi∈RM×Ny_i \in \mathbb{R}^{M \times N}4

where yi∈RM×Ny_i \in \mathbb{R}^{M \times N}5 is a Gaussian response peaking at yi∈RM×Ny_i \in \mathbb{R}^{M \times N}6. In the frequency domain, the closed-form solution is:

yi∈RM×Ny_i \in \mathbb{R}^{M \times N}7

Updates follow the same exponential windows as for translation. Scale detection is accomplished by applying yi∈RM×Ny_i \in \mathbb{R}^{M \times N}8 to the yi∈RM×Ny_i \in \mathbb{R}^{M \times N}9 feature vectors and selecting the scale index as the argmax of the output scores (Danelljan et al., 2016).

3. Decoupled Translation and Scale Estimation

A primary innovation of DSST is the explicit separation of translation and scale estimation, as opposed to exhaustive or joint approaches. The sequence for each frame is as follows:

  • Translation estimation: Apply the spatial filter h=(h1,…,hd)h = (h^1,\ldots,h^d)0 to a patch at previous h=(h1,…,hd)h = (h^1,\ldots,h^d)1 to obtain position h=(h1,…,hd)h = (h^1,\ldots,h^d)2 via maximum response.
  • Scale estimation: Around h=(h1,…,hd)h = (h^1,\ldots,h^d)3, extract scale-sampled patches h=(h1,…,hd)h = (h^1,\ldots,h^d)4 and apply the scale filter h=(h1,…,hd)h = (h^1,\ldots,h^d)5 to obtain the new scale h=(h1,…,hd)h = (h^1,\ldots,h^d)6.
  • Model update: At h=(h1,…,hd)h = (h^1,\ldots,h^d)7, extract new patches for translation and scale, updating h=(h1,…,hd)h = (h^1,\ldots,h^d)8, h=(h1,…,hd)h = (h^1,\ldots,h^d)9 and their scale counterparts with learning rate min⁡h∑i=1N∥h⋆xi−yi∥22+λ∑l=1d∥hl∥22\min_h \sum_{i=1}^N \lVert h \star x_i - y_i \rVert_2^2 + \lambda \sum_{l=1}^d \lVert h^l \rVert_2^20 (Danelljan et al., 2016).

This decoupling avoids the high cost of full 3D or multi-resolution search and allows DSST to operate efficiently while preserving accuracy.

4. Frame-by-Frame Algorithm and Implementation

The framewise operations executed by DSST can be summarized:

  1. Translation
    • Extract a patch around min⁡h∑i=1N∥h⋆xi−yi∥22+λ∑l=1d∥hl∥22\min_h \sum_{i=1}^N \lVert h \star x_i - y_i \rVert_2^2 + \lambda \sum_{l=1}^d \lVert h^l \rVert_2^21, compute min⁡h∑i=1N∥h⋆xi−yi∥22+λ∑l=1d∥hl∥22\min_h \sum_{i=1}^N \lVert h \star x_i - y_i \rVert_2^2 + \lambda \sum_{l=1}^d \lVert h^l \rVert_2^22
    • Compute min⁡h∑i=1N∥h⋆xi−yi∥22+λ∑l=1d∥hl∥22\min_h \sum_{i=1}^N \lVert h \star x_i - y_i \rVert_2^2 + \lambda \sum_{l=1}^d \lVert h^l \rVert_2^23, obtain spatial score min⁡h∑i=1N∥h⋆xi−yi∥22+λ∑l=1d∥hl∥22\min_h \sum_{i=1}^N \lVert h \star x_i - y_i \rVert_2^2 + \lambda \sum_{l=1}^d \lVert h^l \rVert_2^24
    • Set min⁡h∑i=1N∥h⋆xi−yi∥22+λ∑l=1d∥hl∥22\min_h \sum_{i=1}^N \lVert h \star x_i - y_i \rVert_2^2 + \lambda \sum_{l=1}^d \lVert h^l \rVert_2^25
  2. Scale
    • For each min⁡h∑i=1N∥h⋆xi−yi∥22+λ∑l=1d∥hl∥22\min_h \sum_{i=1}^N \lVert h \star x_i - y_i \rVert_2^2 + \lambda \sum_{l=1}^d \lVert h^l \rVert_2^26, extract min⁡h∑i=1N∥h⋆xi−yi∥22+λ∑l=1d∥hl∥22\min_h \sum_{i=1}^N \lVert h \star x_i - y_i \rVert_2^2 + \lambda \sum_{l=1}^d \lVert h^l \rVert_2^27 at scale min⁡h∑i=1N∥h⋆xi−yi∥22+λ∑l=1d∥hl∥22\min_h \sum_{i=1}^N \lVert h \star x_i - y_i \rVert_2^2 + \lambda \sum_{l=1}^d \lVert h^l \rVert_2^28, compute min⁡h∑i=1N∥h⋆xi−yi∥22+λ∑l=1d∥hl∥22\min_h \sum_{i=1}^N \lVert h \star x_i - y_i \rVert_2^2 + \lambda \sum_{l=1}^d \lVert h^l \rVert_2^29
    • Compute ⋆\star0, ⋆\star1
    • Set ⋆\star2, ⋆\star3
  3. Model Update
    • Update ⋆\star4 and scale arrays ⋆\star5 at ⋆\star6 with newly extracted data (Danelljan et al., 2016).

5. Computational Complexity

DSST achieves substantial computational efficiency compared to joint or exhaustive multi-scale search methods. The dominant operations are:

  • Translation filter: ⋆\star7 per update/detection, for ⋆\star8 feature channels and patch size ⋆\star9.
  • Scale filter: λ>0\lambda > 00 for λ>0\lambda > 01 scale levels and feature dimension λ>0\lambda > 02.
  • Comparative costs:
    • Multi-resolution search applies translation filter at λ>0\lambda > 03 scales: λ>0\lambda > 04
    • Full 3D correlation filter: λ>0\lambda > 05
    • DSST, with separate translation and 1D scale filtering, is approximately λ>0\lambda > 06 times faster than multi-resolution and λ>0\lambda > 07 times faster than 3D filtering for reasonable values of λ>0\lambda > 08 and λ>0\lambda > 09 (e.g., M ⁣× ⁣NM \! \times \! N0) (Danelljan et al., 2016).

6. Experimental Results and Benchmarks

Key empirical findings from DSST and its compressed/faster variant fDSST:

Tracker (Configuration) Overlap Precision (%) Distance Precision (%) Speed (FPS)
Baseline DCF (translation only) 57.7 70.8 57.3
Multi-resolution DCF (5 scales) 65.2 74.8 16.9
Joint 3D DCF (33 scales) 63.2 72.1 1.46
Iterative Joint DCF 64.1 74.2 1.01
DSST (translation+scale, 33 scales) 67.7 75.7 25.4
fDSST (PCA+FFT interp) 74.3 80.2 54.3

On the OTB-50 benchmark, DSST achieves a 6.6% AUC improvement over the baseline, and fDSST outperforms the previous best (SAMF) by 2.6%. fDSST demonstrates real-time processing at 54 FPS and robust initialization, consistently outperforming KCF, SAMF, Struck, and similar trackers in both TRE and SRE protocols.

In the VOT 2014 challenge (25 videos), fDSST achieved the best composite rank (accuracy and robustness) among 38 trackers, attaining average overlap M ⁣× ⁣NM \! \times \! N1 and the fewest failures per sequence (M ⁣× ⁣NM \! \times \! N2). Attribute-based analysis showed fDSST winning in 7 out of 11 categories, notably excelling in scale variation, fast motion, and background clutter (Danelljan et al., 2016).

7. Notation Summary and Key Formulas

Key notation and formulae:

  • M ⁣× ⁣NM \! \times \! N3: M ⁣× ⁣NM \! \times \! N4-th feature channel of translation filter and input patch
  • M ⁣× ⁣NM \! \times \! N5: numerator and denominator spectra for translation filter at time M ⁣× ⁣NM \! \times \! N6
  • M ⁣× ⁣NM \! \times \! N7, M ⁣× ⁣NM \! \times \! N8, M ⁣× ⁣NM \! \times \! N9
  • d×dd \times d0
  • d×dd \times d1
  • Filter: d×dd \times d2; Detection: d×dd \times d3
  • Scale filter d×dd \times d4, samples d×dd \times d5, desired responses d×dd \times d6: update and spectrum calculation identical in form, with parameter d×dd \times d7

DSST’s primary methodological contribution is the decoupled scale-adaptive correlation filter framework, providing state-of-the-art accuracy and speed for generic object tracking (Danelljan et al., 2016).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Discriminative Scale Space Tracker (DSST).