Papers
Topics
Authors
Recent
Search
2000 character limit reached

SpaRTAN: Differentiable Fixed-Sparsity Training

Updated 4 July 2026
  • SpaRTAN is a fixed-sparsity training framework that employs soft top-k masking and dual-averaging updates to maintain gradient flow during optimization.
  • It leverages a regularized transportation formulation to compute differentiable masks, enabling smooth interpolation between hard pruning and dense updates.
  • Empirical results on ImageNet-1K demonstrate that SpaRTAN supports both unstructured and block-structured sparsity with minimal accuracy degradation.

SpaRTAN denotes a unified algorithm for training neural networks to a fixed sparsity level kk by smoothly interpolating between two extremes—hard-masking with steep magnitude pruning and dense “dual-averaging” updates that maintain gradient flow to all weights. In "Spartan: Differentiable Sparsity via Regularized Transportation" (Tai et al., 2022), the method is defined by two coupled mechanisms: soft top-kk masking via a regularized optimal-transportation formulation, and dual-averaging-style parameter updates with hard sparsification in the forward pass. This design yields an exploration-exploitation tradeoff over sparsity patterns, supports unstructured, block-structured, and cost-sensitive sparsity allocation, and is reported to produce highly sparse ImageNet-1K models with less than 1%1\% absolute top-1 degradation relative to dense training (Tai et al., 2022).

1. Conceptual definition and limiting regimes

SpaRTAN, expanded in the source as SParsity via Regularized Transportation And dual aNalyzing, maintains a dense weight vector θ∈Rd\theta \in \mathbb{R}^d while enforcing a predetermined sparsity budget. At iteration tt, it computes a soft mask, forms a masked weight vector, applies a hard pruning operator for the forward pass, and updates the dense parameters through a Jacobian induced by the soft mask. The central objective is not merely to prune after training, but to train directly toward a fixed sparse support while preserving gradient flow to all weights during earlier stages of optimization (Tai et al., 2022).

The method is explicitly framed as an interpolation between two known regimes. When β=0\beta = 0, the soft mask is constant, mt≡1dm_t \equiv 1_d, and SpaRTAN reduces to Top-KAST, described in the source as pure dual-averaging with hard forward pruning. As β→∞\beta \to \infty, the soft mask converges to a hard indicator and the method recovers iterative magnitude pruning (IMP). This parameterized transition is fundamental to the algorithm’s interpretation: low sharpness favors exploration of sparsity patterns, whereas high sharpness increasingly fixes the support and emphasizes parameter optimization on that support (Tai et al., 2022).

A plausible implication is that SpaRTAN is best understood as a training-time sparsification framework rather than a post hoc compression heuristic. The source presents it as a unified algorithmic family rather than a narrow implementation for one model class, since the same machinery is reused for per-parameter, blockwise, and linear-cost-constrained sparsity (Tai et al., 2022).

2. Core mechanics: soft top-kk masking and dual-averaging updates

At each iteration, SpaRTAN computes

mt=softtopk(∣θt∣, k, βt)∈[0,1]d,m_t = \mathrm{softtopk}(|\theta_t|,\,k,\,\beta_t)\in[0,1]^d,

then forms the soft-masked weights

kk0

and applies hard pruning in the forward pass,

kk1

The parameter update is

kk2

Because kk3 is dense via the soft-mask Jacobian, even masked-out weights receive gradient updates (Tai et al., 2022).

This separation between forward sparsification and backward densification is the defining operational feature of the method. The forward pass uses exactly kk4 nonzeros after projection, so the loss is evaluated on a strictly sparse model. The backward pass, however, is mediated by the differentiable soft mask, allowing “dormant” weights to remain trainable. The source characterizes this as an exploration regime early in training and an exploitation regime later in training, with the transition controlled by the temperature kk5 (Tai et al., 2022).

At inference time, the soft mask is discarded and only the final hard projection is retained:

kk6

This makes the deployment rule considerably simpler than the training rule. A plausible implication is that the additional algorithmic complexity is concentrated in optimization, not in the resulting inference graph (Tai et al., 2022).

3. Regularized transportation formulation of the differentiable top-kk7 operator

The differentiable masking mechanism is derived from a budgeted linear program. For magnitudes kk8 and positive costs kk9, the objective is to maximize 1%1\%0 subject to 1%1\%1 and 1%1\%2. In transport form, this becomes an optimal-transportation problem over a matrix 1%1\%3 with row sums fixed by 1%1\%4 and column sums fixed by the sparsity budget. The cost matrix is 1%1\%5, and the mask is recovered as 1%1\%6 (Tai et al., 2022).

SpaRTAN introduces entropic regularization with parameter 1%1\%7:

1%1\%8

The resulting smooth problem is solved by Sinkhorn-Knopp iteration. The source gives an 1%1\%9 computation in terms of dual variables θ∈Rd\theta \in \mathbb{R}^d0 and θ∈Rd\theta \in \mathbb{R}^d1:

θ∈Rd\theta \in \mathbb{R}^d2

θ∈Rd\theta \in \mathbb{R}^d3

θ∈Rd\theta \in \mathbb{R}^d4

As θ∈Rd\theta \in \mathbb{R}^d5, θ∈Rd\theta \in \mathbb{R}^d6 concentrates on the top-θ∈Rd\theta \in \mathbb{R}^d7 entries of θ∈Rd\theta \in \mathbb{R}^d8 (Tai et al., 2022).

The significance of this construction is that the sparsity budget is enforced through a differentiable relaxation of a combinatorial top-θ∈Rd\theta \in \mathbb{R}^d9 operator. This suggests that SpaRTAN’s notion of importance is not restricted to raw magnitude: by replacing the default tt0 with a nonuniform cost vector, the same operator ranks parameters by the ratio tt1 under a fixed linear budget (Tai et al., 2022).

4. Annealing schedule and exploration-to-exploitation dynamics

SpaRTAN divides training into three phases over tt2 epochs. During Warmup from tt3 to tt4, the global sparsity budget linearly ramps from tt5 to target tt6, and the soft sharpness tt7 ramps from tt8 to tt9. During the Intermediate phase from β=0\beta = 00 to β=0\beta = 01, the target sparsity is maintained while β=0\beta = 02 continues to increase according to

β=0\beta = 03

During Fine-tuning from β=0\beta = 04 to β=0\beta = 05, the hard mask β=0\beta = 06 is frozen and β=0\beta = 07 continues to be updated under that mask (Tai et al., 2022).

This schedule implements the exploration-exploitation interpretation already built into the masking operator. Early low-β=0\beta = 08 allows many weights to explore; later high-β=0\beta = 09 enforces a nearly fixed mt≡1dm_t \equiv 1_d0-sparse support. The method therefore does not assume that the optimal sparse support is identifiable from initialization or from a single early pruning decision. Instead, it delays irreversible commitment until the soft approximation has been sharpened sufficiently (Tai et al., 2022).

A common misconception about fixed-sparsity training is that it must choose between exact sparse forward computation and gradient access to pruned parameters. SpaRTAN explicitly avoids that dichotomy. The source’s update rule uses a hard projection in the forward pass while backpropagating through the soft mask, so exact budgeted sparsity and dense gradient propagation coexist during training (Tai et al., 2022).

5. Supported sparsity models and cost-aware generalization

SpaRTAN is presented as supporting three sparsity allocation policies through the same regularized transportation operator. In the unstructured case, mt≡1dm_t \equiv 1_d1 for all parameters and the mask is computed per parameter. In the block-structured case, parameters are grouped into blocks mt≡1dm_t \equiv 1_d2, block magnitudes are defined by mt≡1dm_t \equiv 1_d3, costs by mt≡1dm_t \equiv 1_d4, and one mask is computed per block. In the cost-sensitive case, mt≡1dm_t \equiv 1_d5 is chosen to reflect per-parameter FLOPs or memory, and the budget constraint mt≡1dm_t \equiv 1_d6 fixes total cost rather than raw nonzero count (Tai et al., 2022).

This unification is one of the most technically consequential aspects of the method. The source does not introduce separate pruning rules for each structure class; instead, it changes the cost model and aggregation granularity while leaving the soft-top-mt≡1dm_t \equiv 1_d7 mechanism intact. A plausible implication is that the method can express deployment-aware sparsity objectives, provided the objective can be written as a linear model of per-parameter costs (Tai et al., 2022).

The same formulation also clarifies what SpaRTAN is not. It is not limited to unstructured magnitude pruning, and it is not intrinsically tied to a particular architecture family. The paper’s reported experiments span both convolutional and transformer models, and the formalism is organized around weight vectors, groups, and cost vectors rather than architecture-specific motifs (Tai et al., 2022).

6. Empirical performance and relation to other uses of the name

On ImageNet-1K classification, the paper reports the following results. For ResNet-50 with mt≡1dm_t \equiv 1_d8M parameters, 95% unstructured sparsity yields 76.5% top-1 versus 77.5% dense, and 97.5% sparsity yields 74.2% top-1 versus 77.5%, described in the source as a new SOTA. For ViT-B/16 with mt≡1dm_t \equiv 1_d9M parameters, 90% unstructured sparsity yields 81.2% top-1 versus 80.1% dense DeiT-B, and 90% block-structured (β→∞\beta \to \infty0) sparsity yields 79.1% top-1 versus 74.1% prior methods. The source further states that these sparse models reduce inference FLOPs by β→∞\beta \to \infty1–β→∞\beta \to \infty2 while incurring at most β→∞\beta \to \infty3 absolute accuracy loss (Tai et al., 2022).

These results position SpaRTAN within the literature on sparse training at extreme sparsity, particularly where exact budget control and structured deployment constraints matter simultaneously. The reported ResNet-50 and ViT-B/16 numbers are used in the source to support two claims: first, that the method remains accurate at high sparsity; second, that the same optimization framework extends across unstructured and block-structured settings without changing its essential form (Tai et al., 2022).

The acronym itself is reused across unrelated literatures. SPARTAN has also named a sparse hierarchical memory for parameter-efficient Transformers (Deshpande et al., 2022), a sparse Transformer world model for local causation (Lei et al., 2024), a self-supervised spatiotemporal Transformer for group activity recognition (Chappa et al., 2023), a scalable PARAFAC2 algorithm (Perros et al., 2017), a sparse robust addressable P2P overlay (Augustine et al., 2019), a two-dimensional password interface (Helble et al., 2019), and graph-theoretic work on Spartan graphs (Misra et al., 9 Apr 2025). In that broader naming landscape, SpaRTAN in (Tai et al., 2022) specifically denotes differentiable sparsity via regularized transportation and dual-averaging-style updates, rather than a general-purpose sparse architecture or a domain-specific acronym.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to SpaRTAN.