---
title: 'SpaRTAN: Differentiable Fixed-Sparsity Training'
url: https://www.emergentmind.com/topics/spartan
type: topic
---

# SpaRTAN: Differentiable Fixed-Sparsity Training

SpaRTAN denotes a unified algorithm for training neural networks to a fixed sparsity level $k$ by smoothly interpolating between two extremes—hard-masking with steep magnitude pruning and dense “dual-averaging” updates that maintain gradient flow to all weights. In "Spartan: Differentiable Sparsity via Regularized Transportation" [2205.14107], the method is defined by two coupled mechanisms: soft top-$k$ masking via a regularized optimal-transportation formulation, and dual-averaging-style parameter updates with hard sparsification in the forward pass. This design yields an exploration-exploitation tradeoff over sparsity patterns, supports unstructured, block-structured, and cost-sensitive sparsity allocation, and is reported to produce highly sparse ImageNet-1K models with less than $1\%$ absolute top-1 degradation relative to dense training [2205.14107].

## 1. Conceptual definition and limiting regimes

SpaRTAN, expanded in the source as **SParsity via Regularized Transportation And dual aNalyzing**, maintains a dense weight vector $\theta \in \mathbb{R}^d$ while enforcing a predetermined sparsity budget. At iteration $t$, it computes a soft mask, forms a masked weight vector, applies a hard pruning operator for the forward pass, and updates the dense parameters through a Jacobian induced by the soft mask. The central objective is not merely to prune after training, but to train directly toward a fixed sparse support while preserving gradient flow to all weights during earlier stages of optimization [2205.14107].

The method is explicitly framed as an interpolation between two known regimes. When $\beta = 0$, the soft mask is constant, $m_t \equiv 1_d$, and SpaRTAN reduces to Top-KAST, described in the source as pure dual-averaging with hard forward pruning. As $\beta \to \infty$, the soft mask converges to a hard indicator and the method recovers iterative magnitude pruning (IMP). This parameterized transition is fundamental to the algorithm’s interpretation: low sharpness favors exploration of sparsity patterns, whereas high sharpness increasingly fixes the support and emphasizes parameter optimization on that support [2205.14107].

A plausible implication is that SpaRTAN is best understood as a training-time sparsification framework rather than a post hoc compression heuristic. The source presents it as a unified algorithmic family rather than a narrow implementation for one model class, since the same machinery is reused for per-parameter, blockwise, and linear-cost-constrained sparsity [2205.14107].

## 2. Core mechanics: soft top-$k$ masking and dual-averaging updates

At each iteration, SpaRTAN computes
$$
m_t = \mathrm{softtopk}(|\theta_t|,\,k,\,\beta_t)\in[0,1]^d,
$$
then forms the soft-masked weights
$$
\sigma_k^{\beta_t}(\theta_t)=\theta_t\odot m_t,
$$
and applies hard pruning in the forward pass,
$$
\tilde\theta_t = \Pi_k(\sigma_k^{\beta_t}(\theta_t)).
$$
The parameter update is
$$
\theta_{t+1}=\theta_t-\eta_t\,\nabla\sigma_k^{\beta_t}(\theta_t)\,\nabla L(\tilde\theta_t).
$$
Because $\nabla_\theta \sigma_k^\beta$ is dense via the soft-mask Jacobian, even masked-out weights receive gradient updates [2205.14107].

This separation between forward sparsification and backward densification is the defining operational feature of the method. The forward pass uses exactly $k$ nonzeros after projection, so the loss is evaluated on a strictly sparse model. The backward pass, however, is mediated by the differentiable soft mask, allowing “dormant” weights to remain trainable. The source characterizes this as an exploration regime early in training and an exploitation regime later in training, with the transition controlled by the temperature $\beta$ [2205.14107].

At inference time, the soft mask is discarded and only the final hard projection is retained:
$$
\theta^*=\Pi_k(\theta_T),\qquad y=f(x;\theta^*).
$$
This makes the deployment rule considerably simpler than the training rule. A plausible implication is that the additional algorithmic complexity is concentrated in optimization, not in the resulting inference graph [2205.14107].

## 3. Regularized transportation formulation of the differentiable top-$k$ operator

The differentiable masking mechanism is derived from a budgeted linear program. For magnitudes $v_i = |\theta_i|$ and positive costs $c \in \mathbb{R}_{++}^d$, the objective is to maximize $v^T m$ subject to $0 \le m \le 1$ and $c^T m = k$. In transport form, this becomes an optimal-transportation problem over a matrix $Y \in \mathbb{R}_+^{d \times 2}$ with row sums fixed by $c$ and column sums fixed by the sparsity budget. The cost matrix is $C=[-v/c,\;0]$, and the mask is recovered as $m_i = Y_{i1}/c_i$ [2205.14107].

SpaRTAN introduces entropic regularization with parameter $1/\beta$:
$$
\min_{Y\ge0}\sum_{i,j}C_{ij}Y_{ij}-\tfrac1\beta H(Y)
\quad\text{s.t.}\quad
Y1_2=c,\;\;1_d^T Y=[k,\;1^T c-k].
$$
The resulting smooth problem is solved by Sinkhorn-Knopp iteration. The source gives an $O(d)$ computation in terms of dual variables $\mu \in \mathbb{R}$ and $\nu \in \mathbb{R}^d$:
$$
m = \exp\bigl(\tfrac{\beta\,v}{c}+\mu+\nu-\log c\bigr),
$$
$$
\nu\gets\log c-\log\bigl(1+\exp(\tfrac{\beta v}{c}+\mu)\bigr),
$$
$$
\mu\gets\log k -\log\sum_i\exp(\tfrac{\beta v_i}{c_i}+\nu_i).
$$
As $\beta \to \infty$, $m$ concentrates on the top-$k$ entries of $v/c$ [2205.14107].

The significance of this construction is that the sparsity budget is enforced through a differentiable relaxation of a combinatorial top-$k$ operator. This suggests that SpaRTAN’s notion of importance is not restricted to raw magnitude: by replacing the default $c=1_d$ with a nonuniform cost vector, the same operator ranks parameters by the ratio $v_i/c_i$ under a fixed linear budget [2205.14107].

## 4. Annealing schedule and exploration-to-exploitation dynamics

SpaRTAN divides training into three phases over $T$ epochs. During **Warmup** from $0$ to $0.2T$, the global sparsity budget linearly ramps from $1$ to target $k/d$, and the soft sharpness $\beta$ ramps from $1$ to $\beta_{\max}$. During the **Intermediate** phase from $0.2T$ to $0.8T$, the target sparsity is maintained while $\beta$ continues to increase according to
$$
\beta(t)=1+(\beta_{\max}-1)\,\tfrac{t-0.2T}{0.6T}.
$$
During **Fine-tuning** from $0.8T$ to $T$, the hard mask $m_{0.8T}$ is frozen and $\theta$ continues to be updated under that mask [2205.14107].

This schedule implements the exploration-exploitation interpretation already built into the masking operator. Early low-$\beta$ allows many weights to explore; later high-$\beta$ enforces a nearly fixed $k$-sparse support. The method therefore does not assume that the optimal sparse support is identifiable from initialization or from a single early pruning decision. Instead, it delays irreversible commitment until the soft approximation has been sharpened sufficiently [2205.14107].

A common misconception about fixed-sparsity training is that it must choose between exact sparse forward computation and gradient access to pruned parameters. SpaRTAN explicitly avoids that dichotomy. The source’s update rule uses a hard projection in the forward pass while backpropagating through the soft mask, so exact budgeted sparsity and dense gradient propagation coexist during training [2205.14107].

## 5. Supported sparsity models and cost-aware generalization

SpaRTAN is presented as supporting three sparsity allocation policies through the same regularized transportation operator. In the **unstructured** case, $c_i=1$ for all parameters and the mask is computed per parameter. In the **block-structured** case, parameters are grouped into blocks $B_j$, block magnitudes are defined by $v_j=\sum_{i\in B_j}|\theta_i|$, costs by $c_j=|B_j|$, and one mask is computed per block. In the **cost-sensitive** case, $c_i$ is chosen to reflect per-parameter FLOPs or memory, and the budget constraint $c^T m=k$ fixes total cost rather than raw nonzero count [2205.14107].

This unification is one of the most technically consequential aspects of the method. The source does not introduce separate pruning rules for each structure class; instead, it changes the cost model and aggregation granularity while leaving the soft-top-$k$ mechanism intact. A plausible implication is that the method can express deployment-aware sparsity objectives, provided the objective can be written as a linear model of per-parameter costs [2205.14107].

The same formulation also clarifies what SpaRTAN is not. It is not limited to unstructured magnitude pruning, and it is not intrinsically tied to a particular architecture family. The paper’s reported experiments span both convolutional and transformer models, and the formalism is organized around weight vectors, groups, and cost vectors rather than architecture-specific motifs [2205.14107].

## 6. Empirical performance and relation to other uses of the name

On ImageNet-1K classification, the paper reports the following results. For **ResNet-50** with $25.6$M parameters, **95% unstructured sparsity** yields **76.5% top-1** versus **77.5% dense**, and **97.5% sparsity** yields **74.2% top-1** versus **77.5%**, described in the source as a new SOTA. For **ViT-B/16** with $86$M parameters, **90% unstructured** sparsity yields **81.2% top-1** versus **80.1% dense DeiT-B**, and **90% block-structured ($16\times16$)** sparsity yields **79.1% top-1** versus **74.1% prior methods**. The source further states that these sparse models reduce inference FLOPs by $7\times$–$10\times$ while incurring at most $1\%$ absolute accuracy loss [2205.14107].

These results position SpaRTAN within the literature on sparse training at extreme sparsity, particularly where exact budget control and structured deployment constraints matter simultaneously. The reported ResNet-50 and ViT-B/16 numbers are used in the source to support two claims: first, that the method remains accurate at high sparsity; second, that the same optimization framework extends across unstructured and block-structured settings without changing its essential form [2205.14107].

The acronym itself is reused across unrelated literatures. **SPARTAN** has also named a sparse hierarchical memory for parameter-efficient Transformers [2211.16634], a sparse Transformer world model for local causation [2411.06890], a self-supervised spatiotemporal Transformer for group activity recognition [2303.12149], a scalable PARAFAC2 algorithm [1703.04219], a sparse robust addressable P2P overlay [1907.12028], a two-dimensional password interface [1905.08199], and graph-theoretic work on Spartan graphs [2504.06832]. In that broader naming landscape, **SpaRTAN** in [2205.14107] specifically denotes differentiable sparsity via regularized transportation and dual-averaging-style updates, rather than a general-purpose sparse architecture or a domain-specific acronym.

Source: https://www.emergentmind.com/topics/spartan