Papers
Topics
Authors
Recent
Search
2000 character limit reached

Residual Task Vector Quantization (RTVQ)

Updated 15 January 2026
  • RTVQ is a memory-efficient multi-task learning method that decomposes task vectors into a shared base and per-task residuals for precise quantization.
  • It employs asymmetric affine quantization to compress the narrow-range differences in task vectors, significantly reducing storage while controlling error.
  • Empirical results demonstrate that RTVQ achieves up to 92% storage reduction with negligible or improved performance on benchmarks like ViT and ResNet.

Residual Task Vector Quantization (RTVQ) is a method designed for highly memory-efficient model merging in multi-task learning frameworks. It addresses scalability limits imposed by traditional storage of multiple fine-tuned checkpoints by decomposing and quantizing task vector differences with high precision per bit. The technique leverages the statistically narrow range of task vectors to sustain or improve downstream performance while offering substantial reductions in storage requirements (Kim et al., 10 Mar 2025).

1. Formal Definition of Task Vectors and Narrow Range

Given a pre-trained model with parameters θpre∈Rn\theta_{\mathrm{pre}} \in \mathbb{R}^n, and a collection of fine-tuned checkpoints {θftt}t=1T\{\theta^{t}_{\mathrm{ft}}\}_{t=1}^T for tasks t=1,…,Tt=1, \ldots, T, the task vector for task tt is defined as

τt=θftt−θpre.\tau_t = \theta^{t}_{\mathrm{ft}} - \theta_{\mathrm{pre}}.

Empirical analysis demonstrates that the dynamic range of τt\tau_t is approximately an order of magnitude narrower than the range of θftt\theta^{t}_{\mathrm{ft}} itself (Figure 1, (Kim et al., 10 Mar 2025)). In the standard asymmetric quantization protocol, this property bounds the per-element rounding error

∣ϵ∣≤Δ/2,Δ=θmax⁡−θmin⁡2b−1,| \epsilon | \leq \Delta / 2, \quad \Delta = \frac{\theta_{\max} - \theta_{\min}}{2^b-1},

where the smaller range of τt\tau_t ensures significantly reduced quantization noise for bitwidth bb. This holds for all model types explored, including vision transformers (ViT-B/32, ViT-L/14) and convolutional nets (ResNet-50).

2. RTVQ Algorithmic Workflow

RTVQ quantizes each task vector in two parts: (1) a shared base vector, and (2) a per-task residual ("offset" vector). The process is:

  1. Compute Task-Average Weight:

{θftt}t=1T\{\theta^{t}_{\mathrm{ft}}\}_{t=1}^T0

  1. Form Base Vector:

{θftt}t=1T\{\theta^{t}_{\mathrm{ft}}\}_{t=1}^T1

  1. Quantize Base Vector (to {θftt}t=1T\{\theta^{t}_{\mathrm{ft}}\}_{t=1}^T2 bits):

{θftt}t=1T\{\theta^{t}_{\mathrm{ft}}\}_{t=1}^T3

  1. Error Correction:

{θftt}t=1T\{\theta^{t}_{\mathrm{ft}}\}_{t=1}^T4

  1. Per-Task Offset Vector:

{θftt}t=1T\{\theta^{t}_{\mathrm{ft}}\}_{t=1}^T5

  1. Quantize Offsets (to {θftt}t=1T\{\theta^{t}_{\mathrm{ft}}\}_{t=1}^T6 bits):

{θftt}t=1T\{\theta^{t}_{\mathrm{ft}}\}_{t=1}^T7

  1. Storage and Reconstruction:

Store only {θftt}t=1T\{\theta^{t}_{\mathrm{ft}}\}_{t=1}^T8 and reconstruct as {θftt}t=1T\{\theta^{t}_{\mathrm{ft}}\}_{t=1}^T9 at merge time.

Here, t=1,…,Tt=1, \ldots, T0 denotes t=1,…,Tt=1, \ldots, T1-bit asymmetric affine quantization.

3. Quantization Principles and Mathematical Formulation

The quantization function is given as follows for any tensor t=1,…,Tt=1, \ldots, T2: t=1,…,Tt=1, \ldots, T3 Applying to RTVQ, the base t=1,…,Tt=1, \ldots, T4 is quantized to t=1,…,Tt=1, \ldots, T5 bits, and each t=1,…,Tt=1, \ldots, T6 to t=1,…,Tt=1, \ldots, T7 bits. The reconstructed quantized task vector

t=1,…,Tt=1, \ldots, T8

The per-tensor scale and zero-point require negligible additional storage. The algorithm exploits the high quantization-sensitivity of the base, allocating more bits t=1,…,Tt=1, \ldots, T9, while offsets—which have extremely compressed range—can be quantized with tt0 as low as 2 bits with minimal impact.

4. Bitwidth Allocation and Memory Budget Adaptation

The total bit requirement for tt1 tasks is

tt2

Given a per-task memory budget tt3, the bit allocation satisfies tt4. Practical selection involves sweeping over tt5 and tt6 to balance quantization error and resource constraints, as measured by downstream performance. The design principle is that the higher-variance base encodes global features influencing all tasks, while low-variance per-task offsets can be compressed more aggressively.

5. Quantization Error Analysis

Under affine quantization, the maximum tt7 componentwise error for any vector tt8 is

tt9

and the mean squared (τt=θftt−θpre.\tau_t = \theta^{t}_{\mathrm{ft}} - \theta_{\mathrm{pre}}.0) error scales with τt=θftt−θpre.\tau_t = \theta^{t}_{\mathrm{ft}} - \theta_{\mathrm{pre}}.1. For RTVQ,

τt=θftt−θpre.\tau_t = \theta^{t}_{\mathrm{ft}} - \theta_{\mathrm{pre}}.2

by linearity. Empirical results (Fig. 3, (Kim et al., 10 Mar 2025)) show RTVQ reduces overall τt=θftt−θpre.\tau_t = \theta^{t}_{\mathrm{ft}} - \theta_{\mathrm{pre}}.3 quantization error per bit compared to direct single-stage Task Vector Quantization (TVQ) at ultra-low bitwidths (e.g., 2 bits). RTVQ’s error reduction becomes more pronounced as memory constraints intensify.

6. Empirical Performance and Storage Reduction

RTVQ demonstrates performance that matches or surpasses full-precision (FP32) and standard TVQ baselines while drastically reducing memory. Key results:

  • ViT-B/32, 8 tasks (classification)
    • FP32: 9.1 GB (69.2% accuracy)
    • TVQ (4 bits): 1.1 GB (69.1%)
    • TVQ (2 bits): 62% accuracy
    • RTVQ (τt=θftt−θpre.\tau_t = \theta^{t}_{\mathrm{ft}} - \theta_{\mathrm{pre}}.4, τt=θftt−θpre.\tau_t = \theta^{t}_{\mathrm{ft}} - \theta_{\mathrm{pre}}.5): 0.7 GB (70.2% accuracy, +1% relative to FP32)
  • Scaling to 14/20 tasks (ViT-B/32, ViT-L/14)
    • TVQ degradation shrinks with τt=θftt−θpre.\tau_t = \theta^{t}_{\mathrm{ft}} - \theta_{\mathrm{pre}}.6; RTVQ maintains within 1% accuracy of FP32 at τt=θftt−θpre.\tau_t = \theta^{t}_{\mathrm{ft}} - \theta_{\mathrm{pre}}.72.2 bits/task.
  • ResNet-50 NYUv2 (dense prediction)
    • 4 bit TVQ: Segmentation (mIoU), Depth (RelErr), and Normal (AngErr) within 0.1–0.5% of FP32.
    • 2 bit TVQ: Significant drop (normal error rises from 30.6° to 36°)
    • RTVQ (2+2 bits): Within 2° of FP32.
  • Storage Scaling (ViT-L/14, 20 tasks)
    • FP32: 22.8 GB
    • 4 bit TVQ: 2.9 GB
    • 2 bit TVQ: 1.4 GB
    • RTVQ (3+2 bits): 1.7 GB (τt=θftt−θpre.\tau_t = \theta^{t}_{\mathrm{ft}} - \theta_{\mathrm{pre}}.87.5% of FP32)

These results support that memory can be compressed to less than 8% of the original footprint with negligible or even improved merging performance (Kim et al., 10 Mar 2025).

7. Practical Implementation Considerations and Hyperparameters

RTVQ requires strict alignment of the pre-trained backbone τt=θftt−θpre.\tau_t = \theta^{t}_{\mathrm{ft}} - \theta_{\mathrm{pre}}.9 at both training and merge time; only task vectors are quantized. The quantization protocol employs per-tensor asymmetric scaling and zero-points. The error-correction step—adding τt\tau_t0 to the quantized base before computing offsets—is crucial when τt\tau_t1 is low, mitigating drift in reconstructions.

Recommended hyperparameters are τt\tau_t2, τt\tau_t3 as a default, with sweeps over τt\tau_t4 and τt\tau_t5 for specific accuracy/memory trade-offs. The merging frameworks (Task Arithmetic, Ties, EMR, AdaMerging) operate unchanged, substituting quantized task vectors for their full-precision counterparts. For sensitivity tuning, the τt\tau_t6 norm of quantization error, averaged across layers or tasks, is the metric of choice.

RTVQ capitalizes on the inherent statistical structure of task vector spaces, delivering storage reductions of up to τt\tau_t7 without accuracy loss on both classification and dense-prediction benchmarks, and establishes a new benchmark for scalable, memory-efficient model merging (Kim et al., 10 Mar 2025).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Residual Task Vector Quantization (RTVQ).