Papers
Topics
Authors
Recent
Search
2000 character limit reached

Memory-Light Compression of Second Moments

Updated 4 December 2025
  • The paper introduces SlimAdam, which leverages SNR-guided thresholding to compress second moments and achieve up to 98% memory savings without sacrificing stability.
  • The methodology utilizes homomorphic projection and parametric vectorization to compress high-dimensional covariance matrices and bilinear features efficiently.
  • The SMSO approach normalizes and vectorizes second-order statistics, enabling significant memory reduction in deep visual recognition and large-scale learning tasks.

Memory-light compression of second moments refers to methodologies that dramatically reduce the memory burden associated with storing or operating on second-moment statistics—such as covariance matrices, squared-gradient accumulators, or bilinear feature representations—without incurring a significant loss in information or algorithmic stability. These techniques have become indispensable across optimization in large-scale learning, scientific data analysis, and deep visual recognition, where full second-moment storage is often prohibitive.

1. Second Moments in Modern Algorithms and Their Memory Bottlenecks

Second-moment statistics arise ubiquitously—as moving averages of squared gradients in adaptive optimizers (Adam, RMSProp), as empirical covariance in bilinear pooling, or as correlation matrices in time series analysis. In Adam, the optimizer maintains parameter-wise moving averages:

  • First moment: Mt←β1Mt−1+(1−β1)gtM_t \leftarrow \beta_1 M_{t-1} + (1-\beta_1)g_t
  • Second moment: Vt←β2Vt−1+(1−β2)gt2V_t \leftarrow \beta_2 V_{t-1} + (1-\beta_2)g_t^2

Storing both MtM_t and VtV_t doubles the optimizer’s memory compared to first-order methods. For models with billions of parameters, the O(N)O(N) memory for VtV_t can severely restrict batch size or model scaling (Kalra et al., 3 Mar 2025). In vision models, naive second-order pooling leads to O(c2)O(c^2)-dimensional heads; e.g., a c=512c = 512 feature map would require handling 2.6×1052.6 \times 10^5 covariance elements per image (Yu et al., 2018).

In scientific computing, real-time correlation analysis (e.g., XPCS) involves calculation of G=XXTG = XX^T for Vt←β2Vt−1+(1−β2)gt2V_t \leftarrow \beta_2 V_{t-1} + (1-\beta_2)g_t^20, with Vt←β2Vt−1+(1−β2)gt2V_t \leftarrow \beta_2 V_{t-1} + (1-\beta_2)g_t^21. This results in memory and compute requirements that rapidly become unmanageable (Strempfer et al., 2024).

2. Signal-to-Noise Ratio–Guided Compression in Optimizers

SlimAdam introduces a layer-wise, statistically grounded criterion for compressing second-moment (variance) tensors in Adam. For a tensor Vt←β2Vt−1+(1−β2)gt2V_t \leftarrow \beta_2 V_{t-1} + (1-\beta_2)g_t^22 of shape Vt←β2Vt−1+(1−β2)gt2V_t \leftarrow \beta_2 V_{t-1} + (1-\beta_2)g_t^23, compressibility along dimension set Vt←β2Vt−1+(1−β2)gt2V_t \leftarrow \beta_2 V_{t-1} + (1-\beta_2)g_t^24 is quantified as

Vt←β2Vt−1+(1−β2)gt2V_t \leftarrow \beta_2 V_{t-1} + (1-\beta_2)g_t^25

where Vt←β2Vt−1+(1−β2)gt2V_t \leftarrow \beta_2 V_{t-1} + (1-\beta_2)g_t^26 and Vt←β2Vt−1+(1−β2)gt2V_t \leftarrow \beta_2 V_{t-1} + (1-\beta_2)g_t^27 denote mean and variance over axes in Vt←β2Vt−1+(1−β2)gt2V_t \leftarrow \beta_2 V_{t-1} + (1-\beta_2)g_t^28, and Vt←β2Vt−1+(1−β2)gt2V_t \leftarrow \beta_2 V_{t-1} + (1-\beta_2)g_t^29 are the complementary axes. If MtM_t0, entries along MtM_t1 are tightly clustered and the second-moment tensor can be replaced by their mean over MtM_t2 with minimal performance loss. The trajectory-averaged compressibility,

MtM_t3

is used for robust decision-making across training (Kalra et al., 3 Mar 2025).

Compression is guided by a threshold:

  • For each layer MtM_t4, for all MtM_t5, compute MtM_t6.
  • Let MtM_t7.
  • If MtM_t8 and MtM_t9 (default VtV_t0), compress VtV_t1 along VtV_t2; else leave uncompressed.

This procedure realizes up to VtV_t3 memory savings for the second moment in Transformers and ResNets at small learning rates, without compromising convergence or stability. Other low-memory Adam variants frequently fail at high learning rates, but SlimAdam's SNR analysis provides performance that matches full Adam across loss and accuracy curves (Kalra et al., 3 Mar 2025).

3. Homomorphic Linear Compression for Second-Moment Computation

For high-throughput correlation (e.g., in XPCS), the homomorphic compression framework projects the data matrix VtV_t4 into a lower-dimensional space:

VtV_t5

and directly computes the second moment on the compressed domain:

VtV_t6

If VtV_t7 is constructed from the top VtV_t8 right singular vectors of VtV_t9 (as in SVD), the transformation preserves all bilinear forms of the type O(N)O(N)0, i.e., O(N)O(N)1 is a homomorphism for such operations. In the lossless case, O(N)O(N)2, and O(N)O(N)3 is exact; lossy compression with O(N)O(N)4 yields the best rank-O(N)O(N)5 approximation in Frobenius norm. Critically, all calculations can proceed in the compressed space, with no need to decompress to full dimension, enabling O(N)O(N)6–O(N)O(N)7 speedup and memory reduction in practice (Strempfer et al., 2024).

4. Parametric and Statistical Second-Order Compression in Deep Networks

Statistically Motivated Second-Order Pooling (SMSO) compresses O(N)O(N)8 covariance matrices into O(N)O(N)9-dimensional Gaussian-distributed representations:

  1. Parametric Vectorization (PV):

VtV_t0

By design, VtV_t1 distributed (Wishart quadratic forms).

  1. Gaussianization: Each VtV_t2 is normalized using a square root and scaling such that for large VtV_t3, the result is approximately standard normal. A learnable affine VtV_t4 is applied akin to BatchNorm.

VtV_t5

All operations are differentiable. In the alternative, more efficient realization, PV can be mapped to 1x1 convolutions followed by global VtV_t6-pooling.

Empirically, SMSO achieves VtV_t7–VtV_t8 compression of bilinear heads, frequently with VtV_t9 between 64 and 2048, while outperforming both uncompressed and compact random-projection baselines on major vision datasets (Yu et al., 2018).

5. Error Control, Memory Scaling, and Practical Performance

The memory-light compression approaches are characterized by mathematically analyzable tradeoffs:

  • Optimizer compression (SlimAdam): Memory savings up to O(c2)O(c^2)0 for second moments; performance and stability indistinguishable from full Adam over diverse Transformer, ResNet, and ViT architectures (Kalra et al., 3 Mar 2025).
  • Homomorphic sketching: Using O(c2)O(c^2)1, the error in the second-moment estimate is bounded as per the Eckart-Young theorem,

O(c2)O(c^2)2

and is controlled by the neglected singular values. Choice of O(c2)O(c^2)3 targeting less than O(c2)O(c^2)4 residual energy typically suffices (Strempfer et al., 2024).

  • SMSO: For O(c2)O(c^2)5, O(c2)O(c^2)6, SMSO reduces the feature size O(c2)O(c^2)7 over bilinear pooling and outperforms all tested baselines; SMSO-64 features (O(c2)O(c^2)8) are O(c2)O(c^2)9 smaller than CBP-8192 but match or exceed their accuracy (Yu et al., 2018).

The computational cost is reduced correspondingly: slim optimizer states allow larger models or batches; homomorphic sketching reduces c=512c = 5120 to c=512c = 5121; SMSO reduces c=512c = 5122 to c=512c = 5123 in inference and storage.

6. Deployment Guidelines, Insights, and Limitations

  • SlimAdam: Select c=512c = 5124 as the SNR threshold. Rules for which layers and axes to compress should be derived at a small learning rate and retained for the full optimization. Inconsistent compressibility is observed in certain MLP and “Gate” layers, and token embeddings may necessitate uncompressed storage due to low SNR—critical for handling rare tokens (Kalra et al., 3 Mar 2025).
  • Homomorphic approaches: Optimal for second-moment and linear bilinear forms; do not extend to higher-order or nonlinear statistics. For stationary signals, fixed compression operators suffice, but drift in the signal subspace may warrant re-computation of the compression basis (Strempfer et al., 2024).
  • SMSO: The parametric vectorization approach preserves second-order statistical structure and regularizes the compressed features to Gaussianity, ensuring effective gradient flow and compatibility with first-order training paradigms. The value of c=512c = 5125 can be selected to match hardware or deployment constraints, with empirical evidence favoring modest c=512c = 5126 values for most applications. No SVD or matrix-power operations are needed at inference (Yu et al., 2018).

Potential limitations noted include reduced compressibility for fine-tuning or in highly nonstationary datasets (SlimAdam), possible loss of fidelity for low signal-to-noise eigenmodes (homomorphic sketching), and overfitting risk if c=512c = 5127 is chosen too large relative to dataset scale in SMSO. Re-evaluation of the compression rule or basis may be needed mid-training if gradient statistics shift significantly.

7. Summary Table: Major Memory-Light Second-Moment Compression Methods

Method Domain Mechanism Savings (Typical)
SlimAdam Optimizer (Adam) Layer- and axis-wise SNR, aggregate & threshold, mean/broadcast substitution Up to 98%
Homomorphic sketch XPCS, time series SVD-based projection, computations on compressed space Up to 10,000×
SMSO Deep vision pooling Parametric quadratic form, square-root, learnable affine 10–100×

These advances offer principled, practical solutions for compressing second-moment information, enabling large-scale learning and real-time analytics previously infeasible due to memory or computational constraints (Kalra et al., 3 Mar 2025, Strempfer et al., 2024, Yu et al., 2018).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (3)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Memory-Light Compression of Second Moments.