Papers
Topics
Authors
Recent
Search
2000 character limit reached

Successive Randomized Compression (SRC)

Updated 26 December 2025
  • SRC is a framework that uses randomized techniques to compress and approximate data, improving efficiency in statistical learning, tensor network contraction, and lossy coding.
  • It integrates methods like boosting, sparse regression codes, and Khatri–Rao sketches to optimize trade-offs between accuracy, complexity, and generalization.
  • Empirical results demonstrate that SRC methods offer significant speedups and tighter performance bounds compared to conventional approaches.

Successive Randomized Compression (SRC) is a framework and set of algorithmic primitives that leverage randomization for efficient compression, representation, and approximation in statistical learning, tensor network contraction, and lossy coding. SRC principles are instantiated in boosting theory, sparse regression codes for lossy compression, and randomized tensor contractions for quantum simulations. Each instantiation exploits randomized selection, sketching, or greedy column choice to balance trade-offs between accuracy, complexity, and generalization.

1. Formal Definitions and Variants

Randomized Sample Compression Schemes. Formally, a randomized compression scheme operates on labeled data (XƗY)n(X \times Y)^n and consists of:

  • A distribution DĪŗD_\kappa over deterministic encoding maps kk, mapping any input sample SS of size nn to a subsequence k(S,n)āŠ†Sk(S, n) \subseteq S of at most sns_n distinct points.
  • A deterministic reconstruction function ρ\rho that, given any compressed sequence UU of size ≤sn\leq s_n, returns a predictor DĪŗD_\kappa0.
  • Probabilistic consistency: For all DĪŗD_\kappa1 of size DĪŗD_\kappa2, the reconstructed predictor DĪŗD_\kappa3 classifies all points in DĪŗD_\kappa4 correctly with probability at least DĪŗD_\kappa5.
  • Stability: Conditioning on DĪŗD_\kappa6 ensures the conditional law of DĪŗD_\kappa7 matches that of DĪŗD_\kappa8, a property necessary for optimal generalization bounds (Cunha et al., 2024).

Sparse Regression Codes (SPARC, also termed SRC by some authors). In lossy compression, an SRC is constructed by partitioning a Gaussian design matrix into sections and selecting columns through a greedy randomized procedure for approximation (Venkataramanan et al., 2012).

Successive Randomized Compression in Tensor Networks. For matrix product operator (MPO) and matrix product state (MPS) contractions, SRC is a randomized, single-pass sweep that uses Khatri–Rao product sketches for efficient compression of the MPO–MPS product (CamaƱo et al., 8 Apr 2025).

2. Algorithmic Principles Across Domains

In Learning Theory and Boosting

  • SRC for boosting (SRC-Boost) repeatedly subsamples small batches from weighted distributions, trains weak learners on them, and compresses the voting classifier as the concatenation of subsamples DĪŗD_\kappa9 (Cunha et al., 2024).
  • The reconstruction function retrains the weak learners on the selected subsamples and aggregates their votes to form the final classifier.

In Lossy Compression (SPARC/SRC)

  • The codebook is built as a Gaussian design matrix kk0 partitioned into kk1 sections of kk2 columns each.
  • Greedy encoding: At each stage kk3, select column kk4 maximizing the inner product with the residual, and reduce the residual iteratively. The coefficients kk5 for each section are predetermined to minimize MSE (Venkataramanan et al., 2012).
  • The compressed representation consists of the vector of selected indices kk6 transmitted to the decoder.

In Tensor Network Contraction

  • Operating from right to left, SRC applies randomized QB decompositions to local "unfoldings" of the partially contracted network, utilizing Khatri–Rao product sketches at each step (CamaƱo et al., 8 Apr 2025).
  • Output cores are determined by thin QR factorizations of these sketches, and projection reduces the contraction size as the sweep progresses.

3. Theoretical Guarantees and Performance Bounds

Generalization Bounds in Learning

  • For any stable randomized compression scheme of size kk7 and consistency error kk8, the generalization error satisfies kk9 for some universal constant SS0, with probability at least SS1 over samples and random encoding (Cunha et al., 2024).
  • SRC-Boost achieves this bound for voting classifiers, improving over AdaBoost by reducing the dependence on SS2 from two factors to one.
Method Bound on SS3
SRC-Boost SS4
AdaBoost SS5

Lossy Compression (SPARC/SRC)

  • SPARC achieves the optimal rate–distortion function SS6 for i.i.d. Gaussian sources.
  • With complexity per sample SS7, the probability of excess distortion decays exponentially in SS8 for a fixed gap above SS9 (Venkataramanan et al., 2012).
  • Robustness: For ergodic sources of variance nn0, SPARC achieves the same distortion guarantee as for Gaussian sources.

Tensor Network Contraction

  • SRC with Khatri–Rao sketches recovers the compressed MPS product exactly (with probability one) if the bond is sufficiently large for exact representation.
  • Approximate errors for standard Gaussian sketches satisfy nn1.
  • Computational complexity is nn2 when nn3, and SRC is empirically up to nn4 faster than density-matrix and randomized contract-then-compress approaches (CamaƱo et al., 8 Apr 2025).

4. Trade-Offs: Complexity, Accuracy, and Adaptivity

Lossy Compression Trade-offs

  • With nn5 and nn6, encoding complexity is polynomial; reducing nn7 for lower complexity increases the gap above the rate–distortion limit by nn8 (Venkataramanan et al., 2012).
  • "Shannon codebook" (nn9) achieves k(S,n)āŠ†Sk(S, n) \subseteq S0 convergence but at exponential complexity and storage cost.

SRC in Tensor Networks

  • Single sweep compresses the MPO–MPS product; outperforming dense contraction schemes in speed and error when bond dimension k(S,n)āŠ†Sk(S, n) \subseteq S1 is moderate.
  • Empirically, oversampling and adaptive bond selection strategies yield optimal bond with minimized computation.
Method Typical Computational Cost
Basic contract-then-compress k(S,n)āŠ†Sk(S, n) \subseteq S2
SRC (Khatri–Rao sketches) k(S,n)āŠ†Sk(S, n) \subseteq S3

SRC-Boost Parameter Selection

  • Subsample size k(S,n)āŠ†Sk(S, n) \subseteq S4; number of rounds k(S,n)āŠ†Sk(S, n) \subseteq S5; compression size k(S,n)āŠ†Sk(S, n) \subseteq S6 (Cunha et al., 2024).
  • Stable randomization ensures the optimal single-logarithmic dependency on k(S,n)āŠ†Sk(S, n) \subseteq S7 in generalization error.

5. Robustness and Extensions

Robustness Across Sources

  • SPARC/SRC maintains optimal distortion for ergodic sources, provided sample variance converges (Venkataramanan et al., 2012).
  • SRC-Boost maintains stability and consistency even under random subsampling and weak learner error, via margin analysis and union-bound reasoning (Cunha et al., 2024).

Extensions in Tensor Networks

  • SRC with randomized sketches can be executed in parallel across sites, is non-iterative, and applicable in time-evolution, boundary contraction in PEPS, and thermal state computation.
  • Limitation: For small bond dimension or loose tolerance, deterministic methods may suffice or outperform randomized SRC; zip-up compression is sometimes faster at high approximation error (CamaƱo et al., 8 Apr 2025).

6. Empirical Evidence and Practical Guidelines

Lossy Compression Simulations

  • Empirically, SPARC matches k(S,n)āŠ†Sk(S, n) \subseteq S8 at low and moderate rates (especially for k(S,n)āŠ†Sk(S, n) \subseteq S9), and finite-sns_n0 gaps align with theoretical predictions.
  • Variants like norm-squared minimization at each step slightly improve performance over pure inner-product maximization (Venkataramanan et al., 2012).

Tensor Network Benchmarks

  • Synthetic problems (sns_n1, sns_n2): SRC yields up to sns_n3 speedup over traditional density-matrix and randomized contract-then-compress, with matching error.
  • Time evolution for quantum spin systems: SRC inserted in Krylov expansion achieves up to sns_n4 speedup over classical procedures.
  • Heuristics: Oversampling (e.g., sns_n5) and adaptive tolerance strategies yield optimal bond selection and minimal computation (CamaƱo et al., 8 Apr 2025).

SRC-Boost Illustration

  • For sns_n6 data points, sns_n7, sns_n8, the compressed classifier—retrained on the union of subsamples—matches the desired generalization bound sns_n9 in practice, and theoretical choices of ρ\rho0 drive ρ\rho1 small for optimal generalization (Cunha et al., 2024).

7. Comparison to Existing Methods and Future Directions

Domain Existing Method SRC Advantage
Boosting AdaBoost/margin bounds Reduces generalization bound from double-log to single-log in ρ\rho2
Lossy Compression Shannon codebook, scalar quantizer Polynomial complexity for near-optimal rate–distortion trade-off
Tensor Networks Density-matrix, contract-then-compress, zip-up Single-pass, fastest at tight tolerances, parallelizable

Immediate open problems include designing polynomial-complexity encoders whose distortion gap decays like the information-theoretic ρ\rho3, further analysis of randomized sketch structures in tensor networks, and exploring adaptive schemes for broader classes of ergodic and heavy-tailed sources (Venkataramanan et al., 2012, Cunha et al., 2024, Camaño et al., 8 Apr 2025).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Successive Randomized Compression (SRC).