Papers
Topics
Authors
Recent
Search
2000 character limit reached

Lossless Compression for Convex ERM

Updated 7 February 2026
  • The paper demonstrates a lossless compression framework for convex ERM by generalizing color refinement to obtain exact instance compression across models.
  • It preserves all global optima by ensuring that gradients and aggregated feature sums remain constant across equitable partitions.
  • Empirical evaluations on diverse datasets show substantial reductions in data dimensions and runtime while maintaining numerical accuracy.

A lossless compression framework for convex empirical risk minimization (ERM) enables computational reductions in optimization tasks without loss of solution accuracy. The framework introduced by Zhu & Chen is built on color refinement—a combinatorial technique from graph theory—generalized to operate on data or kernel matrices associated with convex, differentiable ERMs. It yields exact instance compression for problems including linear and polynomial regression, (multi)class logistic regression, elastic-net regularization, and kernelized methods, while preserving all global optima and producing measurable practical speedups (Zhu et al., 31 Jan 2026).

1. Convex ERM: Formulation and Lossless Compression

Given feature vectors xi∈RDx_i \in \mathbb{R}^D, labels yiy_i, and positive weights viv_i, convex ERM in standard (primal) form seeks to minimize:

F(w,b)=∑i=1nvi f(xiTw+b;  yi)+R(w),F(w, b) = \sum_{i=1}^n v_i\,f(x_i^T w + b;\;y_i) + R(w),

where ff is convex and differentiable in its first argument, and RR is a convex regularizer (e.g., ridge or elastic-net). The kernelized ERM form utilizes:

  • Kernel matrix Kij=k(xi,xj)⪰0K_{ij}=k(x_i,x_j)\succeq0
  • Dual variables α∈Rn\alpha \in \mathbb{R}^n, bias b∈Rb \in \mathbb{R}

Objective:

F(α,b)=∑i=1nvi f((Kα+b1n)i;yi)+λ2αTKαF(\alpha, b) = \sum_{i=1}^n v_i\,f((K\alpha + b 1_n)_i ; y_i) + \frac{\lambda}{2} \alpha^T K \alpha

A compression mapping yiy_i0 is lossless if the compressed ERM has exactly the same set of solutions as the original problem.

2. Color Refinement and Equitable Partitions

Color refinement operates by representing the data or kernel matrix yiy_i1 as a weighted bipartite graph, where rows (yiy_i2) represent samples and columns (yiy_i3) features/kernels. Color refinement seeks a pair of partitions yiy_i4 (rows, columns) that are equitable:

  • For each sample-block yiy_i5 and feature-block yiy_i6:
    • yiy_i7 is constant over yiy_i8
    • yiy_i9 is constant over viv_i0

The algorithm iteratively refines blocks, initialized by grouping by label, feature statistics, or other coarse signatures, and splits along dimensional aggregates until the finest equitable partition is obtained. The process has time complexity viv_i1 for sparse viv_i2 or viv_i3 when viv_i4 is dense.

3. Compression Mapping and Theoretical Properties

Partition matrices viv_i5 encode feature blocks in viv_i6, with their row-scaled transposes viv_i7 averaging variables within each block. The compression mapping for the primal ERM is:

  • viv_i8
  • viv_i9
  • F(w,b)=∑i=1nvi f(xiTw+b;  yi)+R(w),F(w, b) = \sum_{i=1}^n v_i\,f(x_i^T w + b;\;y_i) + R(w),0
  • Regularizer F(w,b)=∑i=1nvi f(xiTw+b;  yi)+R(w),F(w, b) = \sum_{i=1}^n v_i\,f(x_i^T w + b;\;y_i) + R(w),1 in terms of compressed F(w,b)=∑i=1nvi f(xiTw+b;  yi)+R(w),F(w, b) = \sum_{i=1}^n v_i\,f(x_i^T w + b;\;y_i) + R(w),2

Main Losslessness Theorem: If F(w,b)=∑i=1nvi f(xiTw+b;  yi)+R(w),F(w, b) = \sum_{i=1}^n v_i\,f(x_i^T w + b;\;y_i) + R(w),3 is a reduction coloring (i.e., the gradients of F(w,b)=∑i=1nvi f(xiTw+b;  yi)+R(w),F(w, b) = \sum_{i=1}^n v_i\,f(x_i^T w + b;\;y_i) + R(w),4 and constraints are constant on color-blocks), then:

  • Any optimum F(w,b)=∑i=1nvi f(xiTw+b;  yi)+R(w),F(w, b) = \sum_{i=1}^n v_i\,f(x_i^T w + b;\;y_i) + R(w),5 in the original problem compresses to F(w,b)=∑i=1nvi f(xiTw+b;  yi)+R(w),F(w, b) = \sum_{i=1}^n v_i\,f(x_i^T w + b;\;y_i) + R(w),6 in the reduced problem.
  • Any F(w,b)=∑i=1nvi f(xiTw+b;  yi)+R(w),F(w, b) = \sum_{i=1}^n v_i\,f(x_i^T w + b;\;y_i) + R(w),7 optimal for the reduced problem lifts to F(w,b)=∑i=1nvi f(xiTw+b;  yi)+R(w),F(w, b) = \sum_{i=1}^n v_i\,f(x_i^T w + b;\;y_i) + R(w),8 optimal for the original.

This construction strictly refines any partition from automorphism symmetry ("lifted inference"), guaranteeing at least as much (or more) compression. Proofs leverage convexity (averaging in blocks lowers F(w,b)=∑i=1nvi f(xiTw+b;  yi)+R(w),F(w, b) = \sum_{i=1}^n v_i\,f(x_i^T w + b;\;y_i) + R(w),9) and preservation of feasibility via partition structure.

4. Specialization to Models and Algorithmic Implementation

The framework admits efficient algorithms for several widely used convex models:

Model Equitability Conditions Additional Conditions
Linear regression ff0 on ff1 ff2 is constant on each ff3
Binary logistic regression ff4 on ff5 ff6 const on ff7; ff8 const on ff9
Multiclass logistic RR0 on RR1 Replace RR2 term with RR3 counts
Elastic-net regression Same as linear Convexity (RR4 term) via Jensen argument
Kernel methods RR5 on RR6 RR7 const on RR8; RR9 const in block

Algorithmic steps consist of a color-refinement loop over Kij=k(xi,xj)⪰0K_{ij}=k(x_i,x_j)\succeq00, with model specialization via suitable signature choices to initiate partitions. The closed-form regression solution Kij=k(xi,xj)⪰0K_{ij}=k(x_i,x_j)\succeq01 is preserved upon reduction, and similar properties hold for kernel and logistic models.

5. Empirical Evaluation and Compression Performance

Empirical evaluation on five binary classification datasets (Titanic, skin_nonskin, phishing, a7a, breast-cancer) demonstrates:

  • Substantial reduction in sample and/or feature count without loss of accuracy:
Dataset Samples: Pre→Post Features: Pre→Post
Titanic 1309 → 1147 378 → 367
skin_nonskin 245057 → 51444 unchanged
phishing 11055 → 5849 unchanged
a7a 16100 → 13900 122 → 120
breast-cancer 683 → 675 unchanged
  • Cumulative runtime (reduction plus training) as a fraction of baseline training:
Dataset Runtime (% of baseline)
Titanic 16
skin_nonskin 28
phishing 21
a7a 91
breast-cancer 90
  • Prediction and objective values of compressed solutions are within numerical solver tolerance of the full problem.

6. Key Conditions and Complexity Analysis

The framework's theoretical underpinnings are anchored by precise criteria:

  • Gradient equivalence: Kij=k(xi,xj)⪰0K_{ij}=k(x_i,x_j)\succeq02 for all Kij=k(xi,xj)⪰0K_{ij}=k(x_i,x_j)\succeq03 in the same block
  • Equitable sums: Kij=k(xi,xj)⪰0K_{ij}=k(x_i,x_j)\succeq04 constant over Kij=k(xi,xj)⪰0K_{ij}=k(x_i,x_j)\succeq05; Kij=k(xi,xj)⪰0K_{ij}=k(x_i,x_j)\succeq06 constant over Kij=k(xi,xj)⪰0K_{ij}=k(x_i,x_j)\succeq07
  • Compression runtime: sparse Kij=k(xi,xj)⪰0K_{ij}=k(x_i,x_j)\succeq08 yields Kij=k(xi,xj)⪰0K_{ij}=k(x_i,x_j)\succeq09
  • Compressed problem dimension: α∈Rn\alpha \in \mathbb{R}^n0 samples by α∈Rn\alpha \in \mathbb{R}^n1 features (or α∈Rn\alpha \in \mathbb{R}^n2 dual variables in kernel methods)

By computing the coarsest equitable partition via color refinement, lossless instance compression is achieved for a general class of differentiable convex ERM problems. The approach is strictly at least as strong as group-theoretic symmetry identification and yields measurable end-to-end computational benefits while maintaining exact optimality (Zhu et al., 31 Jan 2026).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Lossless Compression Framework for Convex ERM.