---
title: Core-Elements Subsampling Method
url: https://www.emergentmind.com/topics/core-elements-subsampling-method
type: topic
---

# Core-Elements Subsampling Method

Searching arXiv for the cited papers to ground the article and confirm metadata.
{"query":"1709.04453 Visualization of Big Spatial Data using Coresets for Kernel Density Estimates", "max_results": 5}
{"query":"2509.18024 Core-elements Subsampling for Alternating Least Squares", "max_results": 5}
{"query":"2206.10240 Core-Elements for Large-Scale Least Squares Estimation", "max_results": 5}
Core-Elements Subsampling Method denotes a family of coreset, subsampling, and subset-selection procedures that replace a large dataset by a much smaller subset while attempting to preserve a target mathematical object or downstream computation. Depending on the setting, the “core elements” may be spatial points, training examples, rows of a design matrix, individual matrix entries, convexly weighted support points of an empirical measure, frame elements, or vertices of a simplicial complex. The preserved object may be a kernel density estimate, an ordinary least-squares estimator, an alternating least-squares update, a posterior distribution, a smooth divergence, or the homotopy type of a Vietoris–Rips complex [1709.04453], [2206.10240], [2509.18024], [2203.09675], [2504.20194], [2511.20954].

## 1. Conceptual scope and formal definitions

The literature does not use a single universal formalization. In spatial visualization, a coreset is a subset \(Q \subset P\) such that the kernel density estimate (KDE) of \(Q\) approximates that of \(P\) uniformly over all spatial locations:
\[
L_\infty(\kde_P,\kde_Q) := \max_{x \in \mathbb{R}^d}\big|\kde_P(x) - \kde_Q(x)\big| \le \varepsilon.
\]
This is a worst-case guarantee on the density field rather than a pointwise average guarantee [1709.04453].

In supervised learning, the coreset is a subset \(\mathcal{S} \subset \mathcal{T}\) of the training set
\[
\mathcal{T} = \{(\mathbf{x}_i, y_i)\}_{i=1}^{N}
\]
chosen so that training on \(\mathcal{S}\) is almost as good as training on \(\mathcal{T}\). In HyperCore, the global coreset is assembled classwise as
\[
\mathcal{S} = \bigcup_{c=0}^{C-1} \mathcal{S}_c,
\]
where each \(\mathcal{S}_c\) contains the retained in-class samples for class \(c\) [2509.21746].

In large-scale least squares, the phrase “Core-Elements” can refer either to a core subset of observations chosen by an optimal-design or robustness criterion, or to an element-wise sketch of the design matrix. The element-wise formulation constructs
\[
X^* = S \odot X,
\]
where \(S\) is a binary mask and only selected entries of \(X\) are retained; the corresponding estimator is
\[
\tilde\beta = (X^{*\top} X)^{-1} X^{*\top} y.
\]
This differs from row-wise subsampling because the retained information need not lie in a selected set of rows [2206.10240], [2105.01552].

In Bayesian inference, a coreset is a sparse nonnegative weight vector \(w\) with at most \(M\) nonzeros, defining a coreset posterior
\[
\pi_w(\theta) = \frac{1}{Z(w)} \exp\Big(\sum_{n=1}^N w_n f_n(\theta)\Big)\,\pi_0(\theta),
\]
and the construction problem is posed as minimizing
\[
\mathrm{KL}(\pi_w \,\|\, \pi)
\]
subject to the sparsity constraint on \(w\) [2203.09675].

In topological data analysis, the core is defined through domination in the \(\delta\)-neighborhood graph. A point \(x\) is \(\delta\)-dominated by \(x'\) if
\[
N_\delta(x) \subseteq N_\delta(x'),
\qquad
N_\delta(x) := \{y\in X \mid d(x,y)\le \delta\},
\]
and a \(\delta\)-core is a subsample with no dominated points. At scale \(\delta\), this is exactly the vertex set of a core of \(\mathrm{VR}(X,\delta)\) under strong collapses [2511.20954].

| Setting | Core elements | Preserved target |
|---|---|---|
| Spatial KDE | \(Q \subset P\) | \(L_\infty(\kde_P,\kde_Q)\le \varepsilon\) |
| Supervised learning | \(\mathcal{S}=\bigcup_c \mathcal{S}_c\) | Training quality relative to \(\mathcal{T}\) |
| Least squares | \(X^*=S\odot X\) or subset \(\mathcal{C}\) | OLS, design optimality, or robustness |
| Bayesian inference | Sparse weighted subset | \(\mathrm{KL}(\pi_w\|\pi)\) |
| TDA | \(X_\delta \subset X\) | Strong-collapse core of \(\mathrm{VR}(X,\delta)\) |

A broad synthesis is therefore supported by the sources: core-elements subsampling is best understood as a design principle rather than a single algorithm. The specific target of preservation determines both the definition of “core” and the admissible approximation error.

## 2. Selection mechanisms and algorithmic constructions

One recurrent construction is stratified selection after imposing geometric structure. For large spatial KDE visualization, the data are first ordered by a Z-order curve, then partitioned into contiguous blocks, and one point is randomly selected from each block. The method also supports a priority ordering \(S=\langle s_1,\dots,s_n\rangle\) such that every prefix \(\{s_1,\dots,s_k\}\) is a coreset of size \(k\). This makes the subsample resolution-adjustable without rerunning the selection procedure [1709.04453].

A second construction is deterministic element-wise masking. In large-scale least squares, core-elements are the largest-magnitude entries in each column of \(X\): for each column \(j\), the method keeps the \(r\) largest absolute values and zeros out the rest, producing \(X^*=S\odot X\). The same columnwise idea is transferred to alternating least squares for recommender systems: for each regression matrix \(\widetilde M_{I_i^U}^{(t)}\) or \(\widetilde U_{I_j^M}^{(t+1)}\), the method keeps only the top \(\lfloor |I_i^U|r \rfloor\) or \(\lfloor |I_j^M|r \rfloor\) entries per column, forming sparse sketches used inside ALS updates [2206.10240], [2509.18024].

A third construction is class-wise geometric filtering. HyperCore learns, for each class \(c\), a hypersphere model \(\phi_c(\mathbf{x};W_c)\) centered at the origin. The conformity score is the embedding norm
\[
d_i = \|\phi_c(\mathbf{x}_i)\|,
\]
and the class-specific threshold is chosen by maximizing Youden’s statistic
\[
J_c(\tau)=\mathrm{TPR}_c(\tau)-\mathrm{FPR}_c(\tau).
\]
The retained core elements for class \(c\) are precisely those in-class points with \(d_i \le \tau_c^*\) [2509.21746].

A fourth construction is probabilistic with post-selection reweighting. In fast Bayesian coresets, the subset is chosen by uniform subsampling, but the weights are refined by a quasi-Newton method based on covariance operators under the coreset posterior. In set-based stochastic subsampling, the first stage uses conditionally independent Bernoulli random variables to select candidate elements, and the second stage uses conditionally dependent autoregressive Categorical sampling with set attention to model pairwise interactions among candidates [2203.09675], [2006.14222].

A fifth construction is functional compression. CO2 derives a kernel from the second-order Hadamard expansion of a smooth divergence and then runs kernel compression through Nyström approximation and recombination. In this formulation the selected support points are convexly weighted and are chosen to minimize an MMD-like quadratic form associated with the divergence [2504.20194].

A sixth construction is topological domination pruning. For \(\delta\)-core subsampling, one iteratively removes points \(x\) for which \(N_\delta(x)\subseteq N_\delta(x')\) for some \(x'\), until no dominated points remain. On Vietoris–Rips complexes, this is exactly dominated-vertex deletion in a flag complex, hence an iterative strong collapse [2511.20954].

## 3. Approximation guarantees and preserved structure

The strongest guarantees in the surveyed literature are target-specific rather than generic. For spatial KDEs, the Z-order coreset satisfies
\[
|Q| = O\Big(\frac{1}{\varepsilon}\log^{2.5}\frac{1}{\varepsilon}\cdot \log\frac{1}{\delta}\Big),
\]
and with probability at least \(1-\delta\),
\[
\max_x |\kde_P(x)-\kde_Q(x)| \le \varepsilon.
\]
A random sample also yields an \(\varepsilon\)-KDE coreset with probability \(\ge 1-\delta\), but requires
\[
|Q_{\mathrm{RS}}| = O\Big(\frac{1}{\varepsilon^2}\log\frac{1}{\delta}\Big).
\]
The same paper also derives a safe threshold shift \(T'=T-\varepsilon\), ensuring that regions with \(\kde_P(x)\ge T\) are not omitted when thresholding the coreset KDE, and discusses \((\tau,\varepsilon)\)-nets as a complementary guarantee against visually broken high-density regions [1709.04453].

For element-wise least squares, the core-elements estimator is unbiased:
\[
E(\tilde\beta\mid X)=\beta.
\]
Its conditional error satisfies
\[
E(\|\tilde\beta-\beta\|^2\mid X)=\sigma^2\left\|(X^{*\top}X)^{-1}X^{*\top}\right\|_F^2,
\]
and a Taylor expansion yields the upper bound
\[
E(\|\tilde\beta-\beta\|^{2} \mid X) \leq \sigma^2 \left[ \operatorname{tr}\{(X^{\top}X)^{-1}\}
+ \|(X^{\top}X)^{-1}\|_2^2 \|L\|_F^2\right]\{1+O(\lambda_0)\},
\]
with \(L=X-X^*\). The same framework proves a coreset-like finite-sample guarantee:
\[
\|y-X \beta_{\rm OLS}\|^{2} \leq \|y-X \tilde\beta\|^{2} \leq (1+\epsilon)\|y-X \beta_{\rm OLS}\|^{2}
\]
whenever \(\|X-X^*\|_2\) is sufficiently small relative to \(\|X\|_2\) [2206.10240].

For alternating least squares, Core-ALS provides a per-regression \((1+\epsilon)\)-approximation in residual loss. For users,
\[
\big\|\boldsymbol{R}^\top(i, I_i^U) - \widetilde{\boldsymbol{M}_{I_i^U}^{(t)}}\widetilde{\boldsymbol{u}_i^{(t+1)}}\big\|^2
\le (1+\epsilon_m)\,
\big\|\boldsymbol{R}^\top(i, I_i^U) - \widetilde{\boldsymbol{M}_{I_i^U}^{(t)}}\widehat{\boldsymbol{u}_i^{(t+1)}}\big\|^2,
\]
with an analogous theorem for items, and the regularized loss remains monotonically non-increasing under the stated assumptions [2509.18024].

For Bayesian coresets, the quasi-Newton method is accompanied by a general high-probability bound on the KL divergence of the output coreset posterior, and the paper proves that in finite-dimensional span settings an exact approximation can exist with high probability once the number of sampled core points is large enough. In the asymptotic smooth setting, the minimum KL over the subsampled family is \(o(1)\) under the stated conditions [2203.09675].

For smooth divergences, CO2 proves that second-order Hadamard differentiability reduces coreset selection to MMD minimization. If the spectral tail satisfies \(T(m)=o(1/n)\), then one can construct \(P_m\) with
\[
D(P_m)=D(\mathbb{P}_n)+o_p(n^{-1}).
\]
For the Sinkhorn divergence, the paper verifies the required regularity and obtains the compression theorem
\[
S(\mathbb{P},P_m)=S(\mathbb{P},\mathbb{P}_n)+o_p(1/n)
\]
for \(m=\omega(\log n)\) [2504.20194].

For \(\delta\)-core subsampling in TDA, the guarantee is topological rather than statistical. If \(Y\subseteq X\) is a \(\delta\)-core, then \(\mathrm{VR}(Y,\delta)\) is obtained from \(\mathrm{VR}(X,\delta)\) by strong collapses, so the two complexes are homotopy equivalent at scale \(\delta\). Any two \(\delta\)-cores are unique up to \(\delta\)-equivalence [2511.20954].

Finally, in finite-frame subsampling, a random reweighted subframe with only \(O(m\log m)\) elements can preserve frame bounds with high probability, while deterministic BSS-style spectral sparsification reduces the number of elements to \(O(m)\) with controlled conditioning [2202.12625].

## 4. Computational profile

Core-elements methods often trade expensive full-data algebra for cheap preprocessing plus small dense solves. In spatial KDE visualization, preprocessing is \(O(n\log n)\) time and \(O(n)\) space; once the priority ordering is available, extracting a coreset of size \(k\) is a prefix operation, and KDE evaluation on a grid of \(m\) pixels costs \(O(mk)\) rather than \(O(mn)\) [1709.04453].

In large-scale least squares, the element-wise estimator requires only
\[
O(\mathrm{nnz}(X)+rp^2)
\]
operations, where \(r\) is the number of retained elements per column. This scaling is the central reason the method targets sparse and numerically sparse matrices [2206.10240].

In Core-ALS, the full algorithm over \(n_t\) iterations has complexity
\[
O\Big(n_f\big(\mathrm{nnz}(\boldsymbol{R})\cdot (r n_f)+n_f^2 n_m+n_f^2 n_u\big)\,n_t\Big),
\]
to be compared with standard ALS complexity
\[
O\Big(n_f^2\big(\mathrm{nnz}(\boldsymbol{R})+n_f n_u+n_f n_m\big)\,n_t\Big).
\]
The fast variant further reduces overhead by subsampling \(\widetilde{\boldsymbol{U}}^{(t)}\) and \(\widetilde{\boldsymbol{M}}^{(t)}\) only once per iteration [2509.18024].

In HyperCore, threshold search for each class requires sorting \(D_c^{\mathrm{in}}\), hence \(O(n_c\log n_c)\) time and \(O(n_c)\) memory per class, while the classwise hypersphere models are embarrassingly parallel [2509.21746].

For Bayesian coresets, one iteration of the quasi-Newton refinement scales as
\[
O(M^3 + SM^2 + SN),
\]
with space \(O(M^2)\). This is linear in dataset size \(N\) and sublinear in memory compared with methods that retain per-point auxiliary objects [2203.09675].

In TDA, the direct neighborhood computation for \(\delta\)-cores is \(O(n^2)\) naively, while spatial data structures reduce practical cost. The motivation is the downstream reduction in the number of simplices and therefore in persistent homology runtime [2511.20954].

## 5. Empirical behavior across domains

The reported empirical behavior is heterogeneous but informative. In spatial KDE approximation, the Utah highway example gives the following maximal errors for equal sample size: at size \(830\), random sampling error is \(0.035\) and coreset error is \(0.01\); at \(1890\), \(0.023\) versus \(0.005\); at \(5000\), \(0.014\) versus \(0.002\); and at \(10000\), \(0.010\) versus \(0.001\). The visual comparisons further show that random sampling can introduce spurious local peaks or holes, whereas the coreset KDE preserves the shape and intensity of density features more faithfully [1709.04453].

In noisy supervised learning, HyperCore is evaluated on ImageNet-1K and CIFAR-10. On CIFAR-10, at retained fractions between \(0.1\%\) and \(10\%\), it achieves up to \(5.6\%\) higher accuracy than the best baseline. Under \(10\%\) label corruption, the paper reports, for example, \(41.4\%\) accuracy at \(1\%\) retained versus \(36.8\%\) for the second-best CAL, and \(84.2\%\) at \(20\%\) retained versus \(83.2\%\) for Forgetting. In adaptive mode, the chosen pruning ratio increases with poisoning level while accuracy on the coreset exceeds full-data training; at \(10\%\) poisoning, full data gives \(90.8\%\) and HyperCore gives \(94.8\%\) with \(\alpha_{\text{HyperCore}}\approx 16.4\%\) [2509.21746].

In recommender-system matrix factorization, Core-ALS is consistently reported as closer to full ALS than uniform or leverage-based subsampling on RMSE, PRMSE, Hit@5, and NDCG@10. On a \(5000\times 5000\) rating matrix with \(n_f=50\), full ALS takes \(222.03\)s, while Core-ALS takes \(50.50\)s at \(r=0.10\), \(65.82\)s at \(r=0.15\), \(75.68\)s at \(r=0.20\), and \(83.00\)s at \(r=0.25\). On Netflix-sized data, full ALS takes \(606.15\)s, while Core-ALS ranges from \(137.35\)s at \(r=0.10\) to \(196.95\)s at \(r=0.25\) [2509.18024].

In least squares with sparse predictors, the synthetic and scRNA-seq studies report that core-elements consistently attains the lowest MSE and PMSE among the compared subsampling methods, and on the scRNA-seq dataset CORE and MOM-CORE are described as performing essentially as well as FULL and MOM-FULL while being dramatically faster [2206.10240].

At the same time, empirical evidence is not uniformly favorable for sophisticated subsampling. In logistic regression, the large benchmark in “A Coreset Learning Reality Check” finds that no method shows significant improvement over uniform subsampling for NLL or ROC, and only OSMAC variants show significant improvement over uniform for coefficient MSE, with substantial variability and heavier tails in some regimes [2301.06163].

In TDA, \(\delta\)-core subsampling reduces persistent-homology cost while preserving the dominant topological signatures. For the torus with \(2000\) points and \(\delta=1.4\), the paper reports bottleneck distances \(d_B(D_1,D'_1)=0.050543\) and \(d_B(D_2,D'_2)=0.035024\). For a heterogeneous cube with \(1500\) points, the full computation takes about \(217.55\) s, whereas a \(\delta\)-core with \(978\) points gives about \(43.14\) s total and a \(\delta\)-core with \(598\) points gives about \(26.16\) s total [2511.20954].

## 6. Limitations, misconceptions, and open directions

A first limitation is terminological. The least-squares review explicitly notes that the phrase “Core-Elements” does not appear verbatim there; instead, the corresponding idea is realized by deterministic optimal-design-based subsamples such as IBOSS, GKM, KYM, and robustness-oriented Lowcon. This suggests that “core-elements subsampling” is a cross-domain label imposed on a family of related but non-identical constructions [2105.01552].

A second limitation is that superiority over uniform subsampling is not automatic. The logistic-regression benchmark shows that sophisticated coreset and optimal subsampling methods often do not outperform simple uniform subsampling, especially once regularization is present. This directly counters a common misconception that any non-uniform core-selection rule is necessarily better than random choice [2301.06163].

A third limitation is problem dependence. HyperCore relies on learning meaningful per-class embeddings; the paper notes reduced robustness in extremely low-data or highly imbalanced classes, overhead from separate per-class models, the symmetric treatment of false positives and false negatives in Youden’s \(J\), and static thresholds after training [2509.21746]. For KDE coresets, assumptions on kernel smoothness and boundedness matter, and extreme bandwidths can affect constants in the coreset size [1709.04453]. For CO2, the present guarantees are asymptotic and depend on deriving a Hadamard operator for the divergence, which is analytically nontrivial [2504.20194]. For \(\delta\)-core subsampling, the exact guarantee is at the scale \(\delta\); across an entire filtration the persistence approximation is empirical rather than given by a closed-form bottleneck bound in the reported results [2511.20954].

The published extensions point in several directions. HyperCore mentions multi-label classification, imbalanced datasets, different embedding models, and continual learning; the KDE work discusses more aggressive coreset constructions, streaming data, and combinations with multi-resolution systems such as nanocubes and Gaussian cubes; Core-ALS notes compatibility with parallel ALS implementations, Sherman–Morrison-based acceleration, and improved iteration formulations; CO2 identifies finite-sample guarantees and improved MMD compression as open problems [2509.21746], [1709.04453], [2509.18024], [2504.20194].

Taken together, the literature supports a precise but plural conclusion. A Core-Elements Subsampling Method is not a single canonical algorithm. It is a family of principled constructions that select a small surrogate—possibly weighted, classwise, element-wise, or topologically reduced—so as to preserve a specified objective, estimator, divergence, or structure. The decisive technical question is therefore not whether a method is a “coreset,” but which object it preserves, by what mechanism, and under what approximation regime.

Source: https://www.emergentmind.com/topics/core-elements-subsampling-method