Papers
Topics
Authors
Recent
Search
2000 character limit reached

KT-NW: Supervised Kernel Thinning

Updated 21 February 2026
  • Kernel Thinning (KT-NW) is a supervised coreset construction technique that compresses input-target pairs using a tailored meta-kernel for efficient kernel regression.
  • It employs a deterministic split-and-swap strategy to reduce the maximum mean discrepancy, achieving a quadratic improvement over naïve subsampling approaches.
  • Empirical results demonstrate that KT-NW significantly lowers computational cost while maintaining competitive mean squared error compared to Full-NW methods.

Kernel Thinning (KT-NW) is a distribution compression and coreset construction technique in the context of kernel regression. It extends the Kernel Thinning (KT) framework to supervised learning, specifically for the Nadaraya–Watson (NW) estimator, providing substantial computational and statistical improvements over naïve i.i.d. subsampling. The KT-NW methodology derives its core efficacy from the use of a supervised “meta-kernel,” tailored to jointly compress both input features and targets, enabling high-fidelity, low-cardinality coresets for regression and related tasks.

1. Theoretical Framework of Kernel Thinning

Kernel Thinning (Dwivedi et al., 2021, Dwivedi et al., 2021) addresses the problem of choosing a size-mm subset (or weighted coreset) from an nn-point sample {xi}i=1n\{x_i\}_{i=1}^n such that the empirical averages of all functions in a target RKHS Hk\mathcal{H}_k are closely preserved. The worst-case integration error is measured by the Maximum Mean Discrepancy (MMD),

MMDk(P,Q)=sup⁡∥f∥Hk≤1∣EPf−EQf∣.\mathrm{MMD}_k(P,Q) = \sup_{\|f\|_{\mathcal{H}_k}\leq 1}\left| \mathbb{E}_P f - \mathbb{E}_Q f \right|.

Standard approaches like i.i.d. random thinning or uniform subsampling typically achieve O(n−1/4)O(n^{-1/4}) MMD error with subsampled coresets of size m≈nm\approx\sqrt{n}. By contrast, KT uses a deterministic-plus-randomized split-and-swap strategy guided by a split kernel k~\tilde k, yielding O(n−1/2log⁡n)O(n^{-1/2}\sqrt{\log n}) error (for many common kernels and distributions), representing a quadratic improvement (Dwivedi et al., 2021).

The KT workflow involves recursive halving (split phase) with random assignments favoring reduced MMD, followed by a greedy swap phase to further decrease the discrepancy with respect to the target kernel kk. The algorithm is kernel-agnostic, supporting Gaussian, Matérn, inverse multiquadric, Laplace, sinc, and other kernels (Dwivedi et al., 2021).

2. KT-NW: Supervised Kernel Thinning for Nadaraya–Watson Regression

The KT-NW estimator (Gong et al., 2024) generalizes KT to combine with the Nadaraya–Watson estimator for regression. Given i.i.d. data pairs nn0 with nn1, nn2, the conventional NW estimation is

nn3

requiring nn4 time per query.

KT-NW introduces a supervised “meta-kernel” on data pairs

nn5

and applies KT (specifically, the CompressPP algorithm) to produce a weighted coreset nn6 with weights nn7. The KT-NW estimator is then

nn8

reducing query-time cost to nn9. The meta-kernel structure enables both the numerator and denominator of NW regression to be approximated within the RKHS of {xi}i=1n\{x_i\}_{i=1}^n0.

3. Theoretical Guarantees and Multiplicative-Error Bounds

Multiplicative-Error Approximation

On the event that KT succeeds (with probability at least {xi}i=1n\{x_i\}_{i=1}^n1), the weighted coreset {xi}i=1n\{x_i\}_{i=1}^n2 satisfies, for any {xi}i=1n\{x_i\}_{i=1}^n3 in the RKHS,

{xi}i=1n\{x_i\}_{i=1}^n4

and a multiplicative error bound: {xi}i=1n\{x_i\}_{i=1}^n5 which provides strong relative-error control over the empirical means with respect to kernels and derivatives required for kernel regression (Gong et al., 2024).

Statistical Optimality

Assuming {xi}i=1n\{x_i\}_{i=1}^n6 and the data density {xi}i=1n\{x_i\}_{i=1}^n7 are {xi}i=1n\{x_i\}_{i=1}^n8-Hölder and smooth, and with coreset size {xi}i=1n\{x_i\}_{i=1}^n9, KT-NW achieves MSE: Hk\mathcal{H}_k0 matching the minimax rate of Full-NW up to logarithmic factors, and outperforming uniform Hk\mathcal{H}_k1-subsampling, which only attains Hk\mathcal{H}_k2 (Gong et al., 2024).

4. Algorithmic Implementation and Practical Considerations

KT-NW is implemented via the CompressPP coreset construction, with an overall runtime Hk\mathcal{H}_k3, storage Hk\mathcal{H}_k4, and Hk\mathcal{H}_k5 kernel evaluations per test point. This yields quadratic speedup over Full-NW (Hk\mathcal{H}_k6 per query) and matches the inference cost of naïve subsampling while achieving a superior statistical guarantee (Gong et al., 2024).

Coreset size selection is typically Hk\mathcal{H}_k7, with bandwidth and regularization determined by cross-validation. Empirical results confirm necessity of the supervised meta-kernel Hk\mathcal{H}_k8; alternative choices such as feature-only or concatenation-based kernels yield inferior MSE in regression benchmarks.

5. Empirical Performance and Benchmarking

Experiments across synthetic (Hk\mathcal{H}_k9, MMDk(P,Q)=sup⁡∥f∥Hk≤1∣EPf−EQf∣.\mathrm{MMD}_k(P,Q) = \sup_{\|f\|_{\mathcal{H}_k}\leq 1}\left| \mathbb{E}_P f - \mathbb{E}_Q f \right|.0, MMDk(P,Q)=sup⁡∥f∥Hk≤1∣EPf−EQf∣.\mathrm{MMD}_k(P,Q) = \sup_{\|f\|_{\mathcal{H}_k}\leq 1}\left| \mathbb{E}_P f - \mathbb{E}_Q f \right|.1, Wendland kernel), real regression (California Housing, MMDk(P,Q)=sup⁡∥f∥Hk≤1∣EPf−EQf∣.\mathrm{MMD}_k(P,Q) = \sup_{\|f\|_{\mathcal{H}_k}\leq 1}\left| \mathbb{E}_P f - \mathbb{E}_Q f \right|.2, MMDk(P,Q)=sup⁡∥f∥Hk≤1∣EPf−EQf∣.\mathrm{MMD}_k(P,Q) = \sup_{\|f\|_{\mathcal{H}_k}\leq 1}\left| \mathbb{E}_P f - \mathbb{E}_Q f \right|.3, Gaussian kernel), and large-scale classification (SUSY, MMDk(P,Q)=sup⁡∥f∥Hk≤1∣EPf−EQf∣.\mathrm{MMD}_k(P,Q) = \sup_{\|f\|_{\mathcal{H}_k}\leq 1}\left| \mathbb{E}_P f - \mathbb{E}_Q f \right|.4, MMDk(P,Q)=sup⁡∥f∥Hk≤1∣EPf−EQf∣.\mathrm{MMD}_k(P,Q) = \sup_{\|f\|_{\mathcal{H}_k}\leq 1}\left| \mathbb{E}_P f - \mathbb{E}_Q f \right|.5, Laplace kernel) demonstrate:

  • KT-NW achieves MSE and test/training times nearly identical to subsampled NW (ST-NW) and close to Full-NW, with costs reduced by orders of magnitude compared to full-data inference.
  • In regression, KT-NW's MSE is within a small factor of Full-NW, and KT-NW outperforms RPCholesky thinning by a logarithmic factor in runtime (Gong et al., 2024).
  • In classification, KT-NW processes 4M samples in 1.7s on a single core, with error in-between ST-NW and RPCholesky.
  • Ablation studies show the supervised meta-kernel is statistically optimal for supervised compression.

A typical table from (Gong et al., 2024) illustrates comparative performance:

Method MSE Train (s) Test (s)
Full-NW 0.414 11.11 0.70
ST-NW 0.574 0.002 0.009
RPCholesky 0.350 0.324 0.006
KT-NW 0.558 0.015 0.008

KT-NW compresses preprocessing from 11s to 0.015s and query from 0.7s to 0.008s with negligible accuracy trade-off.

KT-NW is one instantiation of a broader class of kernel-based distribution compression methods developed by Dwivedi & Mackey (Dwivedi et al., 2021, Dwivedi et al., 2021), who introduced multiple kernel-thinning variants:

  • KT-NW (Normalized-Kernel KT): Uses a normalized kernel MMDk(P,Q)=sup⁡∥f∥Hk≤1∣EPf−EQf∣.\mathrm{MMD}_k(P,Q) = \sup_{\|f\|_{\mathcal{H}_k}\leq 1}\left| \mathbb{E}_P f - \mathbb{E}_Q f \right|.6, ensuring all bounds are dimension-free in constant.
  • Target KT: Uses the target kernel directly as the split kernel for tightest single-function error bounds.
  • Power KT: Employs a fractional power split kernel to improve MMD rates for non-smooth kernels like Laplace and Matérn.
  • KT+: Combines target and power kernels for simultaneous best-case single-function and MMD allows.

All these variants are cast in a generalized split-and-swap template, providing a unified theory and practical set of tools for kernel coreset construction.

7. Practical Guidelines and Limitations

Practical implementation recommendations include the use of median bandwidth heuristics (MMDk(P,Q)=sup⁡∥f∥Hk≤1∣EPf−EQf∣.\mathrm{MMD}_k(P,Q) = \sup_{\|f\|_{\mathcal{H}_k}\leq 1}\left| \mathbb{E}_P f - \mathbb{E}_Q f \right|.7), cross-validation on held-out MMD, and always carrying an i.i.d. baseline for comparison. When targeting a single function, Target KT or KT-NW is preferred; for worst-case MMD, Power KT is optimal; for both objectives, KT+ is superior.

Complexity is MMDk(P,Q)=sup⁡∥f∥Hk≤1∣EPf−EQf∣.\mathrm{MMD}_k(P,Q) = \sup_{\|f\|_{\mathcal{H}_k}\leq 1}\left| \mathbb{E}_P f - \mathbb{E}_Q f \right|.8 in kernel evaluations for small MMDk(P,Q)=sup⁡∥f∥Hk≤1∣EPf−EQf∣.\mathrm{MMD}_k(P,Q) = \sup_{\|f\|_{\mathcal{H}_k}\leq 1}\left| \mathbb{E}_P f - \mathbb{E}_Q f \right|.9; memory can be reduced via low-rank decompositions. The method is robust across kernel choices, with fractional-power modifications expanding the feasible kernel class. The core limitation is scalability to extremely large O(n−1/4)O(n^{-1/4})0 without further engineering.

References

Definition Search Book Streamline Icon: https://streamlinehq.com
References (3)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Kernel Thinning (KT-NW).