Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
156 tokens/sec
GPT-4o
7 tokens/sec
Gemini 2.5 Pro Pro
45 tokens/sec
o3 Pro
4 tokens/sec
GPT-4.1 Pro
38 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

Optimal Coreset for Gaussian Kernel Density Estimation (2007.08031v5)

Published 15 Jul 2020 in cs.DS, cs.CG, and cs.LG

Abstract: Given a point set $P\subset \mathbb{R}d$, the kernel density estimate of $P$ is defined as [ \overline{\mathcal{G}}P(x) = \frac{1}{\left|P\right|}\sum{p\in P}e{-\left\lVert x-p \right\rVert2} ] for any $x\in\mathbb{R}d$. We study how to construct a small subset $Q$ of $P$ such that the kernel density estimate of $P$ is approximated by the kernel density estimate of $Q$. This subset $Q$ is called a coreset. The main technique in this work is constructing a $\pm 1$ coloring on the point set $P$ by discrepancy theory and we leverage Banaszczyk's Theorem. When $d>1$ is a constant, our construction gives a coreset of size $O\left(\frac{1}{\varepsilon}\right)$ as opposed to the best-known result of $O\left(\frac{1}{\varepsilon}\sqrt{\log\frac{1}{\varepsilon}}\right)$. It is the first result to give a breakthrough on the barrier of $\sqrt{\log}$ factor even when $d=2$.

Citations (9)

Summary

We haven't generated a summary for this paper yet.