Kernelized Locality-Sensitive Hashing for Semi-Supervised Agglomerative Clustering

Published 16 Jan 2013 in cs.LG, cs.CV, and stat.ML | (1301.3575v1)

Abstract: Large scale agglomerative clustering is hindered by computational burdens. We propose a novel scheme where exact inter-instance distance calculation is replaced by the Hamming distance between Kernelized Locality-Sensitive Hashing (KLSH) hashed values. This results in a method that drastically decreases computation time. Additionally, we take advantage of certain labeled data points via distance metric learning to achieve a competitive precision and recall comparing to K-Means but in much less computation time.