Normalized Information Distance (0809.2553v1)

Published 15 Sep 2008 in cs.IR and cs.AI

Abstract: The normalized information distance is a universal distance measure for objects of all kinds. It is based on Kolmogorov complexity and thus uncomputable, but there are ways to utilize it. First, compression algorithms can be used to approximate the Kolmogorov complexity if the objects have a string representation. Second, for names and abstract concepts, page count statistics from the World Wide Web can be used. These practical realizations of the normalized information distance can then be applied to machine learning tasks, expecially clustering, to perform feature-free and parameter-free data mining. This chapter discusses the theoretical foundations of the normalized information distance and both practical realizations. It presents numerous examples of successful real-world applications based on these distance measures, ranging from bioinformatics to music clustering to machine translation.

Citations (101)

View on Semantic Scholar

Summary

We haven't generated a summary for this paper yet.

Summarize Now

Normalized Information Distance (0809.2553v1)

Summary

Related Papers