2000 character limit reached
Suffix Arrays for Spaced-SNP Databases (1407.0114v1)
Published 1 Jul 2014 in cs.DS
Abstract: Single-nucleotide polymorphisms (SNPs) account for most variations between human genomes. We show how, if the genomes in a database differ only by a reasonable number of SNPs and the substrings between those SNPs are unique, then we can store a fast compressed suffix array for that database.