Sciweavers

RECOMB
2010
Springer

Genomic DNA k-mer Spectra: Models and Modalities

13 years 11 months ago
Genomic DNA k-mer Spectra: Models and Modalities
Background: The empirical frequencies of DNA k-mers in whole genome sequences provide an interesting perspective on genomic complexity, and the availability of large segments of genomic sequence from many organisms means that analysis of k-mers with non-trivial lengths is now possible. Results: We have studied the k-mer spectra of more than 100 species from Archea, Bacteria, and Eukaryota, particularly looking at the modalities of the distributions. As expected, most species have a unimodal k-mer spectrum. However, a few species, including all mammals, have multimodal spectra. These species coincide with the tetrapods. Genomic sequences are clearly very complex, and cannot be fully explained by any simple probabilistic model. Yet we sought such an explanation for the observed modalities, and discovered that low-order Markov models capture this property (and some others) fairly well. Conclusions: Multimodal spectra are characterized by specific ranges of values of C+G content and of Cp...
Benny Chor, David Horn, Nick Goldman, Yaron Levy,
Added 16 May 2010
Updated 16 May 2010
Type Conference
Year 2010
Where RECOMB
Authors Benny Chor, David Horn, Nick Goldman, Yaron Levy, Tim Massingham
Comments (0)