Article
Using reference-free compressed data structures to analyze sequencing reads from thousands of human genomes.
Genome research - 1 Feb 2017
Dolle Dirk D, Liu Zhicheng, Cotten Matthew, Simpson Jared T, Iqbal Zamin, Durbin Richard, McCarthy Shane A, Keane Thomas M
Abstract excerpt
We are rapidly approaching the point where we have sequenced millions of human genomes. There is a pressing need for new data structures to store raw sequencing data and efficient algorithms for population scale analysis. Current reference-based data formats do not fully exploit the redundancy in population sequencing nor take advantage of shared genetic variation. In recent years, the Burrows-Wheeler transform...
Read the complete abstract on PubMedTopics
Share this publication in a Topic to start or enrich a Post.
