Article
BEETL-fastq: a searchable compressed archive for DNA reads.
Bioinformatics (Oxford, England) - 1 Oct 2014
Janin Lilian, Schulz-Trieglaff Ole, Cox Anthony J
Abstract excerpt
MOTIVATION: FASTQ is a standard file format for DNA sequencing data, which stores both nucleotides and quality scores. A typical sequencing study can easily generate hundreds of gigabytes of FASTQ files, while public archives such as ENA and NCBI and large international collaborations such as the Cancer Genome Atlas can accumulate many terabytes of data in this format. Compression tools such as gzip are often...
Read the complete abstract on PubMedTopics
Share this publication in a Topic to start or enrich a Post.
