Article
An efficient algorithm for identifying matches with errors in multiple long molecular sequences.
Journal of molecular biology - 20 Oct 1991
Leung M Y, Blaisdell B E, Burge C, Karlin S
Abstract excerpt
An efficient algorithm is described for finding matches, repeats and other word relations, allowing for errors, in large data sets of long molecular sequences. The algorithm entails hashing on fixed-size words in conjunction with the use of a linked list connecting all occurrences of the same word. The average memory and run time requirement both increase almost linearly with the total sequence length. Some...
Topics
- Algorithms
- Base Sequence
- Consensus Sequence
- Databases, Factual
- Escherichia coli
- Molecular Sequence Data
- Mutation
- Repetitive Sequences, Nucleic Acid
- Sequence Alignment
