Article
Distinguishing word identity and sequence context in DNA language models
2023-07-12
Abstract excerpt
Transformer-based large language models (LLMs) are very suited for biological sequence data, because the structure of protein and nucleic acid sequences show many analogies to natural language. Complex relationships in biological sequence can be learned, although there may not be a clear concept of words, because they can be generated through tokenization. Training is subsequently performed for masked token predic...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- 7e89e1c3-b503-5829-9d96-8517ea98f923
- DOI
- 10.1101/2023.07.11.548593
