Back to search

Article

Distinguishing word identity and sequence context in DNA language models

2023-07-12

Abstract excerpt

Transformer-based large language models (LLMs) are very suited for biological sequence data, because the structure of protein and nucleic acid sequences show many analogies to natural language. Complex relationships in biological sequence can be learned, although there may not be a clear concept of words, because they can be generated through tokenization. Training is subsequently performed for masked token predic...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
7e89e1c3-b503-5829-9d96-8517ea98f923
DOI
10.1101/2023.07.11.548593
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
Distinguishing word identity and sequence context in DNA language modelsDOI 10.1101/2023.07.11.548593
Select a neighboring publication to make it the new centre.