Article
Tokenizing single-cell transcriptomes as a native language for large language models
2025-10-24
Abstract excerpt
Large language models (LLMs) can process diverse forms of information once they are represented as tokens in a shared sequence space. However, single-cell transcriptomes remain a foreign modality to LLMs because they are continuous, high-dimensional molecular profiles rather than discrete linguistic units. Here, we propose CellTok, a tokenized single-cell language modeling approach that converts transcriptomic pro...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- 9aaa2535-7d79-54f9-b17a-5a441e2f2e60
- DOI
- 10.1101/2025.10.22.684047
