Article
NucEL: Single-Nucleotide ELECTRA-Style Genomic Pre-training for Efficient and Interpretable Representations
2025-08-17
Abstract excerpt
Pre-training large language models on genomic sequences has become a powerful approach for learning biologically meaningful representations. While masked language modeling (MLM)-based approaches, such as DNABERT and Nucleotide Transformer (NT), achieve strong performance, they are hindered by inefficiencies due to partial token supervision, pre-training/fine-tuning mismatches, and high computational costs. We intr...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- fae2acdf-3279-561c-84ce-e72b47b517b5
- DOI
- 10.1101/2025.08.17.670700
