Article
SMILES Pair Encoding: A Data-Driven Substructure Tokenization Algorithm for Deep Learning
2020-05-21
Abstract excerpt
SMILES-based deep learning models are slowly emerging as an important research topic in cheminformatics. In this study, we introduce SMILES Pair Encoding (SPE), a data-driven tokenization algorithm. SPE first learns a vocabulary of high frequency SMILES substrings from a large chemical dataset (e.g., ChEMBL) and then tokenizes SMILES based on the learned vocabulary for deep learning models. As a result, SPE a...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- e67bdac6-eb09-5ff7-a49a-e85d06684039
- DOI
- 10.26434/chemrxiv.12339368.v1
