Back to search

Article

SMILES Pair Encoding: A Data-Driven Substructure Tokenization Algorithm for Deep Learning

2020-05-21

Abstract excerpt

SMILES-based deep learning models are slowly emerging as an important research topic in cheminformatics. In this study, we introduce SMILES Pair Encoding (SPE), a data-driven tokenization algorithm. SPE first learns a vocabulary of high frequency SMILES substrings from a large chemical dataset (e.g., ChEMBL) and then tokenizes SMILES based on the learned vocabulary for deep learning models. As a result, SPE a...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
e67bdac6-eb09-5ff7-a49a-e85d06684039
DOI
10.26434/chemrxiv.12339368.v1
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
SMILES Pair Encoding: A Data-Driven Substructure Tokenization Algorithm for Deep LearningDOI 10.26434/chemrxiv.12339368.v1
Select a neighboring publication to make it the new centre.