SMILES gave chemistry a language. But it may not give Transformers the right vocabulary.

TL;DR


Summary:
- The article explores the limitations of the Simplified Molecular Input Line Entry System (SMILES) when used as a tokenization method for training Large Language Models (LLMs) and Transformers in the field of computational chemistry.
- It highlights how the inherent structural ambiguity and non-unique representation of molecules in SMILES strings can hinder the ability of AI models to accurately predict chemical properties and synthesize new compounds.
- The author advocates for the development of more robust, graph-based, or canonical representations that better align with the underlying topological nature of molecular structures, rather than relying on linear string-based notations.

Like summarized versions? Support us on Patreon!