New grammar-based molecular representation captures ring structures and improves generation and prediction tasks
Chemistry models learn from how molecules are written down. Standard ways of writing molecules, such as simple sequences or graphs, can miss important higher-order structure like ring systems and repeating motifs. This paper introduces Higher-order Grammar Representation (HGR), a new way to describe molecules that directly encodes those higher-order features and makes them usable by common machine learning models.
HGR works by lifting a molecule into a combinatorial complex — a mathematical object that can record how atoms and bonds form rings and larger connected pieces — and then parsing that object with a context-free higher-order grammar. In plain terms, the molecule is broken down into a compact sequence of production rules. Those rules are easy to feed into standard sequence models, so the representation keeps the rich topological information but avoids the heavy computation needed by some previous higher-order encodings.
To test the idea, the authors built a large, ring-focused benchmark called RingDiv with 1.18 million molecules and a curated subset called RingDiv300k. They also introduced the ring diversity index (RDI) to measure how well a dataset covers different ring systems. These resources were designed to reduce bias in benchmarks toward simple ring structures and to stress-test models on more realistic ring chemistry.
The results reported are promising. For molecular generation, HGR-based models produce molecules that are valid by construction — the grammar only generates chemically valid structures — and they achieved top performance on distributional alignment across five generation benchmarks, ranking first on a commonly used metric (FCD) that measures how well generated molecules match target distributions. For representation learning, a model called HGR-FM achieved the highest mean area under the curve (AUC) across seven MoleculeNet benchmarks under two transfer protocols, improving over the strongest baseline by 8.3 AUC points when using probing and by 3.3 AUC points with full fine-tuning.