A new benchmark and model aim to teach AI the difference between a reaction and a reaction that only happens because an enzyme made it happen.
Researchers introduced VenusRX-Bench, a benchmark covering forward reaction prediction, single-step retrosynthesis, and Enzyme Commission (EC) number prediction, built from multiple biochemical databases with standardized curation and leakage-controlled splits. Testing existing chemical and enzymatic models on it exposed a clear gap: models trained on ordinary organic chemistry struggle once an enzyme's catalytic role enters the picture. The team then built VenusRX, a T5-style sequence-to-sequence model trained in two stages - first on millions of template-expanded reactions, then on real biochemical ones - with optional EC conditioning and a decoding method constrained to a molecule library to keep generated biomolecules chemically sane. Across the benchmark's hardest generalization splits, VenusRX matched or beat existing chemical and enzymatic baselines on most tasks.
This matters because most reaction-prediction AI treats chemistry as pure structure, ignoring that enzymes are often the entire reason a reaction happens in a living system. Get the catalytic context wrong and predictions for drug metabolism or synthetic-biology pathways fall apart. The finding that EC classification and reaction prediction improve each other suggests these were two sides of one problem that prior work had split apart for no good reason.
A unified, leakage-controlled benchmark is the less glamorous but more durable contribution here - plenty of reaction-prediction models get announced, few survive contact with a standardized test built to catch them cheating.