Given a reaction, which enzyme catalyses it? Given an enzyme, what can it perform? A geometric foundation model that answers both
Of the ~250 million protein sequences in UniProt, fewer than 0.3% have been manually curated for function. Meanwhile, 40–50% of known enzymatic reactions lack any associated enzyme sequence—orphan reactions. Traditional approaches rely on EC number classification, which groups distinct reactions under the same code, or sequence homology tools like BLASTp, which fail when similarity is low. Neither directly models whether a specific enzyme structure can catalyse a specific reaction.
Yong Liu and coauthors introduce EnzymeCAGE, a geometric foundation model trained on ~1.5 million structure-informed enzyme–reaction pairs across 3,273 species. The key architectural choice is to focus on the catalytic pocket rather than the full protein. A GNN encodes pocket geometry—backbone coordinates, dihedral angles, side-chain torsions—extracted via AlphaFill from AlphaFold structures, while ESM Cambrian embeddings capture global evolutionary context. On the reaction side, SchNet encodes 3D substrate and product conformations, with a reacting-area weight matrix that upweights atoms at the reaction centre. Geometry-enhanced cross-attention then models pocket–reaction interactions to output a catalytic compatibility score.
On unseen enzymes, EnzymeCAGE achieves 58% top-10 success rate—a 45% improvement over baselines including CLIPZyme, ESP, and MMseqs2. For orphan reactions, enzyme retrieval improves by 41%. It works even when test enzymes share less than 30% sequence identity with training data, where homology methods break down. An emergent capability is catalytic site identification: attention weights consistently highlight experimentally validated active-site residues, despite this never being a training objective.
In two case studies—withanolide biosynthesis and glutarate pathway reconstruction—EnzymeCAGE correctly retrieves catalytic enzymes where all baselines fail, ranking positive P450s within the top 6–13 among 107 candidates at only ~40% sequence similarity to training proteins.
The design principle: by decomposing catalysis into pocket geometry, reaction centre chemistry, and their 3D interaction—rather than relying on sequence similarity or coarse EC labels—the model learns transferable representations of catalytic compatibility that generalize across enzyme families.
Paper: https://t.co/GWME8Dw9rY
I’m excited to share the first preprint out of the Altemose Lab! This stems from a heroic effort by Dr. Matt Franklin @matt_franklin_ (who's on the job market!), who made surprising discoveries regarding some of the most mysterious regions of the genome. https://t.co/rH3plvRIpA
(1/16) Check out our @SimonMJGaudin paper https://t.co/ZcCmKI4CFh from the Canzio lab out in @ScienceMagazine today on how tuning local cohesin trajectories enables differential readout of the clustered Protocadherin (Pcdh) locus across different types of neurons.
A green alga cell can swim 100 microns per second.
Scientists put little baskets on a wheel. As algae became trapped, they pushed and spun the wheel around.
How to make sense of tons of transcriptomic data in cell biology?
@Laurent_Guyon managed to correlate A CELLULAR METRIC (the centrosome position) with changes in the AMOUNT OF RNA in multiple cells lines.
@murugan_chicago@RRavasio @kabir8husain @SzostakLab Love the waddington analogy. Would it be accurate to say that nonequilibrium systems use erode a flat energetic landscape that constrains trajectories? Like how the Colorado river created the Grand Canyon?
Lab-kept bumble bees roll small wooden balls around for no apparent purpose other than fun, revealed a 2022 study.
Learn more on #WorldBeeDay: https://t.co/72GH2Ptkru @NewsfromScience
A "metal umlaut" is a phonetically irrelevant diacritic used gratuitously or decoratively over letters in the names of some heavy metal bands
A journal's impact factor could skyrocket with similar creative marketing
@nature@TheLancet@CellCellPress@nejm@ScienceMagazine
8,803 cardiovascular disease sufferers were randomized to semaglutide (Ozempic) and another 8,801 were randomized to a placebo.
Over the next four years, the semaglutide group had 20% fewer cardiovascular deaths, nonfatal infarctions, and nonfatal strokes.
If you are not satified with ChimeraX built-in CLI, try this!
I made an enhanced, intelligent CLI widget "CliX", which supports command suggestion, completion, syntax highlighting and multi-line execution!
https://t.co/y2M79eEwqp
A pet peeve of mine: I dislike the use of the word "incremental" to describe research in negative terms.
All research is incremental. Even Ramanujan studied Carr's "A Synopsis of Elementary Results in Pure and Applied Mathematics".
Ingenious strategy to repurpose a bacterial E1-like enzyme to emulate ubiquitin-like cascades for ATP-driven protein ligation. Its like a battery pack for powering bioconjugation reactions 🔋!
Ingenious strategy to repurpose a bacterial E1-like enzyme to emulate ubiquitin-like cascades for ATP-driven protein ligation. Its like a battery pack for powering bioconjugation reactions 🔋!
Our paper on swimming active protein droplets is finally out! 🎊🎉
https://t.co/KAraORBem2
The main result is in the title: "Phase-separated droplets swim to their dissolution".
Explanations below 🧵👇
@SoftLiv_Cornell@ETH_Materials@Cornell@CNRSingenierie@NatureComms