Hi there! I'm a graduate researcher at @TAMU's Biochemistry and Biophysics department, also a #graphicdesigner#3Dartist. Follow me for updates on my research
A Unified 3D Generative Model for Synthesizable Structure-Based Drug Design
1. The paper introduces LDDM (Large Drug Discovery Model), a unified 3D generative framework that jointly samples atom types, covalent bonds, and 3D coordinates conditioned on a protein pocket—so the same model can do de novo design, fragment growing/linking, (non)covalent docking, and even non-canonical peptide side-chain docking.
2. A key methodological idea is fragment-based masked modeling (inspired by masked language modeling): molecules are BRICS-fragmented, and during training different fragments are masked in different “roles” (design: graph+coords missing; docking: coords missing but graph known; context: fully known). This trains one model that naturally supports many structure-based tasks by changing what is masked/conditioned.
3. The generative core combines equivariant flow matching for continuous 3D coordinates with Markov bridge models for discrete atom/bond types, plus per-atom uncertainty estimates. That uncertainty becomes a practical scoring signal: it correlates with pose RMSD (Spearman ρ = 0.59) and can highlight which substructures are unreliable.
4. Training data is scaled via a curated synthetic complex dataset (836,259 protein–ligand complexes; 454,662 unique molecules) assembled from CrossDocked, BindingNet v2, and BigBind (BigBind ligands docked with Gnina), then heavily filtered with PoseBusters + GenBench3D and outlier filters. The model’s samples cover the training chemical space while staying novel: ~90% of generated molecules have <0.5 max Tanimoto similarity to the training set.
5. Benchmarking shows “unified but competitive” performance: on PoseBusters local docking, LDDM reaches 79.5% top-1 success at 2 Å RMSD (valid cases), and on a covalent docking benchmark it achieves 80.1% top-1 success at 2 Å—outperforming specialized covalent docking baselines in the cited comparison.
6. Beyond one-shot generation, the paper proposes programmable generation: an iterative fragment-tree search that alternates global filtering (e.g., QED thresholds) with local fragment-level filtering (pose plausibility, hydrogen-bond satisfaction, target interactions), allocating sampling budget with a UCB-style rule (Monte Carlo tree search flavor). This improves sampling efficiency for local, structure-dependent objectives versus SMILES RL (REINVENT), especially for hydrogen-bonding criteria.
7. A second practical innovation is synthesizable generation: programmable generation constrained by building blocks + reaction templates (demonstrated with Enamine REAL). Instead of generating “anything” then hoping it’s makeable, LDDM generates full molecules but only accepts expansions that match precomputed reaction products; partial matches use an MCS-anchored hybrid docking step. In comparisons with equal sampling budgets, this approach yields far more candidates that are both filter-passing and present in REAL space.
8. Experimental validation spans five targets and multiple modalities, emphasizing “few syntheses, high hit rates” rather than massive screening. In a VHL–CDO1 molecular glue campaign, LDDM redesigned two CDO1-contacting flanks while keeping the VHL anchor fixed; across both flank campaigns, 67% of synthesized LDDM modifications retained nanomolar ternary-complex affinity by SPR.
9. LDDM also optimized a covalent non-natural peptide inhibitor of cathepsin S (CTSS): docking non-canonical side chains (54 building blocks; 270 single mutants; 200 conformers each) led to one improved design (J2) with 3.5× better potency, improving IC50 from 70.4 nM to 19.9 nM in FRET assays—despite LDDM not being trained on peptides.
10. De novo hit-finding case studies show micromolar binders with orthogonal validation and pose accuracy: for PGK1, 4/6 purchasable one-shot designs bound (best KD 14.7 µM) with NMR CSPs supporting binding at the targeted nucleotide pocket; for BRD4, synthesizable designs produced binders (e.g., Z967 ~40–60 µM) with NMR CSPs consistent with the acetyl-lysine pocket; for SARS-CoV-2 Nsp3 Mac1, 6/88 synthesized candidates showed competition in HTRF (best IC50 46.2 µM), and X-ray structures for two designs showed ligand RMSDs of 2.35 Å and 1.43 Å vs designed poses.
💻Code: https://t.co/OYEHi2NWgi.
📜Paper: https://t.co/0S2pgret2F
#ComputationalDrugDiscovery #GenerativeAI #StructureBasedDrugDesign #DiffusionModels #FlowMatching #MolecularDocking #FragmentBasedDesign #SyntheticAccessibility #DrugDesign #ChemicalBiology
“Codex, please design some small molecules to fight malaria.”
In this tool calling demo, Codex with GPT-6 Astra used LDDM to generate candidate molecules which might bind Plasmodium falciparum's dihydroorotate dehydrogenase, an enzyme the malaria parasite needs to make pyrimidines. Disrupting this enzyme can stop the parasite from replicating.
Codex used LDDM in two ways:
1) de novo molecule generation
2) extending a fixed fragment of the existing dehydrogenase inhibitor, DSM265
DSM265 was reported by Coteron et al in 2011. This demo uses its experimentally determined binding pocket in PDB 4RX0. The attached vid shows the LDDM generation steps.
The LDDM generated candidates had no exact matches in Astra's PubChem and ChEMBL searches, and the de novo molecule had no >85% similar hits.
Codex used retrosynthesis tools (AiZynthFinder, Syntheseus, and SynPlanner) to search for routes to make these, but no complete routes were found during the demo.
Codex used RDKit and PoseBusters to check geometry and protein/cofactor clashes (no severe clashes). Codex also used AutoDock Vina to score fit in the pocket and these computational candidates scored similarly to DSM265.
If you work in Life Sciences and haven't downloaded our biological viewer plugins yet, you're missing out!
Go to the plugin marketplace in ChatGPT desktop and look for the "Scientific Research" section. You'll see them there!
Some cool examples attached below
Science is entering a new era - one where AI agents can do scientific work.
🧬 Today NVIDIA is launching the BioNeMo Agent Toolkit - an open, agent-ready toolkit that gives any AI agent callable tools for protein structure prediction, molecular docking, generative chemistry, genomic analysis, and more.
(1/2)
1/ In my final project as part of the @CorryLab, we show the bacterial TAM complex facilitates the spontaneous flow of phospholipids from the IM to the OM in MD simulations..
https://t.co/74LBAaoP7I
🚀 The largest-ever open‑source protein‑complex treasure trove - 1.7 million of AI‑predicted complexes now live in the AlphaFold Database
In collaboration with @emblebi, @GoogleDeepMind, and @SeoulNatlUni, we have added millions of predicted complexes to the AlphaFold Database to accelerate global health research.
🧵👇
ProteinTTT is now easy to run on Hugging Face Spaces and Google Colab. We’ll also be presenting the paper at ICLR 2026 🇧🇷
🤗 Hugging Face Space: https://t.co/tq4lWuTqVJ
⚙️ Google Colab: https://t.co/nUSN6dEd5o
🧵👇
From code → complex structures 🧬
Meet OpenFold3 NIM: A production-ready model for predicting proteins, nucleic acids & ligands at lightning speed.
Dive in 👉 https://t.co/CzZYOPaTDZ
BoltzGen is now live on @proteinbase! We validated the new all-atom protein design model extensively in the lab and now you can access all the experimental data!
The full dataset with over 400 lab-validated proteins is open-source, same as the model itself
Introducing LiteFold Dynamo.
For the first time, you can now run 100ns+ molecular dynamics simulations at LiteFold. Molecular dynamics is one of the last steps before researchers validate in wet labs.
At LiteFold, run multiple simulations, manage them, get insights, and view trajectories in real time within the platform.
Check it out, it's free. Links in the comments.
🚀Introducing fully automated data processing for repeat-target #cryoEM.
Using new tools in #CryoSPARC, it is now possible to obtain resolutions & map quality equal to or better than manual processing, with zero user intervention.
Preprint: https://t.co/pt9aRc8p5L
✨Run the world’s fastest protein structure prediction with #NVIDIARTXPRO 6000 Blackwell Server Edition to fold proteins up to 4.8x faster than L40S.
96 GB GDDR7 keeps MSAs + OpenFold2 on-GPU, cutting cloud spend.
#AI4Science#DrugDiscovery
Boltz v2.2.1 out. A few improvements including support for .pdb templates, better treatment of stereochemistry in guidance potentials, and improved documentation. As always a great thank you to all those in the community who contributed via PR, raising issues or directly reporting issues & improvements to us.