@thsottiaux Looking forward to it! It was frustrating to send one prompt, not even finish the task, and already hit the 5-hour limit on Pro. I get why letting agents finish after 0% was removed due to abuse, but couldn’t that extra usage simply count against the remaining weekly limit?
@thsottiaux Happy to hear that. It was frustrating to send one prompt, not even finish the task, and already hit the 5-hour limit on Pro. I get why letting agents finish after 0% was removed due to abuse, but couldn’t that extra usage simply count against the remaining weekly limit?
For anyone doing cryo-EM, I built a small Mol* extension for creating focused masks directly in the browser. Started as something for my own workflow, but I figured others might find it useful too.
Would love any feedback if you give it a try.
https://t.co/iuHiQ8Mb2F
🚀 We’ve launched a new CryoCloud website!
CryoCloud has grown significantly over the past few years – more users, more deployments, and continuous technical innovation. It was time for a website that reflects that evolution.
🌐 Take a look: https://t.co/QxrmSG6RhA
#cryoEM
Big news from Boltz today: we’re launching Boltz Lab, a new platform with new small-molecule + protein design agents, announcing Boltz PBC and a $28M seed round, and sharing a multi-year partnership with Pfizer. More below! 🚀
Excited to release BoltzGen which brings SOTA folding performance to binder design! The best part of this project has been collaborating with many leading biologists who tested BoltzGen at an unprecedented scale, showing success on many novel targets and pushing its limits! 🧵..
De-novo design of a random protein walker
🚀 New preprint from David Baker!🚀
1. A team of researchers has achieved a significant milestone in protein engineering by designing a random protein walker that diffuses along a designed protein track. This is a major step forward in the field of molecular machines, as it demonstrates the potential for creating dynamic protein systems.
2. The study presents the design and characterization of a protein walker that can move along a protein track without the need for fuel. The walker consists of homo-oligomers with reversible binding feet, allowing it to diffuse along the track. Cryo-EM experiments confirmed the structure of both the track and the walkers.
3. The researchers designed micro-meter long fibres to serve as tracks and developed a method to rigidly decorate them with arbitrary proteins. This method, called Fibre H-fuse, enables the attachment of various proteins to the fibres, providing a versatile platform for future applications.
4. The study tested multiple heterodimer interfaces for reversibility and designed six walkers with different numbers of feet. Single molecule tracking experiments revealed that walkers with more feet diffuse faster along the track, highlighting the influence of multivalency on movement.
5. The system represents a tunable starting point for future powered protein molecular machines. The walkers' diffusion rates were influenced by the number of feet and the binding affinity of the feet to the track. This provides a powerful approach to fine-tune the system's behavior under various conditions.
6. The researchers also explored the effects of environmental conditions on the walkers' movement. Higher temperatures, lower viscosity, and lower salt concentrations resulted in faster diffusion rates. This suggests that the system can be further optimized for specific applications.
7. The study's findings open up new possibilities for the development of programmable and robust de novo designed protein nanorobots. These could have significant implications for fields such as precision medicine and material science.
📜Paper: https://t.co/EJMbEpHoPH
#ProteinEngineering #MolecularMachines #DeNovoDesign #Biophysics #SyntheticBiology
De novo Design of All-atom Biomolecular Interactions with RFdiffusion3
🚀 New preprint from David Baker!🚀
1. A groundbreaking study introduces RFdiffusion3 (RFD3), a novel diffusion model that generates protein structures in the context of ligands, nucleic acids, and other non-protein atoms. This model represents a significant leap forward in the field of protein design, offering a more detailed and accurate approach to creating proteins with specific interactions.
2. RFD3 stands out by modeling all polymer atoms explicitly, allowing for more precise conditioning on complex sets of atom-level constraints. This is crucial for enzyme design and other applications where specific interactions are required. The model achieves improved performance compared to previous methods, with a notable reduction in computational cost.
3. The study demonstrates the broad applicability of RFD3 through in silico benchmarks and experimental validation. The model successfully designs DNA-binding proteins and cysteine hydrolases, showcasing its potential for creating functional proteins. The ability to generate protein structures guided by complex atom-level constraints expands the range of achievable protein functions.
4. RFD3’s architecture includes a transformer-based U-Net with sparse attention blocks, enabling effective coupling between atom-level and residue-level features. This design allows for joint sampling of backbone and side-chain coordinates, resulting in more diverse and accurate protein structures. The model also incorporates classifier-free guidance to improve adherence to conditioning information.
5. The training of RFD3 involves a hierarchical procedure using the Protein Data Bank (PDB) and high-quality AlphaFold2 structures. This approach prevents overfitting and ensures the model can handle a wide range of design tasks. The model’s performance is evaluated across various challenges, including protein-protein binding, protein-DNA binding, small molecule binding, and enzyme design.
6. Experimentally, RFD3 designs are validated through the creation of DNA-binding proteins with measurable binding affinity and cysteine hydrolases with high catalytic efficiency. These results highlight the model’s potential for generating functional proteins in vitro, demonstrating the practical impact of its improved atomic conditioning capabilities.
📜Paper: https://t.co/iIOeIMjvYP
#ProteinDesign #DeepLearning #BiomolecularInteractions #RFdiffusion3 #ComputationalBiology
BoltzDesign1: Inverting All-Atom Structure Prediction Model for Generalized Biomolecular Binder Design
1. BoltzDesign1 introduces a new framework that inverts the Boltz-1 all-atom structure prediction model, enabling direct protein binder design across diverse molecular targets—without model fine-tuning or diffusion backpropagation.
2. By optimizing the distogram (probability distribution of atomic distances) instead of sampled structures, BoltzDesign1 significantly reduces computational cost while maintaining design fidelity and enhancing structure diversity.
3. Unlike prior models like RfDiffusionAA, BoltzDesign1 leverages flexible ligand modeling during optimization, generating binders that adapt to ligand conformation at each iteration—a major advantage for unknown or dynamic targets.
4. The method employs only the Pairformer and Confidence modules from Boltz-1, skipping the memory-intensive diffusion module, yet still achieves high in silico success rates for small molecule binders such as IAI, FAD, SAM, and OQO.
5. BoltzDesign1 surpasses RfDiffusionAA in design diversity, AlphaFold3 success metrics (pLDDT > 0.7, iPAE < 10), and structural consistency across models—while also generating binders with higher sequence recovery at key interfaces.
6. A modular four-stage optimization scheme enables exploration of continuous sequence space before converging on confident, one-hot encoded sequences—further enhanced with LigandMPNN for interface-aware refinement.
7. The framework generalizes beyond small molecules: BoltzDesign1 designs binders for metal ions (iron, zinc), B-DNA, and post-translationally modified proteins (e.g., PCNA-Y211, Smad2, CD45 glycosylation), all validated with AlphaFold3 and AllMetal3D.
8. Designed binders exhibit expected biochemical properties—e.g., known coordination residues for metal binding and electrostatic complementarity for nucleic acid interactions—demonstrating biological plausibility and precision.
9. BoltzDesign1 achieves highest design success without recycling steps, suggesting that simpler optimization pipelines can prevent overfitting or adversarial confidence inflation during structure prediction.
10. Distogram-based loss correlates well with confidence metrics like pLDDT and inter-pAE, validating its use as a structural proxy for both intra- and inter-molecular contacts in protein-ligand design.
11. The framework offers a generalizable, architecture-efficient approach for generating functional binders in protein engineering, drug discovery, and diagnostics—by making full use of pretrained structure models without retraining.
12. Future directions include integrating nucleic acid MSAs, using templates for constrained design, and expanding support for flexible RNA/DNA targets and multivalent or multi-modified biomolecular systems.
💻Code: https://t.co/h1GU6mURmu
📜Paper: https://t.co/97uLU2I4ER
#ProteinDesign #BinderDesign #AlphaFold3 #StructurePrediction #MolecularInteractions #ComputationalBiology #DeepLearning #SyntheticBiology #Bioinformatics #ICLR2025
I’m excited to share our significantly-updated preprint on de novo antibody design, where we now demonstrate the structurally accurate design of scFvs (in addition to VHHs) with RFdiffusion! https://t.co/WdYKu0s1Uc
I'm usually not too emotional about paper acceptances, but this one justifies it. 🥹 It doesn't feel real, but PepPrCLIP is now published at @ScienceAdvances! Let me tell you its story. 📕
📜: https://t.co/x0nlBSJ8cJ
💻: https://t.co/uQftSZGdgb
PepPrCLIP (🌶️📎) began as Cut&CLIP, before I even came to @DukeU in 2022, where my first @Harvard undergrad mentee @kalyanmpalepu (now at @DEShawResearch ) would cut ✂️peptides from interacting partner proteins, and throw them thru a trained peptide-protein CLIP 📎 model to predict which ones were specific to the target (cool use of DALL E 2's architecture, right?). We had good initial experimental results, but they weren't robust enough for publication. 🤔
While the team tried to get our CLIP model better for experimental testing, my amazing other @Harvard undergrad @garykbrixi came up with a really unique way of cutting peptides from the ESM-2-predicted binding sites of partners (our SaLT&PepPr model🧂). We changed the focus to SnP (get it, SnP ✂️?), and spent over a year through numerous good/bad review processes, and FINALLY got it out last October in @CommsBio: https://t.co/DavIss1CXP 🥳
@bhat_suhaas, a @Harvard undergrad, later a Rhodes Scholar (so proud of you!! 🥹🫵), and @kalyanmpalepu's roommate, who stuck with me as I transitioned to @DukeU, kept working on it with Kalyan, and not only trained a beautiful CLIP model on peptides and proteins with ESM-2 embeddings (goodbye ESM-1b and MSA transformer!), but also devised an incredibly clever way of generating new peptides via Gaussian perturbation of peptide embeddings for screening with CLIP, creating PepPrCLIP! 🌶️📎 The experimental validations were beautiful, and it became our first de novo peptide model in the lab! 👩🔬
We went into submission at another top journal, and worked hard to iron out IP issues, but at the end of the day, it was just too difficult to convince structure-focused reviewers how this was better than RFDiffusion 😑, even though we showed strong data showing PepPrCLIP-generated peptides allowed us to extend to conformationally disordered targets that RFDiffusion/structure-based methods could not access (we even did better on structured targets than RFDiffusion in vitro!). Frustrating. 😒
Still, we had incredible data! As you can see in our manuscript, PepPrCLIP-generated peptides can serve as inhibitory peptides to enzymes (i.e. biotin ligases, like UltraID) as well as degraders to transcription factors (β-catenin) and EVEN heavily-disordered fusion oncoproteins (SS18-SSX1)! 🧫 We showed extensive binding, inhibition, and degradation data in endogenous cellular settings and have now extended PepPrCLIP to more undruggable targets! And after over 2 years since our first submission, @ScienceAdvances, who published my first first-author paper back in graduate school, decided to accept it! 🌟 Full circle. ⭕️ Incredible. 😭
To the most important part: I am forever, forever grateful to @bhat_suhaas and @kalyanmpalepu's incredible dedication to this work, even after I moved to @DukeU and even after they graduated from @Harvard. 🫂 I love you guys. 🫶 And I couldn't be more thankful to our incredible collaborators @SoderlingLab @DeLisaGroup and @anideshpandelab (as well as the amazing experimentalists in my lab) who believed in our algorithm and did the painstaking experiments to prove it out. 🙏 Just an INCREDIBLE team-wide effort to get this over the finish line. I am just so lucky and so grateful -- I don't deserve such an amazing collaborators. 🥲
Finally, with the generous support of my company @UbiquiTxINC, we have made the code FREELY AVAILABLE to academics after signing a non-commercial license! I can guarantee you that PepPrCLIP peptides will work for your target! 😉
Feel free to read our paper and provide your thoughts! 💡We just got multiple other big paper acceptances and will be sharing those shortly as they come online! 🪇
We're thrilled to present ESM3 in @ScienceMagazine. ESM3 is a generative language model that reasons over the three fundamental properties of proteins: sequence, structure, and function. Today we're making ESM3 available free to researchers worldwide via the public beta of an API for biological intelligence.
Trained with over a trillion teraflops of compute, this is the first time a model of this scale has been trained for biology, pushing the frontier of AI for biological discovery and engineering.
ESM3 learns to represent the immense complexity of protein biology, learning from billions of natural proteins. From this training it developed the capability to design proteins, responding to complex prompts combining atomic level details and high level instructions to generate new proteins.
ESM3 can explore protein space far beyond natural evolution. We prompted ESM3 to generate a fluorescent protein at a far distance from any known fluorescent proteins, searching an unknown region of protein space, to discover a new fluorescent protein.
We estimate this is equivalent to simulating five hundred million years of evolution.
Excited to share our latest pre-print! 🌟 Collaboration between @UWproteindesign & @NIH tackles the global threat of #influenza & the need for #H5N1 countermeasures. #CryoEM reveals how conserved motifs & water molecules help achieve broad protection.
https://t.co/uRsPcnkisD
Super excited to preprint our work on developing a Biomolecular Emulator (BioEmu): Scalable emulation of protein equilibrium ensembles with generative deep learning from @MSFTResearch AI for Science.
#ML#AI#NeuralNetworks#Biology#AI4Science
https://t.co/yzOy6tAoPv
We are all born with a genetic lottery. Millions of T cell receptors are what we have with a hope to defend all cancer and viruses. What if that's not enough? Hope our work can give an interesting answer to you. https://t.co/o7HhX553RY