This one is really special to me! As someone who probably spends more time on @uniprot than on Google, I’m incredibly proud that SeqHub search is now integrated with UniProt, bringing together UniProt’s rich protein annotations with SeqHub’s view of native genomic context across diverse microbial lineages.
This gives researchers a more complete picture of proteins of interest, connecting what we know about their function with the genomic neighborhoods and evolutionary contexts in which they occur.
🧬 Today, SeqHub is live!
Researchers have sequenced terabases of DNA across the globe, but a vast majority of these sequences remain poorly understood.
SeqHub is the place to understand your proteins and genomes of interest, with context-aware sequence search, genomic language model-driven functional annotations and SeqHub Agent to help uncover new patterns.
#SeqHub #AIforScience #AIinBiology #OpenScience #Launch
🚀 AlphaFold3's ipSAE_min is the new default for de novo protein binder triage.
Across 3,766 designs/15 targets, ipSAE_min had +1.4× AP vs ipAE (add ΔG/ΔSASA or shape complementarity to boost precision). 🔥
Interface‑focused confidence + physics is overtaking “one‑score‑fits‑all” heuristics.
Quantifying Protein-Protein Interaction with a Spatial Attention Kinetic Graph Neural Network
1.This paper introduces SAKE-PP, a physics-inspired spatial-attention equivariant graph neural network designed to directly predict interface RMSD (iRMSD) of protein-protein interactions without requiring native reference structures.
2.SAKE-PP combines a force-field-like spatial attention mechanism with Laplacian eigenvector orientation to integrate local inter-residue forces and global protein topology, enabling more accurate characterization of protein interfaces.
3.Trained on a curated dataset of over 15,000 docking conformations sampled from 736 protein complexes using a novel hierarchical iRMSD-guided sampling strategy, SAKE-PP effectively learns to prioritize near-native protein-protein conformations.
4.On a benchmark of 176 heterodimers with AlphaFold3 (AF3)-generated decoys, SAKE-PP outperforms the native AF3 ranking score by 13.75% in iRMSD accuracy and 12.5% in DockQ score, demonstrating better hit rates, overlap, and correlation metrics.
5.Importantly, SAKE-PP shows strong zero-shot generalization on 139 antibody-antigen complexes, improving the score-iRMSD correlation by 0.4 over AF3 ranking scores despite never having been trained on antibody complexes.
6.SAKE-PP’s architecture uniquely updates both node features and 3D coordinates in an equivariant manner, simulating inter-residue spatial kinetics, which supports physically realistic modeling of protein interfaces.
7.The hierarchical sampling approach balances the training dataset across diverse iRMSD ranges, substantially improving training stability and predictive accuracy compared to uniform or top-score sampling.
8.Molecular dynamics simulations confirm that conformations selected by SAKE-PP have superior binding energetics and dynamic stability compared to AF3-selected decoys, highlighting its potential in structure-guided drug design.
9.SAKE-PP is an independent scoring framework that complements existing structure predictors like AlphaFold, overcoming biases in their internal confidence scores that often favor structurally deviated decoys.
10.Future extensions of SAKE-PP could include applications to more complex biomolecular assemblies and integration with free energy and adaptive sampling methods to enhance protein complex prediction workflows.
💻Code: https://t.co/c12sMM9Fix
📜Paper: https://t.co/d0zWYuWeWN
#ProteinProteinInteraction #GraphNeuralNetwork #DeepLearning #StructuralBiology #AlphaFold #DrugDiscovery #ComputationalBiology
🚨Data leakage is real🚨 in protein LLMs.
Pretrained pLLMs often boost PPI prediction scores—not because they generalize, but because they memorize.
SqueezeProt proves it: strip test-set proteins from pretraining, and scores drop sharply.
Benchmarking AlphaFold3-like Methods for Protein-Peptide Complex Prediction
1. The paper systematically evaluates the performance of AlphaFold3 and its replication models (Protenix, Chai-1, and Boltz-1) in predicting protein-peptide complex structures, highlighting their improvements over AlphaFold2-multimer (AF2m).
2. Next-generation models achieve significantly higher prediction success rates under stringent criteria (Fnat ≥ 0.8), increasing from 53 percent for AF2m to 70-80 percent, with Protenix performing best at 80.8 percent accuracy.
3. A multi-method combination strategy, particularly AF3 and Protenix, raises the high-quality prediction success rate to 89 percent and covers 97 percent of cases under moderate criteria (Fnat ≥ 0.5), demonstrating synergy among models.
4. Unlike traditional docking methods, AlphaFold3 introduces a diffusion-based optimization process, replacing iterative refinements with a noise-adding and denoising mechanism for enhanced structure prediction.
5. Chai-1, a replication of AF3, performs comparably to AF3 and benefits from multiple sequence alignments (MSAs), achieving a success rate of 78.8 percent with MSAs and 70.7 percent without them.
6. Protenix outperforms other models in predicting accurate binding pockets, achieving the highest proportion of high-precision pocket predictions (DCC < 1 Å), correlating with its superior success rate in complex structure modeling.
7. Evaluations using metrics such as pLDDT, ipTM, RMSD, and the number of interactions between peptides and protein receptors (#PPI) indicate that inter-model consistency is a strong predictor of successful predictions.
8. Despite their advancements, these models still face challenges in predicting highly flexible peptide interactions and cyclic peptides, where binding site misidentification and incorrect conformations remain common failure modes.
9. Combining structural consistency metrics (DockQ) with multiple modeling methods improves prediction reliability, providing an effective strategy for selecting the best structural model among multiple outputs.
10. Future work will focus on further refining interfacial modeling, incorporating molecular dynamics simulations, and improving binding mode predictions, especially for cyclic and disordered peptides.
📜Paper: https://t.co/LRnUZLJ3j2
#ProteinPeptideInteractions #AlphaFold3 #MolecularModeling #AIforScience #ComputationalBiology
🚀 Thrilled to launch our PhD-level agent: AI-Researcher! 🎓 Exploring Innovation Capabilities of LLM Agents 🔍
The complete source code 💻 and papers 📝 generated by AI-Researcher are now available at: https://t.co/O3arx6M0bs
"From Concept to Publication" 🚀
This fully automated system eliminates the need for manual intervention throughout the entire research lifecycle, enabling seamless scientific discovery through every critical phase:
📚 Literature Review & Idea Generation
🧪 Algorithm Design & Implementation
💻 Algorithm Validation & Refinement
📊 Result Analysis
✍️ Manuscript Creation
Our AI-Researcher empowers scientists with:
🎯 Full Autonomy: Complete end-to-end research automation
🔄 Seamless Orchestration: Integrated workflow across all research phases
🧠 Advanced AI Integration: Powered by cutting-edge LLM agents
🚀 Research Acceleration: Streamlined scientific innovation
Mapping Targetable Sites on the Human Surfaceome for the Design of Novel Binders
1. A groundbreaking study maps the human surfaceome, identifying 4,500 targetable sites across 2,886 cell-surface proteins. This resource unlocks new therapeutic opportunities for precision medicine.
2. The study leverages MaSIF, a geometric deep-learning framework, to predict protein-protein interaction sites. Nearly 3 billion docking runs were performed to generate high-quality binder "seeds" for targeted design.
3. A novel web platform, SURFACE-Bind, is introduced. It provides open access to predicted binding sites, corresponding binder seeds, and data visualization tools. This resource aids drug discovery and protein design.
4. Experimental validation highlights three critical targets—FGFR2, IFNAR2, and HER3. De novo-designed binders showed high success rates, targeting key interfaces with nanomolar to micromolar affinities.
5. The team optimized protein design pipelines by integrating ProteinMPNN and AlphaFold2, improving biophysical properties like stability and binding affinity, achieving 11-fold higher success rates in subsequent rounds.
6. A peptide design pipeline was also developed. Interface motifs from mini-protein binders were stabilized as cyclized peptides, yielding 5 target-specific peptides with demonstrated binding activity.
7. This work underscores the power of integrating deep learning and physics-based methods to advance de novo protein and peptide design for therapeutic applications, targeting underexplored surfaceome regions.
8. By combining computational innovation with experimental validation, this research sets a new benchmark for precision protein engineering and opens the door to the next generation of biologics.
@befcorreia@hamed_khakzad@yangche7@SiFulle@J_Damjanovic_
💻Code: https://t.co/S8qD7CnRcg
📜Paper: https://t.co/FLmbHhA5qg
#ProteinDesign #DeepLearning #Surfaceome #ComputationalBiology #DrugDiscovery #MachineLearning
We updated our EvoBind paper with the cyclic binder results: it turns out we can design cyclic peptide binders in a single shot with sub nM affinity only from sequence information (75% success rate). Protein binder design is getting very good. Stay tuned. https://t.co/8TmFCRF8Tf
Nobel Prize is NOT about h-index or citations.
It is about the emergence of big new fields.
So many posts discuss Nobel awardees.
And so many misunderstand the Nobel Prize.
📍 A bit of clarification from my side:
1⃣ Nobel Prize is NOT about how useful your work is.
It’s about how useful it WILL BE.
Science is not about real-world impact.
It is about new knowledge, new understanding.
It’s about nucleating new ways of thinking.
Applications can come decades after the discovery.
2⃣ Nobel Prize is NOT about just doing risky research.
Many of us take on risky projects. But most stay as niche studies that could have been done by others.
It’s about doing what others are AFRAID to do.
It's about looking like a reckless scientist.
It’s about succeeding where others have failed (despite numerous attempts).
3⃣ Nobel Prize is NOT about a small study.
It’s about nucleating a BIG research direction.
It’s about "OMG, I didn't know it's even possible!"
Yes, sometimes it takes decades to recognize a scientist. But in many cases, the prize was given to those who published the "nucleating studies" and pushed hard to grow the new field.
4⃣ Nobel Prize NOT about a lot of citations.
Metrics doesn't matter.
Forget this "Stanford top-2% ranking".
Your peers' opinion is what really matters.
Are you recognized by your peers SO MUCH that they want to see you as a Nobel Laureate?
Do they see you as someone who created their field and made their research possible?
Do they see you as their thought leader?
▫️
My HUGE congratulations to all Laureates.
AI has changed science a lot.
Those who made it possible deserve this recognition.
#science #AcademicChatter
The Protein Design Competition results are in!
🧬 200 designs tested in our lab
🌍 90 protein designers from around the world
💎 5 novel binders found
🎯 2.5% hit rate (vs 0.01% on previous EGFR work)
1st place: @MartinPacesa & Lennart Nickel
2nd place: @khRRustamov
3rd place: Adrian Tripp (@tugraz)
A beautiful thread from @rohitsingh8080!! 🧵 It really captures my feelings/frustration about current "trends" in protein design. As Rohit mentioned, the aftermath of sequencing the human genome has been a huge huge disappointment. Genomics has really not lived up to the hype, and has underperformed in terms of effect on therapeutic development. Rather, it's been more holistic "go around" methods like GWAS that has allowed us to find mutations that, when fixed with CRISPR, can help treat diseases. Still, I will argue that that's more on the inherent programmability of CRISPR than the utility of GWAS.
Rohit's point on the analogous potential for protein design to not live up to the hype resonates with me so much. Making protein binders is super important -- literally, that's what my lab currently does, and we bind really tough, disordered, disease-driving targets (not obvious structured targets that AlphaFold methods get to, like EGFR). But you know what we also do? We show that our binders work in RELEVANT experimental situations, not just solving a useless crystal structure to demonstrate interaction, but actually to get rid of these targets in relevant cellular (and soon, animal) models 🧫🐁. I mean, look at the targets we've degraded in our recent preprints with some (relatively simple) pLM tricks:
-SS18-SSX1, Beta-catenin with contrastive learning: https://t.co/AACMO2DpFK
-MSH3, mutant Huntington, pandemic viral proteins with span masked language modeling:
https://t.co/zcgSbeH4PP
Why am I writing this? Because I care about protein design being useful beyond "tricks" and being useful to people who should use them. Sometimes I feel the random complexity of structure-based design, with all the weird SE(3) equivariance, harmonic self-conditioned flow matching, etc. -- it's like "omg, look how smart we are to get such meaningful structures." But has AlphaFold created a real therapeutic? No. And based on the @adaptyvbio results, not really anytime soon. For us, a simple span masking on pLM embeddings got us degraders of Nipah viral phosphoproteins with a 6/20 hit rate. 5/153 structure-based designs kind of bound highly-structured EGFR? And what's that going to do? We can easily get scFvs and nanobodies raised or phage displayed to those targets. It's like, what, are we doing? Just being performative at this point? 😒
This is also where I'm also baffled at how much capital has been raised by general protein design "startups". Most of them don't even know what functional thing they're going to do. It's just like, let's play around with new AI until some pharma gives us a (commercally viable) indication and maybe, at that point, we can go after it (as long as there is a structure of it, of course). 🙄 Until then, remake HER2 antibodies forever or unconditionally design de novo CRISPR proteins or run hackathons and have some fun. Like really? Like, legit, what are you doing? And these companies are raising capital from VCs like there's no tomorrow. 🤦♂️ And here I am, hoping for a $100k grant to computationally/experimentally/translationally/clinically go after a rare pediatric disease. Just how the world works, I guess. 🤷♂️
I get it: it's research. But we're not studying biology. AlphaFold3 does NOT model BIOLOGY -- it's more like a semi-okay, slightly irrelevant snapshot of it. We must develop models that will actually provide useful things for society. Otherwise, why? So we can look at pictures of predicted proteins in physiologically irrelevant conditions? As a protein design field, we need to stop doing things just to do it.
Okay, that's enough of a rant. I didn't choose to use pLMs because I think they are cool. I chose for my lab to use and develop them because, as Rohit said, I believe they capture more relevant information and are the "go around" that allows to solve important bioengineering problems. Undruggable protein binding, heavy metal sequestration, environmental bioremediation -- these are complex problems that we CAN solve if we don't hyperfocus on only local information. And for what it's worth, I'm grateful we've received financial support from foundations, non-profits, and the NIH to solve them, and my lab will continue to work hard to do so. 🙏
Congrats to @HubrichFlorian & @sanathrajkk0 for this exciting story finally being published. Thanks for the fruitful collaboration with @chekanlab. If you are interested in the chemical space of #RiPP#biosynthetic#enyzmes and #isoprenoids, click here:
https://t.co/LGXC9ier8Z
LevSeq: Rapid Generation of Sequence-Function Data for Directed Evolution and Machine Learning
🚀Groundbreaking work from the lab of Nobel laureate Frances H. Arnold, a pioneer in the field of directed evolution @francesarnold 🚀
1. The article introduces LevSeq, a new method combining nanopore sequencing with a dual barcoding strategy, enabling rapid and comprehensive sequence-function data generation, ideal for protein engineering workflows.
2. LevSeq’s integration into directed evolution workflows ensures full-length gene sequencing with minimal cost, optimizing both time and resources while enhancing accuracy in variant detection.
3. The system provides real-time quality control before the screening phase, offering significant reductions in library screening burden by filtering out low-quality sequences early.
4. One of the key innovations is its application in generating machine learning-compatible data, streamlining protein engineering by coupling sequence information with functional data to guide further optimization.
5. LevSeq has been demonstrated on two protein engineering projects, where it accurately detected key mutations and epistatic interactions in enzyme engineering for new-to-nature chemistries.
6. By utilizing long-read nanopore sequencing, LevSeq enhances the ability to explore the protein fitness landscape, providing vital insights that were previously inaccessible with traditional methods.
7. LevSeq is designed with user-friendly, open-source software, enabling researchers to easily analyze mutagenesis libraries and incorporate machine learning for improved directed evolution outcomes.
8. It enables the collection of essential sequence-function data, which serves as a foundation for advanced machine learning models to predict high-performance protein variants, a major leap in data-driven protein design.
9. LevSeq offers a robust tool for protein engineers to explore and optimize vast sequence spaces, with potential applications in industrial, environmental, and pharmaceutical biocatalysis.
💻Code: https://t.co/MQk5BvXBI7
📜Paper: https://t.co/du91WnBYrb