It's been a year since we first started writing this primer, and so much has changed already. Always amazed by how fast the field grows🌱
Also check out the rest of this @NatureBiotech issue focused on protein engineering. Many more primers, reviews, and news & views 👇
Our primer "Generative models for protein structures and sequences" is now live https://t.co/dF377XhcZZ @chloehsu0@seafann (free version: https://t.co/KWSJFGUYlh)
It's been a year since we first started writing this primer, and so much has changed already. Always amazed by how fast the field grows🌱
Also check out the rest of this @NatureBiotech issue focused on protein engineering. Many more primers, reviews, and news & views 👇
Our primer "Generative models for protein structures and sequences" is now live https://t.co/dF377XhcZZ @chloehsu0@seafann (free version: https://t.co/KWSJFGUYlh)
Had fun putting together this review! Lots of good surveys of the literature lately so mostly we used it as a chance to look carefully at the key impacts on de novo design, and sneak a few hot/cold(?) takes. Also we figured out a way to reference elephants 🐘 in the text 😂
Pre-print: https://t.co/Jl5vNT1z0v
Code: https://t.co/pfFMdckDMG
Twitter threads from me and @alexrives: https://t.co/qaQqht3Cvr
https://t.co/SRUBAtO29M
Excited to share our new ESM-IF1 inverse protein folding model.
The result of scaling inverse folding with millions of predicted structures.
Paper: https://t.co/2EioLHKfCo
Model: https://t.co/aUWt2xpgIB
Our new inverse folding model (protein structure -> sequence) trained w/ 12 million predicted structures is now available on Colab.
Special thanks to @adamlerer and @TomSercu for open sourcing efforts. Joint work w/ @alexrives and @MetaAI Protein Team.
https://t.co/mIVceY26Bk
This large scale inverse folding project wouldn't have been possible without @adamlerer. Very grateful for the mentorship during the internship and the opportunity to work together. Also would love to see this gets used for designing new proteins.
It was a pleasure to work with @chloehsu0 during her internship last year! Our preprint is out describing our large scale inverse protein folding model, i.e. predicting sequence from 3D structure, trained on millions of sequences using AlphaFold2 predicted structures.
Excited to share our new ESM-IF1 inverse protein folding model.
The result of scaling inverse folding with millions of predicted structures.
Paper: https://t.co/2EioLHKfCo
Model: https://t.co/aUWt2xpgIB
The evolutionary velocity paper ended on a cliffhanger: protein language models could predict evolution retrospectively, but could they also run evolution forward to prospectively design new proteins? So, I retrained as a protein biochemist to find out...
https://t.co/v6S7UqPAdG
Here’s what we learned from inverse folding on millions of #AlphaFold structures. Exciting time to bring a 800x new scale to #proteindesign. ESM-IF1 more accurately designs sequences to fold into desired structure, also unlocking new design capabilities.
https://t.co/WT0ARI2c7C
The ESM-IF1 model uses GVP-GNN encoder layers to extract geometric features, followed by a generic autoregressive encoder-decoder Transformer. We found that this simple architecture is sufficient to learn inverse folding at scale.
Model weights & code: https://t.co/pfFMdckDMG
Existing inverse folding models are limited by the relatively small number of experimentally determined structures. Larger models especially benefit from these 12M new predicted structures. Grateful that such a scale is possible at all today, and curious to see what comes next.
We next show that ESM-IF1 is an effective zero-shot predictor of mutational effects. Examples: mutational effects on the binding affinity of SARS-CoV-2 RBD to human ACE2, AAV packaging (gene delivery), stability of de novo mini proteins, and more.
Beyond existing benchmarks, we also make the sequence design task more challenging along three dimensions: (1) introducing masking on coordinates; (2) generalization to protein complexes; and (3) conditioning on multiple conformations. Our new training data help with all three!
The best model trained with predicted structures improves native sequence recovery by 9.4 percentage points (51.6% vs 42.2%) over the previous state-of-the-art model. Sequence recovery (accuracy) measures how often sampled sequences match the native sequence at each position.