One prediction of this claim, for example, might be that we would expect that when a "rigorous" model makes a formal error in its proof, that error propagates, while a "post-rigorous" model would quickly realize that something is intuitively wrong, and correct it.
Per this other Terry Tao blog post, for example: https://t.co/XEZooRdUyK
As models get better and better at mathematical proofs, I've been thinking about this classic blog from Terence Tao, on the stages of mathematical maturity, from pre-rigorous (intuitive, unable to use formalism) to rigorous (able to use formalism to manipulate mathematical objects), to post-rigorous (intuitive understanding of mathematical objects, formalism is effectively subconscious).
My impression is that models today, although seemingly excellent at proving mathematical theorems, are still in the rigorous stage, and are doing more formalism manipulation than intuitive mathematics. And I wonder whether RL-ing on lean proofs can lead to models that think "intuitively" about math. Does this match the understanding of folks in the field? And if so, is there a way to precisely define what "mathematical intuition" would look like in a model?
Blog: https://t.co/4W8t9JMx4w
We used AI to predict the failure of a Phase 3 trial before the results were announced. Today, we're publishing 10 more predictions for the future.
Thread 🧵
This was the culmination of my undergraduate research. It has been such a gift to have been mentored by @pranamanam -- working with him is among the best decisions I have ever made.
I'm usually not too emotional about paper acceptances, but this one justifies it. 🥹 It doesn't feel real, but PepPrCLIP is now published at @ScienceAdvances! Let me tell you its story. 📕
📜: https://t.co/x0nlBSJ8cJ
💻: https://t.co/uQftSZGdgb
PepPrCLIP (🌶️📎) began as Cut&CLIP, before I even came to @DukeU in 2022, where my first @Harvard undergrad mentee @kalyanmpalepu (now at @DEShawResearch ) would cut ✂️peptides from interacting partner proteins, and throw them thru a trained peptide-protein CLIP 📎 model to predict which ones were specific to the target (cool use of DALL E 2's architecture, right?). We had good initial experimental results, but they weren't robust enough for publication. 🤔
While the team tried to get our CLIP model better for experimental testing, my amazing other @Harvard undergrad @garykbrixi came up with a really unique way of cutting peptides from the ESM-2-predicted binding sites of partners (our SaLT&PepPr model🧂). We changed the focus to SnP (get it, SnP ✂️?), and spent over a year through numerous good/bad review processes, and FINALLY got it out last October in @CommsBio: https://t.co/DavIss1CXP 🥳
@bhat_suhaas, a @Harvard undergrad, later a Rhodes Scholar (so proud of you!! 🥹🫵), and @kalyanmpalepu's roommate, who stuck with me as I transitioned to @DukeU, kept working on it with Kalyan, and not only trained a beautiful CLIP model on peptides and proteins with ESM-2 embeddings (goodbye ESM-1b and MSA transformer!), but also devised an incredibly clever way of generating new peptides via Gaussian perturbation of peptide embeddings for screening with CLIP, creating PepPrCLIP! 🌶️📎 The experimental validations were beautiful, and it became our first de novo peptide model in the lab! 👩🔬
We went into submission at another top journal, and worked hard to iron out IP issues, but at the end of the day, it was just too difficult to convince structure-focused reviewers how this was better than RFDiffusion 😑, even though we showed strong data showing PepPrCLIP-generated peptides allowed us to extend to conformationally disordered targets that RFDiffusion/structure-based methods could not access (we even did better on structured targets than RFDiffusion in vitro!). Frustrating. 😒
Still, we had incredible data! As you can see in our manuscript, PepPrCLIP-generated peptides can serve as inhibitory peptides to enzymes (i.e. biotin ligases, like UltraID) as well as degraders to transcription factors (β-catenin) and EVEN heavily-disordered fusion oncoproteins (SS18-SSX1)! 🧫 We showed extensive binding, inhibition, and degradation data in endogenous cellular settings and have now extended PepPrCLIP to more undruggable targets! And after over 2 years since our first submission, @ScienceAdvances, who published my first first-author paper back in graduate school, decided to accept it! 🌟 Full circle. ⭕️ Incredible. 😭
To the most important part: I am forever, forever grateful to @bhat_suhaas and @kalyanmpalepu's incredible dedication to this work, even after I moved to @DukeU and even after they graduated from @Harvard. 🫂 I love you guys. 🫶 And I couldn't be more thankful to our incredible collaborators @SoderlingLab @DeLisaGroup and @anideshpandelab (as well as the amazing experimentalists in my lab) who believed in our algorithm and did the painstaking experiments to prove it out. 🙏 Just an INCREDIBLE team-wide effort to get this over the finish line. I am just so lucky and so grateful -- I don't deserve such an amazing collaborators. 🥲
Finally, with the generous support of my company @UbiquiTxINC, we have made the code FREELY AVAILABLE to academics after signing a non-commercial license! I can guarantee you that PepPrCLIP peptides will work for your target! 😉
Feel free to read our paper and provide your thoughts! 💡We just got multiple other big paper acceptances and will be sharing those shortly as they come online! 🪇
An exciting @biorxivpreprint update on our recent PepPrCLIP model!🌶️📎 Hopefully, you all remember PepPrCLIP, where we apply Gaussian perturbations to the peptidic latent space of ESM-2 to de novo generate naturalistic peptides, and then input these new peptide sequences into a CLIP model we trained on peptide-protein interactions, to identify peptides that bind to the target protein. All you need to give it is a protein sequence, and you get prioritized binder sequences. 🌟
In our preprint update, we optimize PepPrCLIP and show that model-generated peptides can target conformationally diverse proteins, and can thus serve as inhibitory peptides to enzymes (i.e. biotin ligases, like UltraID) as well as degraders to transcription factors (β-catenin) and heavily-disordered fusion oncoproteins (SS18-SSX1)! 🧫 We show extensive binding, inhibition, and degradation data in endogenous cellular settings and are very excited to extend our models to more undruggable targets. 🦾
This is the result of brilliant algorithmic development by my students, @bhat_suhaas and @kalyanmpalepu, and wonderful experimental collaborations from @anideshpandelab, @SoderlingLab, and @DeLisaGroup! Truly a team effort and shows how widely-useful PepPrCLIP is. 🙌
We're still trudging through the submission/revision process, but please read our updated manuscript, and we will soon make the code available open-source to academics with a free non-commercial license! Feel free to DM me if you're interested in trying some of our peptides.🤗
Paper: https://t.co/CCoOyZWaGq
Code will be available here: https://t.co/uQftSZFFqD
SaLT&PepPr is published in @CommsBio! 🥳 Here, we fine-tune the ESM-2 pLM to identify peptidic binding sites on target-interacting partner sequences. We fuse these "guide" peptides to E3 ubiquitin ligases to degrade disease-causing proteins! 💻➡️🧫 (1/n) https://t.co/Dqp0Q5MtZL
Our new preprint is up -- and it's completely de novo! 🤩Our Peptide Prioritization via CLIP (PepPrCLIP) model is the first de novo binder design algorithm only needing the target amino acid sequence (no 3D structure needed). 🌶️📎Take a read here! https://t.co/qpmuoTZM6w
Our new preprint is up -- and it's completely de novo! 🤩Our Peptide Prioritization via CLIP (PepPrCLIP) model is the first de novo binder design algorithm only needing the target amino acid sequence (no 3D structure needed). 🌶️📎Take a read here! https://t.co/qpmuoTZM6w
BREAKING: After a decade of constant pressure by students, faculty, and alums, @HARVARD IS FINALLY DIVESTING FROM FOSSIL FUELS.
It’s a massive victory for our community, the climate movement, and the world — and a strike against the power of the fossil fuel industry. (THREAD)