Today in @NatureBiotech we report the use of phage-assisted evolution (PACE) to reprogram botulinum neurotoxin (BoNT) proteases to trigger cancer cell death. We evolved proteases to cleave procaspase-1 or gasdermin D, inducing cell death and slowing tumor growth in a highly drug-resistant mouse model.
https://t.co/OTUVxl1FpG
1/13
Only about a quarter of human diseases have an approved therapy. By some estimates it's a few percent. Most of those treatments slow a disease rather than stop it.
Closing that gap is what the AI-cures-everything story promises. Build a system smart enough and the cures hidden in what we already know will fall out. I have worked at the intersection of machine learning and biology for three decades, and I believe #AI will eventually transform human health. The capabilities arriving now are extraordinary. But the promise rests on an assumption that is simply false: that we already understand human biology well enough for a clever enough reasoner to find the answers in it. We don't.
More than 90% of drugs entering clinical trials fail, a number that has barely moved in decades. In the large majority of those failures the molecule was engineered just fine. The mechanism it targeted was wrong. We are doing a pretty good job at manufacturing keys, but they are generally for the wrong locks. And because nobody wants to fail in the clinic, the industry has retreated to the locks it already trusts: 38 targets now have more than 50 programs against each of them, while the number of novel targets advanced per year fell from roughly 100 in 2015 to about 30 in 2024.
AI will not reason its way past this. Biology wasn't engineered. It is the product of billions of years of messy, stochastic evolution, and the variation that produced is too vast and too idiosyncratic to work out in the abstract. You have to measure it. Aimed at a biology this thinly sampled, AI will mostly help us generate failures faster.
I founded @insitro because getting to the right locks requires a different kind of system. We generate multimodal human and cellular data at scale, use machine learning to find causal drivers of disease, and test those hypotheses experimentally. Virtual Human™ is built for causal discovery; TherML™ turns what it finds into the right therapeutic intervention. It is working: first-in-class programs internally and with partners, three #ALS targets that Virtual Human™ identified and @bmsnews nominated, and additional collaborations with @EliLillyandCo and @GileadSciences .
Today we are launching Deep Phenotype: Scaled Biology, Deep Causality. Issue one is "Drug Discovery Has No Magic Wands," the first half of a two-part essay on the magical thinking currently running through our field and what I think it will actually take. After that you will hear from insitro's own scientists and engineers, people who work across computation and experiment because the problem requires both.
Getting this right is hard, and we do not have all of it worked out. I hope you will follow along and think it through with us.
https://t.co/MFTNRNfbDl
Our paper in @Nature today 🥳 We tracked 6,438 mice from puberty to death and mapped the genetics of *when* you die, not just whether a gene associates with lifespan.
https://t.co/EoeexqJoHk
59 loci. Two decades of data. Thread 👇
#Longevity#Aging#Genetics#Healthspan
Evo 2 is out in Nature today, showing that genome language models can predict and design across the full complexity of life, from phages to eukaryotes.
A few surprises from the project, including how ignoring trillions of nucleotides was key to getting a good model. 🧵
🚨 Today in @Nature, we report GEMINI—a genetically encoded intracellular memory device that writes cellular dynamics into tree-ring-like fluorescent patterns within cytoplasmic protein assemblies.[1/n]
https://t.co/eVchPCiK6f
Delighted to share new @arcinstitute work from our group on AI-accelerated lab-in-the-loop, in @ScienceMagazine today
One of the most remarkable things about biology is that it's digital. DNA, RNA, proteins: these are all sequences, and their function is directly encoded in their sequence of letters. But a protein of length N has 20^N possible variants and the vast majority are non-functional. Evolution spent billions of years finding the functional needles in this haystack through random exploration and natural selection. For modern biomedicine, we need to solve this in days to weeks.
Excited to share VIPerturb-seq!
New tech from my lab which aims to improve the cost, data quality, and efficiency of single-cell CRISPR screens so that they are accessible to any lab - even at genome-wide scale
Preprint and 🧵 (1/): https://t.co/m8nleniSUD
We're running a @CompoundVC Research Day on Rapid in vivo Iteration in SF in late Feb.
In vivo evidence is already arguably the core bottleneck to drug development. This will intensify in a world where we have orders of magnitude more drugs to screen. From safety to efficacy to solubility to distribution, we need to assess many drug parameters as rapidly as possible. We will discuss the many emerging strategies to scale these parameters with in vivo evidence (including animal models), human clinical trial speed ups, and increased data modality collection and predictive validity.
Please come to this if you're:
- Frustrated with the speed of drug development
- Seeking different strategies for knowing if your drug works
- Interested in contributing to the future of a more expeditious drug discovery system
Thanks to @mackenziejem for co-organizing!
Chemistry writes the story of life.
Every cell, every second, it tells it again - powered by molecular design.
At the core of this process lies the citric acid cycle, the central hub of metabolism.
Here, acetyl-CoA derived from carbohydrates, fats, or proteins enters a precise series of reactions that convert fuel into energy.
Each step transfers electrons, drives ATP synthesis, and sustains the continuous renewal of life at the cellular level.
What may seem invisible is, in fact, the most constant motion in existence; the quiet rhythm of biochemistry that powers everything we do.
🧬 🍀 Gene-edited four-leaf clovers???
There is only one four-leaf clover for every 5,076 three-leaf clovers.
So far, we only know that the gene responsible for this rare trait is recessive for the quadruploid plants, and that luck favors warm conditions, approximately two-fold when it comes to growing four leaves.
The precise genotype of the iconic four-leaf clover, however, remains a mystery 🧵
*Note also no *APCs* at hybrid journals so those journals will be an option if publishing a paywalled article along with associated fees, which is a bummer
Have $10M and want to cure some diseases?
It costs ~$1B to invent a new medicine…but $10M buys you Project Encore, to see if ANY existing drug might be repurposed for ~100 diseases that have no treatments.
Contact me - be a hero for desperate patients!
https://t.co/9mKWkkRWYy
The NIH has decided that scientists can only submit 6 grants a year. OK. But you need a score that’s < 10% at least to get funded. Meaning, at best, 1 out of 10 grants you submit will have a chance (YMMV of course). 🫠
Can we use RNA to insert kilobases of DNA into the human genome, opening up a new way to engineer cells or treat disease?
Excited to share our latest work turning a retrotransposon (a “jumping gene” with an RNA intermediate) from songbirds into a genome editing tool. 🧵 (1/7)
🚫 No dyes. No bleaching.
🔬 Just AI + label-free microscopy = vivid virtually stained images
New in @NatMachIntell: A deep learning model that enables robust virtual staining across microscopes, cell types & conditions. #CZBiohubSF@mattersOfLight explains:
We made a huge poster that illustrates all of the major genome editing tools in one place.
You can download a copy for free from the @AsimovPress website.
https://t.co/fI2mWvUA7f
I had a similar experience with this N of 1
Why current deep learning sequence models of gene expression struggle to predict counterfactual effects of variants on expression.
https://t.co/INbFHZ6fDc
Also asked for suggestions for improvements
https://t.co/qL6rM5Oj9d