In new work, we lay out a vision for a high-level programming language for generative biology, called Proto.
Proto composes generative and predictive models spanning DNA, RNA, proteins, ligands, and their interactions, which we use to design complex biological functions. 1/n
A techbro told me that biology is easy because DNA is just code, right?
I told him that DNA is 4 billion year old, completely undocumented, vibe-coded spaghetti, built by a blind evolutionary algorithm, which codes for its own compiler and runtime environment.
His confidence popped like a balloon.
Anthropic is going big on biology. In April they paid $400M for a biotech team. Now they hired the Nobel prize winner who built AlphaFold.
In 2024, @DarioAmodei wrote that AI could compress 100 years of biological progress into 10. He definitely meant it.
I really think biology is the most important use case in AI right now. I don't know much about it but I think it's pretty incredible what humans have managed to do with drug discovery.
Now imagine humans + AI at that same problem. Things just accelerate! Helping people live longer and better.
1. Je n'ai jamais dit que les LLMs étaient inutiles ou n'avaient pas de valeur commerciale.
2. Mais beaucoup de technologies sont utiles et à haute valeur sans être particulièrement "intelligentes"
3. Les LLMs sont essentiellement inutiles et sans grande valeur pour ce qui est du pilotage de processus industriels ou de compréhension de données continues, de haute dimension, et/ou bruitées.
4. Il y a énormément d'applications de haute valeur dans le pilotage de processus industriels ou de compréhension de données continues, de haute dimension, et/ou bruitées.
5. A long terme, l'intelligence, la compréhension du monde, et la capacité à planifier est ce qui compte vraiment.
Together with my co-founders Michael @MichaelPoli6, Stefano @Massastrello and Armin @athmsx, I am excited to announce @RadicalNumerics is emerging from stealth with a $50M seed round to build general biological intelligence.
We’re also sharing an early preview of our new model Omnii, the most powerful genome language model to date.
Omnii preview link:
https://t.co/ouikMtRVwf
At Radical Numerics, our mission is to master the code of life, and to drive the frontier of biological AI for both design and defense.
This is our dual mandate, which comes from something our own team helped make possible.
Our founding team trained Evo and Evo 2, the largest biological AI models (40B params) trained on DNA sequences. Trillions of tokens across all of life, from microbes to mammals. It’s fully open source, and created the field now known as generative genomics.
Last year, scientists used Evo to generate the world’s first complete genome from scratch using AI. Turns out it was a bacteriophage—a type of virus. It functioned in the real world, and in this case it was harmless. But for us, it was a clear turning point.
It showed that AI is no longer just analyzing biology. It is on the cusp of generating functional lifeforms. Eventually, AI will have the power to design and control life itself.
That should make all of us incredibly excited, and incredibly uneasy. (Anyone can design DNA with a new function, and have it synthesized and delivered, like something from Amazon Prime).
The same technology that will help us cure cancer is the very technology that might create the next global pandemic, or worse, allow the creation of bioweapons that can wipe out populations.
We believe these forces are inseparable. If you work on the frontier of biology, you have to build technology to safeguard it from its misuse. Existing biosecurity tools are sorely losing the arms race, relying on outdated “have I seen this exact thing before?” style algorithms.
We founded Radical Numerics to turn the tide.
And we can’t do that by training on textbooks and natural language. We must understand the language of biology from the raw physical data itself, to reason across every molecule and modality, from DNA to proteins.
The next frontier for AI goes far beyond chatbots or video generators to models that can understand and engineer life.
Today, we’re previewing Omnii, which is already far surpassing Evo 2, and will continue improving as we scale and add new modalities (training now).
1. For human health, Omnii can read and write whole genomes (more on writing later). It’s state of the art (SOTA) on detecting causal variants for disease, and can rank Alzheimer's mutations zero-shot. We’re partnering with a diagnostics company to use Omnii for early cancer detection (pancreatic and multi-cancer).
2. For defense, Omnii is SOTA at detecting AI-generated pathogens. We benchmarked existing detection tools, and they simply can’t detect the AI-generated ones (“deepfake viruses”). We’re partnering with a US national lab to pilot Omnii for detecting the next pandemic, both natural and AI-generated.
We have a data center full of Blackwells in construction now to build the most powerful biological AI models ever. This mission takes a new kind of AI lab that can actually scale on physical, biological data: new alignment research (mid/post training), scaling long context, building out mech interp teams to dissect what these models learn, new architectures and systems designs, all from the ground up.
Our team is made up of AI researchers and scientists from top labs and institutions (e.g. Stanford, MIT, Google DeepMind), but more importantly, we all share the belief that this is the most important challenge of our lifetime. If you feel similarly, we are hiring. We aim to bring the brightest minds in AI and science together to save lives.
Thanks to our partners on this journey, led by Emergence Capital @emergencecap, with Obvious Ventures @obviousvc, Triatomic @TriatomicCap
, and Patrick Collison @patrickc. Our advisors include Eric Horvitz @erichorvitz, CSO of Microsoft, Chris Re @HazyResearch of Stanford, George Church @geochurch of Harvard, and Andrew Weber @AndyWeberNCB, former Assistant Secretary of Defense for Nuclear, Chemical and Biological Defense Programs.
Fortune article: https://t.co/L3f3f1329T
Jobs: https://t.co/EzsHSMcGJ1
This sort of mindset is probably why xAI failed to catch up to other frontier labs.
If he wants to make SpaceXAI into a frontier lab, hope he changes his mindset.
Though being a cloud provider is probably something they can easily excel in anyway lol (Colossus is impressive!)
We are happy to announce Precigenetics India, a subsidiary of Precigenetics, is being established this month. It will be opening up our first data foundry in New Delhi.
North India sees some of the highest rates of breast cancer, lung cancer and a variety of cancers of the GI tract. We will be setting up digital pathology using our microscopy expertise to accelerate surgeries for hospitals in India by hours.
LUMA, our data factory for LIVING cells, will also be set up in year two in India.
With the help of our partners at leading hospitals in New Delhi, we will be collecting pathology data and tissue to enrich our models, and entering partnerships with frontier labs to do some data sharing.
Much of the data used to train foundation models comes from a small set of the global population.
This is not just a disservice to all other patients—this is the reason why our models are so weak across the world today at actually deciphering cancer.
Further announcements and news announcements are forthcoming.
With this, we are doubling down on our promise to create datasets that change the field of medicine.
Step two will be to create life saving cures with the intelligence we are actively making.
the drugs being designed right now will fail for the wrong reason.
the models are arriving. the data infrastructure that makes them useful is not built. most biological data was never generated to be trained on. clinical records were designed for billing. assay results were designed for one research question.
the data AI actually needs- systematic, multi-modal, longitudinally tracked, designed from the beginning to close the feedback loop between model and experiment- largely does not exist. @insitro built their entire architecture around generating it deliberately. most companies are training on whatever already exists and inheriting every limitation of that data.
secondly, the data is not usable even when it exists. batch effects- systematic differences introduced by when and where an experiment was run- mean a model trained on data from one lab fails on data from another.
thirdly, the data covers the wrong people. South Asians are 25% of the world's population and 0.8% of genomic study participants. the drugs being designed right now are optimised against the wrong population genetics for two billion people. better models trained on the same data will not fix this.
to cap it all, AI makes the incentive problem worse by raising the economic value of the data, which increases the incentive to hoard it.
the most durable position in AI biology is owning the data generation infrastructure that produces data nobody else can replicate. geographically specific. biologically specific.
Can we program cells like computers — using RNA?
Two years ago, our group trained the first language model to decode the regulatory grammar of 5′ UTRs in mRNA, published in Nature Machine Intelligence.
Today, we’re excited to share the next step, also in Nature Machine Intelligence:
“Programmable RNA translation through deep learning-driven IRES discovery and de novo generation.”
We built an AI engine to discover, predict, optimize, and generate IRES elements — RNA control modules that regulate translation initiation.
This brings us closer to programmable RNA systems that control when, where, and how strongly proteins are produced inside cells.
AI is no longer just helping us read biology.
It is beginning to help us write it and harness it.
The future of computing may not only run on silicon — it may also run inside living cells.
#AIForBiology #LLM #AI4S #AI #RNA #MachineLearning #Bioengineering
This isn't in the trial phase.
The entire China International Consumer Products Expo in Hainan, recently, used only these materials for signage, food containers, and more.
This is getting scaled for mass use.
Love seeing Silico (@GoodfireAI ) used to probe our EchoJEPA's representations! this is exactly the kind of interpretability work that's been missing for JEPA-style models.
One thing that makes EchoJEPA particularly interesting to interpret: unlike MAE-based approaches, it never reconstructs pixels. The model learns entirely in latent space through masked prediction, so you can't just look at decoder outputs to understand what it captured. Attribution onto a temporally aligned 3D mesh is a much more honest probe of what the representations actually encode.
What we found in building EchoJEPA: training on 18M echo videos across 300K patients, the model learns to disentangle cardiac anatomy from ultrasound noise (speckle, reverberation artifacts) almost entirely through self-supervision. With 1% labeled data it already outperforms supervised baselines trained on 100%. The latent space is doing real anatomical work, but until you can visualize it like this, "real anatomical work" is mostly a claim.
Paper + code: https://t.co/BFDoHrsZ7n | https://t.co/0DFWi3lzVA
Prometheus deserves way more credit as a top-tier SciFi movie.
Genuinely beautiful and breathtaking shots. An interesting story. Great cast. It’s got it all.
Every time you breathe, saliva droplets are released into the air. These droplets contain DNA, which can be captured from the air and sequenced for the next ~24 hours. So-called "AirDNA" is a relatively new way to do environmental monitoring; you can figure out who entered a room, for example, even if they never touched anything or dropped hair, etc.
DNA eventually settles onto surfaces, and becomes part of dust. You could presumably take the dust from a room and build a genomic record of all the people who have entered that room over the span of many years.
Privacy concerns for all this, of course, but also extremely useful for ambient environmental monitoring / figuring out where pathogens are spreading / tracking animals in the wild.