I’m really excited to share our new preprint, out today, where we built a fully autonomous system for enzyme engineering — integrating a self-driving lab with a generative protein language model that worked together over a month to engineer substrate specificity in glycoside hydrolase (GH1) enzymes.
https://t.co/TPnh2IAASX
[1/n]
The outputs of different protein ML models often provide orthogonal signals for optimization. In reality, you may want to engineer towards several of these signals concurrently.
Navigating multidimensional model output space with utopia point optimization is effective, data efficient, and produces enzymes optimized for multiple characteristics at once 🧬
I am excited to share our preprint from @romerolab1 and @Boehringer collaboration!
We combined 3 ML models to engineer the ketoreductase Gre2, and we found 11 of 15 designs improved at least 3 measured properties in a single design–test cycle.
https://t.co/w3PBKixo3g
I am excited to share our preprint from @romerolab1 and @Boehringer collaboration!
We combined 3 ML models to engineer the ketoreductase Gre2, and we found 11 of 15 designs improved at least 3 measured properties in a single design–test cycle.
https://t.co/w3PBKixo3g
We’re partnering with @Anthropic to launch the biggest Protein Design Competition in the world, challenging people around the world to use AI to design new potential drug candidates for diseases that affect millions of lives.
The competition will feature five challenges, each focused on a specific disease or biological mechanism. Compared to previous competitions, it will be a big step-up in complexity and scale to push the boundaries of AI-driven protein design.
Together with Anthropic, we’re sponsoring over $1 million in experimental validation, making it possible to test more than 5,000 protein designs in our automated lab at no cost to participants. Anthropic is providing an additional $1 million in Claude credits.
All experimental results will be published openly on @Proteinbase, including designs that didn’t work, so anyone can access the data and build on what we learn.
The competition is open to everyone and free to enter. It will feature 3 tracks:
- Track 1 is aimed at expert protein designers, with up to 20 teams to be selected.
- Track 2 is targeting life science academics and industry researchers.
- Track 3 is open to everyone from tech enthusiasts to high-school students.
By combining Anthropic’s models with access to our automated lab, we want to make it possible for anyone with a laptop and an internet connection to join the global effort to advance human health with AI.
A big thanks to @Modal for contributing compute for protein design and to @TwistBioscience for contributing the DNA for the experimental validation!
Sign up link below -
We built high-throughput materials labs in Menlo Park to create a loop between experiments and models. The labs generate fresh data, the models learn from it, and then help us decide what to try next.
Using only 1,300 H200s, plus months of our experimental data, we mid-trained and RL’d an open-source model to surpass GPT-6 Astra on our analysis benchmark. We call it Neon.
This is real footage from our lab. We’re focusing first on hard problems in materials science, including superconductors, magnets, and semiconductor materials.
Read our blog posts below.
We built high-throughput materials labs in Menlo Park to create a loop between experiments and models. The labs generate fresh data, the models learn from it, and then help us decide what to try next.
Using only 1,300 H200s, plus months of our experimental data, we mid-trained and RL’d an open-source model to surpass GPT-6 Astra on our analysis benchmark. We call it Neon.
This is real footage from our lab. We’re focusing first on hard problems in materials science, including superconductors, magnets, and semiconductor materials.
Read our blog posts below.
We in @romerolab1 think a lot about how best to imbue deep learning models with some idea of experimentally-grounded protein function.
@NathanielBlalo2 and co made the connection that you could steer protein language models with experimental data in the same way you fine tune normal language models with human-preferred responses.
It’s been fun watching this idea evolve over the years. Huge congrats!
really excited to share our work is in press @NatureComms today. when i met with @romerolab1 back in sep 2022, hoping to rotate in the lab, one of my first questions was “have you thought about RLHF?” his eyes lit up and he said “YES!”. fast forward to sep 2026, i can truly say this project challenged me in every way possible. this was a real labor of love. i learned so much about myself, the field, and where it’s heading working on this with @NathanielBlalo2 and the rest of my amazing colleagues. i hope you find this work insightful if you happen to give it a read.
https://t.co/M4ZpVAafgU
really excited to share our work is in press @NatureComms today. when i met with @romerolab1 back in sep 2022, hoping to rotate in the lab, one of my first questions was “have you thought about RLHF?” his eyes lit up and he said “YES!”. fast forward to sep 2026, i can truly say this project challenged me in every way possible. this was a real labor of love. i learned so much about myself, the field, and where it’s heading working on this with @NathanielBlalo2 and the rest of my amazing colleagues. i hope you find this work insightful if you happen to give it a read.
https://t.co/M4ZpVAafgU
Thanks! The fridge is actually super simple - basically a minifridge with a linear actuator bolted on it, some 3d printed plate holders inside, and controlled by an arduino. You can find some more details in the GitHub under environment/hardware/auto_fridge https://t.co/7jmhAU6RHZ
I’m really excited to share our new preprint, out today, where we built a fully autonomous system for enzyme engineering — integrating a self-driving lab with a generative protein language model that worked together over a month to engineer substrate specificity in glycoside hydrolase (GH1) enzymes.
https://t.co/TPnh2IAASX
[1/n]
Protein design used to feel like a gacha game. Generative models could guarantee that the probability of a functional sequence was not zero, but they could not guarantee the property you actually cared about. You synthesized, assayed, and hoped. Few-shot learning, reinforcement learning, and other post-training methods can steer directed evolution to some extent, yet the rate-limiting step remains the same: making the proteins in the physical world and feeding experimental outcomes back to the model. That loop is slow and the information gain per cycle is limited.
A growing body of work now targets exactly this bottleneck: faster synthesis and wet-lab validation at a scale that can close the loop. Large industrial automation facilities already generate continuous streams of data. The robotic system here is more modest, but it still changes the tempo of a research lab. While a human team is still drawing the next set of cards, the platform can finish a plate of assembled, expressed, and activity-tested variants and return those measurements to the shared model for the next design round.
Closed-loop automation may arrive faster than many of us expected. Reading this paper is therefore both pressure and excitement.
We raised a $40M Series A to build the automated lab for agentic biology.
Last week, we shared the work we did with @Anthropic: Claude designed proteins, sent them to our automated wet lab, and got real experimental data back.
We believe this is where biology is going: AI agents designing experiments, running them in automated wet labs, learning from the results and iterating.
But AI can only move as fast as the experiments behind it. To use the potential of AI to cure all diseases, we need to build high-throughput wet lab infrastructure: a “biological gigafactory”.
That’s exactly what we’re doing at Adaptyv.
Over the past year we’ve grown our lab throughput by over 5x and onboarded more than 100 customers, ranging from bio AI labs like @chaidiscovery and @boltz_bio to pharmas like @Roche and @novonordisk to dozens and dozens of new startups that use AI to radically speed up drug discovery.
Binder design is nice and all, but here we have three agents sharing data and autonomously controlling a lab to design enzymes with shifted substrate scopes and high activity!
@cobanbrooks@NotinPascal@romerolab1
What if AI could interact directly with biology? Congrats to @cobanbrooks who gave AI the ability to run experiments & learn from feedback. Over 25 autonomous rounds, it learned how protein sequence controls enzyme specificity & discovered enzymes w/ new substrate preferences. 1/
Here, the assay had already been developed and optimized before it was handed to the robot. So the assay became essentially a static program it could run and read out. This is great when you’re testing variants of the same enzyme on the same reaction.
Many more degrees of freedom open up when an automated system has to first discover an assay and optimize it before it can be used. So it’s not so much hardware-limited but rather time and cost and reagents. But you can imagine a future where a network of self-driving labs delegates assay development, expression optimization, etc to different units and basically works in concert as a fully-staffed academic lab might work.
Many drugs work by binding to a specific target in the body and blocking or changing what it does. An important first step in the drug development process is designing a molecule that can bind tightly to its target. Traditionally, that's meant weeks or months of expert work per target, sifting through a large number of candidates to identify the few that work.
We wanted to test if Claude could successfully design novel protein binders from scratch (also called de novo design). With a protein design prompt written by a human expert, Claude autonomously designed protein binders against 14 out of 15 targets.
We then worked with Adaptyv Bio and Twist Bioscience, who independently built and tested the proteins Claude designed.
Anthropic benchmarked their newest Claude models on protein engineering tasks, and we at @adaptyvbio ran the wet lab work behind it.
They picked 16 targets from our past protein design competitions on @proteinbase and sent us an anonymised list of designs, so we had no idea which model produced which sequence. We then ran those protein designs through our automated wet lab workflow to characterize protein expression and binding affinities.
Here are the results:
- 95% of the designs expressed, which three years ago would have been the headline on its own. That matches the best expert expression rates from our competitions
- 354 of 1,320 designs bound their target on SPR, a 26.8% hit rate overall. Claude got binders on 14 of the 15 targets we could evaluate. Compared to our own data from public competitions, Claude beat de novo hit rates, for one target by more than three times the rate on Proteinbase
So can AI now solve all diseases? Well no, not yet.
This case study shows that Claude is at least expert-level at orchestrating protein design tools. That’s great news, since protein design tools are hard to use. Before AI, even setting up a protein structure prediction tool could take hours of debugging opaque conda errors.
Of course, those proteins that we tested here are not real therapeutics. They completed only the first step of the process: demonstrating that they can function as binders. Still, this study shows a path towards making actual therapeutics with AI.
Imagine making a drug is like climbing a mountain. We have clearly been able to climb some mountains, as humanity has made many drugs already. But the way to the top is a dangerous narrow path and climbing it takes many years and costs billions of dollars (and the lives of many biotechs).
The goal of AI for drug discovery is turning this narrow mountain path into a highway, making it easier and cheaper to get to the top so that we can develop 100x more therapeutics than we have right now. Similarly, writing code was a more of a high-expertise craft before LLMs, now it’s mostly automated and it has made generating software accessible for anyone.
With the recent news about the expansion of cloud labs around the US, we really think there’s an incredible opportunity for scaling here, where networks of autonomous labs could share experimental experience, split scientific questions into complementary subproblems, and test competing hypotheses in parallel. Ultimately, we believe strongly that rooting biological AI systems in the real world through experimental interaction is necessary in establishing a new era of biological discovery.
[5/n]
We propose a way out of this regime: instead of being observational, biological AI should be interactive. It should be able to perturb biological systems, measure the response, and learn from feedback. This, of course, is how the vast majority of knowledge in history has been generated, the same process having gone by many different names: the scientific method, the engineering design cycle, the Socratic method. Even evolution by natural selection operates a cycle of design and implementation.
[3/n]
Aside from just producing engineered proteins, we can also use the collected data to generate mechanistic hypotheses about enzyme function -- we can look at the chimera fragments to see which ones had the largest effects on substrate preference and further decompose these into residue-level effects on the protein structure.
[4c/n]