Today we're introducing continuous variational synthesis.
Spearheaded by @AlanNawzadAmin, cVS brings to manufacturing-aware architectures more of what we expect from modern generative model architectures, including pretraining and fine-tuning, and flexibility wrt hardware.
Introducing: continuous variational synthesis
📄 https://t.co/d0G6FEMQef
We’re pleased to announce a new advance in our ability to synthesize generative model-designed DNA sequences at petascale.
@SynBioBeta@jura_bio I am grateful that Chinese coverage of an American company is worth reporting and sharing, @SynBioBeta. But why not cover us yourselves sometime?
At long last, SynBioBeta decided to report on @jura_bio… or at least share the reporting that has happened in China. For those who’d like to read about our work but are stuck with English:
https://t.co/jA9LYrnUAt
@jura_bio reports a biological-AI scaling law built through lab-generated data. Its system screened a large sample from ~10¹⁶ antibody designs against 100 pHLA targets, with model performance improving predictably as data grew. Findings currently cover a specific antibody–pHLA task.
https://t.co/MRf5viDqyr
Subscribe to our newsletter and get the biggest biotech news straight to your inbox 🧬
https://t.co/uNrxaa7z6R
A lot happened in AI x bio over the past week. Here’s what you might’ve missed. 🧵 (1/8)
☑️ Demis Hassabis stepped back from running NanoBanana to go all in on science, AGI and Isomorphic. And one explosive report says he actually wanted to leave Google.
☑️ Meanwhile, four Google legends actually did leave - to build a startup that aims to automate scientific discovery.
☑️ The AI system Biomni spent five days doing months of biology-model research.
☑️ AI found a hidden chemotherapy benefit inside a failed cancer trial.
☑️ Jura Bio says it found a scaling law for biological AI.
☑️ The interpretability masters at Goodfire used a protein model to predict not only whether 2.1 million mutations may be harmful, but what they may break and why.
Nice! I think adding the linear baseline actually makes things MORE compelling because there is a flattening out of performance that your transformer based method seems impervious to.
So in some ways you discovered another scaling law but for linear models.
I wonder if the hype around linear baselines was simply that those evals were overdetermined and still in the regime of comparable performance!
On model ablations: Linear models plateau on this data (dashed), even with foundation model representations (colors). We needed big transformers (solid) to scale.
There's been a lot of talk about data scaling in bio lately but in a kind of underbaked way, as if just measuring more "stuff" will solve everything. Interesting to see a nice example in practice!
Very excited about our new results, showing variational synthesis + large scale measurements + co-designed training -> robust scaling laws for sequence-activity models.
Clearest evidence I've seen that a company's proprietary high-throughput data is *useful*. Would love to see other groups show the same evidence/ablations!
At @jura_bio we built some of the largest therapeutic antibody binding datasets. But does scaling data generation efforts guarantee proportional returns in model quality? We are thrilled to show that for our platform the answer is yes!
More biological details, sensitivity analysis, and details on the scaling rules and strategies we used to get this working are in our blog post:
https://t.co/jA9LYrnUAt
Does biological AI improve predictably as experimentally generated data grows? In new JURA results, held-out model loss follows a power law across nearly three orders of magnitude of designed laboratory data. The gains extend to therapeutic design tasks.
https://t.co/BHo9LTNhTd
Can we achieve the same scaling laws in biological AI as we have in the rest of machine learning?
We found that @jura_bio's mix of designed data generation and training produce robust scaling laws: more data reliably lead to better predictions over many orders of magnitude.
We were just talking about this today! I think we’re going to try to fit a variational synthesis model, at least.
Does someone want to donate the 8000 targets? We’ll handle the binders.
300K proteins, 30M data points — a good start! At @jura_bio we’re privileged to have made 100M, 10B data points a weekly routine. Everyone — come to where the scale is.