Introducing: continuous variational synthesis
📄 https://t.co/d0G6FEMQef
We’re pleased to announce a new advance in our ability to synthesize generative model-designed DNA sequences at petascale.
At long last, SynBioBeta decided to report on @jura_bio… or at least share the reporting that has happened in China. For those who’d like to read about our work but are stuck with English:
https://t.co/jA9LYrnUAt
On model ablations: Linear models plateau on this data (dashed), even with foundation model representations (colors). We needed big transformers (solid) to scale.
Excited to share TERRA, a tissue world model 🧬
Over ~1.5 years we ran a large data-generation + modelling effort to build a world model for human tissues, pretrained on 112M cells from spatial transcriptomics (mostly Xenium 5000-plex + public data).
It's built on one of the largest human spatial transcriptomics corpora assembled to date, spanning 20 tissues across development, health and 26 disease conditions, ~two-thirds newly generated in-house.
Why a "world model" for tissue? Images have universal representations (ViT/DINOv3), so do proteins (ESM, @alexrives) and pathology (UNI, @AI4Pathology). We've worked hard to build something similar for human tissue: one model that captures its multi-scale logic, genes → cells → their native microenvironments.
Like the JEPA approach @ylecun has championed, TERRA learns by prediction in embedding space, but for human tissue. How it works: it tokenises each cell together with its nearest neighbours into one sequence while keeping every gene's identity, then masks part of a neighbourhood and predicts the representation of the hidden part, not raw noisy counts. From one backbone it reads out three scales, gene embeddings (what a gene is doing in a cell and its niche), cell embeddings (cell type and state) and neighbourhood embeddings (the niche), and because it keeps gene-level resolution it can knock a gene out in silico and predict the response. Applied entirely zero-shot, TERRA maps and perturbs human tissue across unseen organs, diseases and technologies, outperforming existing spatial approaches.
Three take-homes:
1️⃣ One model, any tissue. A single pretrained backbone provides tissue representations zero-shot, handling genes, cells and niches across organs and platforms, off the shelf.
2️⃣ New biology, development to clinic. We built a new spatial atlas of the developing human pancreas and found an islet-associated capillary state that looks like a precursor of mature islet vasculature. In kidney, TERRA's in silico knockouts predicted the tissue-injury programme from cancer immunotherapy (checkpoint blockade), confirmed in treated kidneys, detected in blood, and linked to declining kidney function.
3️⃣ A grammar of tissue architecture. By coupling each cell's state to its niche, TERRA defines recurring cross-organ "archetypes" of macrophage neighbourhoods, including a tumour-boundary niche that tracks poor survival in kidney cancer.
TERRA is already in use: it powered our recent skin atlas of hidden immune-memory niches (https://t.co/LeWxOKkgMt), with more studies coming soon.
This was an amazing collaboration between clinicians, machine-learning scientists and cell biologists 🙏 Led by @SebastianBirk_, @ValiSanian@AmirhVahidi, Samuel Ogden, @daniyal_jafree, @Adib_m_, @CarloLeonardi7 and Arpit Merchant, with Lassi Paavolainen, Menna Clatworthy, @bayraktar_lab, @Muzz_Haniffa, Tom Mitchell and @bakhti_mostafa. Huge thanks too to everyone who shared data and helped along the way.
What excites me most is seeing how the community builds on this. The model, code and tutorials are all public, so anyone can run TERRA on their own tissues, extend it, or build new models on top. Huge thanks to the whole team across @sangerinstitute and our many collaborators.
📄 Paper: https://t.co/AOOIcwuTnq
💻 Code: https://t.co/ZBwTmd6SNI
🤗 Model: https://t.co/gyQA79lVfW
#SpatialTranscriptomics #SpatialGenomics #FoundationModels #AI4Science #MachineLearning #ComputationalBiology #SingleCell #WorldModels
Can we achieve the same scaling laws in biological AI as we have in the rest of machine learning?
We found that @jura_bio's mix of designed data generation and training produce robust scaling laws: more data reliably lead to better predictions over many orders of magnitude.
Thrilled to release two new preprints on intelligent labs for driving science and innovation. This is in close coordination with Aviv Regev, Jian Ma (@jmuiuc), Michelle Lee (@michellearning), and the teams at @Genentech, @SCSatCMU, and @Princeton University.
In our perspective, we argue that the next generation of labs should be human-in-the-lead and AI-empowered, integrating human intent, machine reasoning, and physical experimentation through scientific world models and an agentic harnessing layer.
Done responsibly, these systems have the potential to make scientific discovery more programmable, reproducible, adaptive, and scalable while enabling scientists to focus on higher-level scientific reasoning and discovery.
It's been a privilege to pursue this Perspective with @MengdiWang10 and an outstanding group of scientists and innovators advancing the intersection of computation, AI, science, and medicine. We're excited to continue exploring where this vision leads.
Our preprint: https://t.co/aKEkUIaNwb
Preprint led by Jian and Aviv team: https://t.co/yOzQ6cI9XD
This also kicks off a new Gladstone-Stanford AI Hub efforts, led by Katie Pollard at @GladstoneInst and myself, with an amazing team of scientists including Emma Lundberg (@Prof_Lundberg), Brian L Trippe (@brianltrippe ), Anshul Kundaje (@anshulkundaje ), Barbara Engelhardt (@BeEngelhardt), Christina Theodoris (@TheodorisLab), Catherine Tcheandjieu (@ines_catherine), Bruce Conklin, Alexander Marson, Stacie Dodgson (@StacieDodgson), Seth Shipman (@seth_shipman), Vijay Ramani, Danielle Swaney (@dlswaney), and Nevan Krogan. Excited to be building together across two great institutions!
Introducing Discovery Loop ♾. We’re building AI to run experiments at unprecedented scale and to solve the biggest bottlenecks in science and engineering.
Our founders @JeffDean, @Sanjay_Ghemawat, @quocleix, @OriolVinyalsML helped build the AI infrastructure and models that scaled the modern world. Now, we are scaling discovery itself.
The future can't wait.
Learn more at: https://t.co/6Fr8dX5dDB
if you compare @nvidia Proteina-Complexa
(1M designs, 127 targets) paper w @jura_bio Vista (209M designs, 100 targets), it looks like jura wins for capturing data (expected) but also for designing binders over targets?
IIUC jura did not train on PDB.