Prioritizing and interpreting disease-associated genetic variants remains one of the greatest challenges in human genetics. Today, we’re thrilled to introduce AlphaGenome Atlas 🧬, a genome-wide platform providing precomputed predictions for the regulatory effects of all ~9 billion possible single-letter changes and >100M observed indels in the human genome.
Here is what Atlas delivers:
1. Variant Prioritization via AVI
To prioritize variants, we developed the AlphaGenome Variant Impact (AVI) score. AVI predicts a unified score per variant, where higher values indicate greater disruption. It achieves state-of-the-art performance across diverse benchmarks. As a proof-of-concept with our collaborators at @broadinstitute, AVI prioritized a deep-intronic variant in DNM1, helping solve a previously unexplained rare epileptic encephalopathy case by revealing a brain-specific cryptic splice site.
2. Multi-Layer Molecular Interpretation
Variant prioritization is only half the battle; researchers also need to understand why a variant matters. Atlas decomposes variant effects across multiple interpretable layers:
•Feature Attributions which decompose each variant’s score into specific biological modalities driving the impact.
•Cell-Type Specificity: Precomputed predictions with AlphaGenome across hundreds of biosamples reveal the exact cellular context in which a variant acts.
•Regulatory Grammar: Over 2,600 de novo DNA motifs (and >250B genome-wide instances) show when variants directly disrupt critical regulatory binding "words".
Explore the resource:
🌐 Interactive browser & precomputed data: https://t.co/RtXWwMoUhv
- 🎥 Video: https://t.co/KAQk11KIFI
- 📖 Blog: https://t.co/2cJSgVF0aG
- 📄 Preprint: https://t.co/JATG5dESX2
New work I'm very excited about: Physics of Agents!🤖
As AI agents become more prevalent, they'll interact and influence one another.
Can we predict the collective behavior that emerges?
Surprisingly, their dynamics follow compact, predictive laws of statistical physics🧵
A few days at RECOMB - interesting research, learned a ton, and ran into an old friend I hadn't seen in 3 years. Conferencing by the sea in Thessaloniki is truly beautiful.
Today we're announcing ESMFold2, an open scientific engine to power prediction, design, and discovery across protein biology.
The new model delivers state of the art performance on protein interactions, especially antibodies, a critical modality for therapeutics.
We have designed and validated miniprotein binders and single chain antibodies across five therapeutic targets that are important in cancer and immunology. We are seeing very high success rates, and affinities at levels consistent with therapeutic activity.
We’re also releasing an atlas of 6.8 billion proteins, and 1.1 billion predicted structures.
ESMFold2 is built on a state of the art language model that has been trained on billions of protein sequences.
A world model of protein biology emerges through language modeling.
We’ve used the techniques of mechanistic interpretability developed to understand large language models to understand the concepts ESM uses to represent proteins.
The model’s representation space has a compositional organization of features across scales, levels of complexity, and abstraction, that reflects and mirrors the understanding of protein biology developed through a century of empirical science.
This understanding emerges without prior knowledge, just from language modeling of protein sequences.
Language models are becoming a powerful substrate to understand and program biology.
The design of protein interactions is one of the most fundamental problems in biophysics, and has critical implications for the discovery of new medicines. A simple gradient based search with the model was able to discover high-affinity protein binders.
I'm excited by the potential this has to accelerate basic science and the understanding of proteins. And especially for the new avenues it opens up for therapeutic design and medicine.
2026 may be the year AI starts to truly reason about biology.
AlphaFold helped close the sequence → structure gap.
The next frontier is sequence → functions.
Today, together with @genophoria and the team at @arcinstitute , we’re releasing BioReason-Pro — the first multimodal reasoning model for protein function prediction.
I made a Claude Code skill that generates conference posters 🛠️
Instead of a static PDF, it outputs a single HTML file — drag to resize columns, swap sections, adjust fonts, then give your layout back to Claude. 🔁
🔗 Skill 👉 https://t.co/KhYV8anbxL
AlphaGenome paper and models are out today! https://t.co/WHejipHfu7. We have seen great community engagement since we released the API. We hope to open more use cases with the model weights! https://t.co/TUfKPeOYyY
We updated the supplementary note with some very interesting case studies. I encourage everyone using AlphaGenome for non-coding variant interpretation to read it.
It was a huge team effort, I’m very proud of all of our co-authors @googledeepmind!
New paper “Proteome-wide model for human disease genetics” is now live at Nature Genetics: https://t.co/3UKcPlepDV
popEVE (https://t.co/HuxeGfe0g0) finds the needles in the haystacks of human genetic variation:
Announcing our new protein design server https://t.co/nHeMlJVba2:
• End-to-end protein design for everyone!
• Analyze your generated library interactively and on 3D structures
• Export codon-optimized DNA sequences for experimental testing.
Developed in collaboration between @deboramarks, @thomas_a_hopf, @SteineggerM, Simon d'Oelsnitz, Chris Sander, Artem Gazizov,@haysunny_hi, Milot Mirdita, Sergio Garcia Busto, Jake Reardon
We are pleased to share our new preprint: “De novo design of RNA and nucleoprotein complexes”.
This work extends the principles of de novo protein design to RNA and DNA, enabling the generative design of complex multi-polymer structures! (1/6)
https://t.co/QPqpNZr0er
Welcome to the age of generative genome design!
In 1977, Sanger et al. sequenced the first genome—of phage ΦX174.
Today, led by @samuelhking, we report the first AI-generated genomes. Using ΦX174 as a template, we made novel, high-fitness phages with genome language models. 🧵
Deep Generative Models Design mRNA Sequences with Enhanced Translational Capacity and Stability @ScienceMagazine
1. A groundbreaking study by He Zhang et al. introduces GEMORNA, a deep generative model that leverages Transformer architectures to design mRNA sequences with significantly enhanced translational capacity and stability. This innovation could revolutionize mRNA therapeutics by improving protein expression and durability.
2. GEMORNA addresses a critical challenge in mRNA design: the vast sequence space and complex interdependencies between optimization metrics. By using Transformer models tailored for mRNA coding sequences (CDSs) and untranslated regions (UTRs), GEMORNA generates sequences that outperform conventional designs in both in vitro and in vivo experiments.
3. The study demonstrates that GEMORNA-designed full-length mRNAs achieve up to a 41-fold increase in firefly luciferase expression compared to optimized benchmarks in vitro. Additionally, therapeutic mRNAs designed by GEMORNA show up to a 15-fold enhancement in human erythropoietin (EPO) expression and significantly higher antibody titers in mice.
4. GEMORNA’s versatility extends to circular RNA (circRNA) design, where it substantially enhances circRNA expression and boosts anti-tumor cytotoxicity in CAR-T cells. This highlights the potential of GEMORNA to improve the potency of mRNA drugs across various therapeutic areas.
5. The effectiveness of GEMORNA-generated sequences is confirmed through extensive experiments. The model autonomously learns codon and nucleotide usage, resulting in sequences with higher codon adaptation index (CAI), GC content, and lower rare codon rate, aligning with principles for designing therapeutic mRNAs.
6. GEMORNA also optimizes mRNA structure by balancing secondary structures, leading to enhanced stability and translational efficiency. The model’s ability to generate sequences with high naturalness scores further contributes to its success in creating mRNAs with strong performance.
7. The study provides detailed insights into the training and fine-tuning processes of GEMORNA models, emphasizing the importance of high-throughput data and accurate experimentation in validating optimal designs for specific applications.
📜Paper: https://t.co/yooxyzQu20
#mRNAdesign #DeepGenerativeModels #GEMORNA #mRNAtherapeutics #AIinBiology #Biotechnology