A benchmark earns trust when you test everyone and publish the results, whoever wins.
GPT-6 Astra: #1 on three DDD tasks, ahead of every other frontier general-purpose LLM.
Leaderboard: https://t.co/2d408zSKP2
#insilicoSOTAFM
We put GPT-6 Astra through our DDD benchmarks.
On 3 tasks, it crushed it. 🔥
It beat every other frontier model on chemical synthesis and antibody developability.
@sama, @gdb, @thekaransinghal Ready for the full suite? 👀
#insilicoSOTAFM
I trained a fly brain to 🧪synthesize Rentosertib, our anti-aging drug that had its first patient dosed in Phase III yesterday. 🪰
Even a fly uses #insilicoSOTAFM models
Accepted to the #EMNLP2026 Industry Track 🎉
@InSilicoMeds built MMAI Gym for Science. @liquidai built the Liquid Foundation Model. We combined them.
400+ drug-discovery tasks. Chemistry-native tokens. SFT, then RL.
Not a bigger model. A better gym. 🏋️
#insilicoSOTAFM 🧵👇
Yesterday, the first patient was dosed with Rentosertib in Phase III. 🧬
In Phase II, Rentosertib showed something unexpected: across 6 independent aging clocks, patients treated with the drug showed a shift toward a younger predicted biological age. 👀
Now, the first generative-AI-discovered drug to reach Phase III is taking its next step.
And so are we.
We’re introducing a suite of Insilico SOTA+ models for drug discovery, achieving top performance across 70+ drug discovery tasks. 🚀
We made the first step.
Now it’s your turn.
Test your models. Explore the benchmarks.
#insilicoSOTAFM
GPT-6 Astra is now the best-performing frontier model for antibody developability prediction in our benchmark. 🧬
Finding an antibody that binds is only the beginning. Will it aggregate? Will it stay stable? Can it actually behave like a drug?
GPT-6 Astra predicts these developability properties better than the other frontier models we tested. 🤯
And yes, the interactive visualization showing how an antibody actually works was also built with Astra in about 1 hour.
This model looks seriously powerful.
Congrats @gbt@sama
#insilicoSOTAFM
The detail worth noticing: it's a 2.6B model. Trained through MMAI Gym, it beats the best specialist by 7% and frontier models by 27% on single-step retrosynthesis on molecules it had never seen.
Check the results: https://t.co/2d408zSKP2
Can a language model become SOTA in chemical synthesis? 🧪🔥
We put a 2.6B-parameter @liquidai model into MMAI Gym.
It now sets the state of the art in single-step retrosynthesis. 🚀
#insilicoSOTAFM
HUGE!🚀Our specialist chemistry models trained in MMAI Gym are SOTA-level/superior vs Chemeleon/alike on ADMET/IC50 panels.
🏆On 50 tasks, a small language model (no tools) beats such strong baselines! Traditional ML's so over
SLM: S was for small but now for specialist-level💪
The real story here is the shift from general-purpose models to scientific specialists. Trained through MMAI Gym, evaluated across 70+ drug-discovery benchmarks.
Full results: https://t.co/2d408zSKP2
🚀 5 new drug-discovery specialist LLMs. SOTA-level performance across 70+ benchmark tasks.
Insilico Medicine has released a series of frontier Specialist Language Models for chemistry and biology, trained through MMAI Gym.
They cover drug safety, potency prediction, chemical synthesis, and biology.
🧬 From general-purpose models to scientific specialists.
I’m excited to share the follow-up paper to our #ChemCensor framework for evaluating single-step retrosynthesis 🧪
Our C3LM model was already competitive with, and in some cases superior to, general-purpose foundation models, but a gap remained between LLMs and conventional retrosynthesis models.
We finally closed and surpassed that gap 🚀
The key ingredients were… 👇
We should understand that neither Claude Opus nor Mythos themselves designed those binders. They used specialized models for this, and these are normal hit rates for BindCraft or Boltzgen, especially for previously seen proteins that have appeared in articles and competitions along with the binders.
🧬At @InSilicoMeds , one of our drugs, Rentosertib, works by switching off a protein called TNIK.
Here's the fun part: a 2025 @Nature study found that 12 weeks of resistance training left a similar fingerprint in human muscle, less TNIK activity, just from lifting weights.
Same target, hit two completely different ways. One is a molecule designed by #AI. The other is a workout.
Drug discovery meets exercise science, with a real overlap into longevity research too.
Simulation on one side. A TT bike on Hudayriyat🇦🇪 on the other (just a few hours before completing another #DDDBenchmarks task)
@TeamEmiratesUAE - thanks for the inspiration to keep riding even through hot summer mornings
#AIDrugDiscovery #TNIK #ExerciseScience #Longevity #ProteinSimulation #Cycling #TTBike #Hudayriyat #AbuDhabi
Pocket-conditioned ligand generation is like finding the right key for a lock. 🔐
The protein pocket is the lock 🔒, and the ligand is the key 🔑. The key needs the right shape and size to fit into the pocket without colliding with its walls.
But fitting the pocket is only the first part of the challenge. 📐 The teeth of the key need to line up with the right parts of the lock - just like a ligand needs to position the right functional groups at the right locations to make specific interactions with protein residues.
In practice, chemists specify these spatial requirements using different abstractions. An anchor fragment 🧩 is like a part of the molecule that you already know has to be there. A pharmacophore point 🎯 marks where a particular type of chemical feature needs to be positioned. An interaction ⚡ represents a specific contact the ligand needs to make with the pocket.
Diffusion models 🌊 have traditionally dominated 3D molecular generation, but integrating multiple heterogeneous conditions is often non-trivial. LLMs 🤖 offer a compelling alternative: they can naturally combine these diverse 3D instructions through a common language interface.
We explore this in 3D-Fit 🧬, a benchmark for 3D ligand generation under realistic spatial constraints. The results show that modern foundation models can follow local 3D constraints surprisingly well, but satisfying these constraints does not always translate into physically plausible poses or strong docking scores.
Read more in the #Bench3DFit paper: https://t.co/yTs2OFstV8
Explore recent results in #DDDBench: https://t.co/oYrb7wpEcS
Insilico Medicine introduces the Virtual Aging Cell
We think of age as a number. 🔢 Biology doesn’t. 🧬
The Virtual Aging Cell (VAC) is built to explore how cellular states change across time ⏳ and how those changes shape response to intervention.
The goal isn’t just to measure aging. It’s to predict what comes next 🔬
🧵1/5
🧬 Grok 4.6 just took 3rd place in our protein antibody stickiness benchmark.
Grok 4.6 joins GPT-5.5 and GPT-5.6 Sol among the top performers, with 6 of 8 frontier models scoring above 0.5 Spearman correlation. @elonmusk
HAC retention time in this benchmark measures how strongly a protein interacts with a hydrophobic surface, providing a useful proxy for protein stickiness.
🚀 Frontier models are getting better at predicting properties of proteins that matter for real-world drug development.
#DDDBenchmarks
🚀 Opus 5.0 marks a breakthrough for frontier model-generated 3D ligands.
After performance stalled and even regressed from Opus 4.6 to 4.8, Opus 5.0 delivers a dramatic leap on the 3D-Fit benchmark's PLINDER test set:
🧪 Far more poses pass PoseBusters checks
⚡️ UniDock scores move into much more optimized space
🎯 Stronger alignment with the physical principles of chemistry
This is not just an incremental update. Opus 5.0 appears to make a real advance in generating plausible ligand poses. Impressive progress from @AnthropicAI! 🤯
👇 Explore #Bench3DFit at #DDDBench by @InSilicoMeds
What happens to a pill 💊 after it reaches the liver and why it is so important?
The liver breaks drugs down so the body can get rid of them.
Swallowed → Absorbed → Broken down → Cleared
One way we measure this is Clint MLM (intrinsic clearance in mice liver microsomes) -- how quickly a drug is broken down by mouse liver enzymes in a lab test 🧪. The faster 🏃♂️clearance, the less drug is left in the body, the less the effect of the drug.
Can frontier models predict this from molecular structure alone with no tools/orchestration❓
Opus 5 came out on top 🏆 for Clint MLM prediction, with a Spearman correlation of 0.425. Check our #DDDBench site to see how better is Opus 5 compared to other frontier models.
Can frontier models reason beyond a protein's sequence?
At @InSilicoMeds, we tested Qwen3.8-Max on bovine rhodopsin.
✅ Correctly identified that E113→Q is disruptive because E113 is the counterion that stabilizes the protonated retinal Schiff base (reasoning in the comments.).
✅Correctly inferred that L112, particularly L112I, is relatively mutation-tolerant.
⁉️Missed the key structural insight: L112 faces the lipid bilayer rather than the retinal-binding pocket, providing the mechanistic explanation for its higher mutation tolerance.
Strong biochemical intuition (great job @Alibaba_Qwen), but still room to improve understanding of the 3D structural context that governs protein function.
#AI #StructuralBiology #GPCR #ProteinDesign #Bioinformatics #LLM #Qwen #DDDBench #StructBioBench
🚀 How do frontier AI models differ across drug discovery?
We just updated our Drug Discovery & Development Benchmark portal with two new frontier models: GPT-5.6-Sol and Claude Opus-5.
The surprising result: there is no single winner across all categories.
🧪 Claude Opus-5 leads in clinical trial prediction @AnthropicAI
🧬 GPT-5.6-Sol dominates large therapeutic molecules @OpenAI
⏳ Grok 4.5 takes longevity @SpaceXAI
However, in many categories, frontier models have not yet surpassed specialist models.
Explore the benchmark, compare performance, and test your own models.
💊Claude Opus 5 leads drug absorption prediction across Frontier models.
For a drug taken as a pill 👅, one of the first challenges is getting past the gut wall and into the bloodstream, so it can reach the right organs and targets to treat disease.
This is a critical step in making medicines orally available instead of requiring injections 💉or other forms of delivery.
🚀 Opus 5 did an excellent job here, outperforming other frontier models.
Congrats @ClaudeAI, @mgdurrant, @AlecTPhD 👏
#DDDBenchmarks
Sometimes, I get a bit frustrated when people in AI who are far removed from chemistry don’t understand why we benchmark AI’s chemical capabilities. So let me try to bridge that gap by simplifying and explaining what chemical synthesis 🧪 is all about.
How does a molecule go from an idea on a computer to the lab, and perhaps one day to the pharmacy 💊?
This is where chemical synthesis comes in 🧪: the science of building molecules piece by piece, a bit like working with microscopic LEGO 🧱.
Imagine receiving a complicated LEGO set without the instruction booklet. You have a collection of smaller pieces and a picture of what you want to build, but you have to figure out which pieces to connect, in what order, and how to connect them🤔.
Chemistry makes this much more challenging⚗️.
Unlike LEGO bricks, the pieces used to build molecules do not simply click together. Some connect only at certain positions or in a particular sequence. Even small changes, such as the temperature, the solvent, or whether air is present, can determine whether the process works.
Chemists therefore have to write their own instruction booklet: a careful, step-by-step plan for building the molecule. Without chemical synthesis planning and execution, promising molecules would remain pictures on a screen. By making them real, chemists allow researchers to study them, test them, and explore whether they could one day help treat disease.
What chemistry term should we simplify next?
Curious how well AI can write these molecular "instruction booklets"?
We gave Claude Opus the molecular LEGO challenge: "Here's the target molecule. Now figure out how to build it." See how it performed in our #DDDBench portal in comparison with other frontier models! ⬇️