Yesterday, I built a cell model. 🧬 Today, it showed me how Ozempic works.
I asked it to explain the mechanism.
1 hour later... THIS. 🤯
GPT-6 Astra is insanely powerful. 🚀
I can’t stop @DeryaTR_@gdb
Now I’m wondering... can it model the rentosertib mechanism? 👀
Diversity is fuel for efficient single-step retrosynthesis.
Fine-tuning @liquidai's #LFM on 46M reactions yields higher unique reaction coverage than general-purpose LLMs, which is crucial for navigating retrosynthetic trees.
Full benchmark comparison on #DDDBench
It was a pleasure to work with #LFM by @liquidai@ramin_m_h to train our new #insilicoSOTAFM model for single-step retrosynthesis 🧪on ~46M reactions! It provides plenty of diverse reactions that only partially intersect with the reaction space by previous SOTA small models! 🔎👀
2.6B parameters beating purpose-built models in single-step retrosynthesis! 🚀
High time to ensemble these fast SOTA LLMs with search tree algorithms to power the ultimate multi-step retrosynthesis engine. #DDDBench#ChemCensor
🪶 What can a 2.6B-parameter language model achieve in retrosynthesis?
With a leak-proof, broad-coverage drug discovery benchmark #DDDBench, it becomes easier to evolve general-purpose models into SOTA-level specialists.
#insilicoSOTAFM
🚀 #O3DC is live: Stop scrolling 40 repos
The Open #DrugDiscovery & Development Consortium just dropped its first shared resource:
-115 benchmarks indexed across 10 categories and 15 consortia.
-Live GitHub status checks + honest notes on the caveats papers usually hide.
Open to academics, industry, and independents. Free to join.
Link in the comments
Can frontier models reason beyond a protein's sequence?
At @InSilicoMeds, we tested Qwen3.8-Max on bovine rhodopsin.
✅ Correctly identified that E113→Q is disruptive because E113 is the counterion that stabilizes the protonated retinal Schiff base (reasoning in the comments.).
✅Correctly inferred that L112, particularly L112I, is relatively mutation-tolerant.
⁉️Missed the key structural insight: L112 faces the lipid bilayer rather than the retinal-binding pocket, providing the mechanistic explanation for its higher mutation tolerance.
Strong biochemical intuition (great job @Alibaba_Qwen), but still room to improve understanding of the 3D structural context that governs protein function.
#AI #StructuralBiology #GPCR #ProteinDesign #Bioinformatics #LLM #Qwen #DDDBench #StructBioBench
The "one model to rule them all" narrative doesn't hold up in drug discovery so far.
Our updated #DDDBench leaderboard shows a highly fragmented landscape where different frontier models dominate entirely different stages of the pipeline
🚀 How do frontier AI models differ across drug discovery?
We just updated our Drug Discovery & Development Benchmark portal with two new frontier models: GPT-5.6-Sol and Claude Opus-5.
The surprising result: there is no single winner across all categories.
🧪 Claude Opus-5 leads in clinical trial prediction @AnthropicAI
🧬 GPT-5.6-Sol dominates large therapeutic molecules @OpenAI
⏳ Grok 4.5 takes longevity @SpaceXAI
However, in many categories, frontier models have not yet surpassed specialist models.
Explore the benchmark, compare performance, and test your own models.
Sometimes, I get a bit frustrated when people in AI who are far removed from chemistry don’t understand why we benchmark AI’s chemical capabilities. So let me try to bridge that gap by simplifying and explaining what chemical synthesis 🧪 is all about.
How does a molecule go from an idea on a computer to the lab, and perhaps one day to the pharmacy 💊?
This is where chemical synthesis comes in 🧪: the science of building molecules piece by piece, a bit like working with microscopic LEGO 🧱.
Imagine receiving a complicated LEGO set without the instruction booklet. You have a collection of smaller pieces and a picture of what you want to build, but you have to figure out which pieces to connect, in what order, and how to connect them🤔.
Chemistry makes this much more challenging⚗️.
Unlike LEGO bricks, the pieces used to build molecules do not simply click together. Some connect only at certain positions or in a particular sequence. Even small changes, such as the temperature, the solvent, or whether air is present, can determine whether the process works.
Chemists therefore have to write their own instruction booklet: a careful, step-by-step plan for building the molecule. Without chemical synthesis planning and execution, promising molecules would remain pictures on a screen. By making them real, chemists allow researchers to study them, test them, and explore whether they could one day help treat disease.
What chemistry term should we simplify next?
Curious how well AI can write these molecular "instruction booklets"?
We gave Claude Opus the molecular LEGO challenge: "Here's the target molecule. Now figure out how to build it." See how it performed in our #DDDBench portal in comparison with other frontier models! ⬇️
💊Claude Opus 5 leads drug absorption prediction across Frontier models.
For a drug taken as a pill 👅, one of the first challenges is getting past the gut wall and into the bloodstream, so it can reach the right organs and targets to treat disease.
This is a critical step in making medicines orally available instead of requiring injections 💉or other forms of delivery.
🚀 Opus 5 did an excellent job here, outperforming other frontier models.
Congrats @ClaudeAI, @mgdurrant, @AlecTPhD 👏
#DDDBenchmarks
@sumrexromanus Seeing Qwen catch up to Grok 4.5 level on synthesis planning shows how fast the gap is shrinking. Still some way to go to catch Sol 5.6, but the rate of progress is impressive🤯
Big update to the Insilico #DDDBenchmark (Drug Discovery and Development) portal.🧪
The portal now includes purpose-built specialist models alongside the latest frontier models, including Opus 5 and the GPT 5.6 family (☀️Sol, 🌎Terra, 🌓Luna), giving a more complete view of how general-purpose and domain-specific AI compare across drug discovery tasks.
The results are especially striking on our multi-step synthesis planning #URSAbench. GPT 5.6 Sol (High) is now the top-ranked LLM and ranks 🥉 #3 overall, behind only two dedicated retrosynthesis planners, @maggie_hott@OpenAI.
General-purpose models are closing the gap fast. But the specialist leader still sets the pace: #RetroChimera by @marwinsegler@MaziarzKris from @Microsoft remains the only model to break 30 points and continues to hold the 🥇#1 spot 💪. #AiZynthFinder by @AstraZeneca takes 🥈.
Remarkable evolution from Anthropic. This leap from Opus 4.8 to 5.0 on #URSAbench shows how rapidly LLMs are improving at synthesis planning.
The specialized tools still hold a lead.
Will next-generation models finally close the gap, or will hybrid approaches remain essential?
I am not a specialist in all of these disproved/proved conjectures and Erdős problems. But as a chemist I can see decisive improvements by the frontier models by @AnthropicAI to solve synthesis planning without ANY tools/orchestration.
Opus 5.0 has jumped 90% ! (10 to 19) on our hardest #URSAbench set from Opus 4.8! Amazing progress 🚀! Kudos to @mgdurrant@nc_frey@dagarfield@AlecTPhD 👏
Fresh update on our #DDDBench for Synthetic Chemistry. #GPT 5.6 Sol High claims #1 in #URSA benchmark, but with a top score of 38.7 (and an $840 price tag), general LLMs still have a long way to go in multi-step synthesis.
Meet our #DDDBench for synthetic chemistry 🧪! Here are new results for Opus 5.0, Kimi K3, 5.6 Sol ☀️, Luna 🌒 and Terra🌎 ! The Insilico Index is heavily dependent on our #URSAbench with primary focus on multi-step synthesis planning. Kudos 🏆 to @sama@maggie_hott@joyjiao12
Think your foundation model is good at drug discovery?🧪
Prove it on clean data and earn your place on the leaderboard.
Every frontier model is already there. None of them wins everything.
Insilico DDD Benchmark. Submissions are open.
1/ 🧵
Turns out "High reasoning" is sometimes a luxury setting.
We tested GPT-5.6 Sol High vs. Medium on antibody developability (Spearman).
Medium
💸 $85
📈 0.260
High
💸 $427
📈 0.326
5× the cost. +0.066 median Spearman.
The benchmark moved. The invoice moved more.
🧬 At @InSilicoMeds, we give frontier AI models a protein and ask one deceptively simple question. Actually, thousands.
Which parts can change, and which absolutely can't?
Some amino acids are tolerable for mutation. Others are so important that changing just one can break the protein.
Our Structural Biology benchmark measures how well AI can tell the difference using only the protein's structure and sequence context.
🏆 Grok led this run at predicting which changes proteins can tolerate, well done @grok
🛡️ Opus was close behind and stood out at identifying the amino acids most critical to protein function, strong result @AnthropicAI
Plenty of room for the leaderboard to change with the right structural and evolutionary training data. You know where to find it. 😉
We see this quite often: newer model iterations frequently regress on domain-specific tasks like chemistry and biologics. Newer doesn’t always mean better for specialized #benchmarks
🏆 GPT-5.5 isn’t done yet: still dominating biologics developability
The newer 5.6-sol model brings fresh capabilities, but historical developability prediction benchmarks show the older GPT-5.5 model leading across most tasks.
🧬 Scientific AI is about fit, not just scale.
Congrats @OpenAI@joyjiao12 🚀
This is also a major issue in retrosynthesis #benchmarks. General-purpose LLMs often perform remarkably well on existing test sets, mainly due to data leakage from USPTO, but underperform on novel chemistry from our #URSA benchmark
🧪 How do you know an AI is any good? You test it on things it has never seen, just like a real #exam.
For drug discovery AI, that means predicting properties of new molecules, not ones it already learned from.
We checked whether the standard "exams" for small molecule activity prediction actually contain new molecules.
Mostly, they don't. 👀
Continue reading👇
#Benchmark #InsilicoBench
🧪 How do you know an AI is any good? You test it on things it has never seen, just like a real #exam.
For drug discovery AI, that means predicting properties of new molecules, not ones it already learned from.
We checked whether the standard "exams" for small molecule activity prediction actually contain new molecules.
Mostly, they don't. 👀
Continue reading👇
#Benchmark #InsilicoBench
Excited to share a project I’ve been working on with my colleagues — we are introducing the URSA #benchmark to evaluate the real-world validity of synthetic routes!
🧪 Retrosynthesis #benchmarks are broken.
For ~9 years, they’ve rewarded patent matching or “reaching” purchasable building blocks, even through chemistry no chemist would trust.
The real test is simple: would a chemist take this route to the lab?
That’s why we built #URSA 🧵
Continue reading 👇