GPT-6 Astra⭐️outperformed 2 specialist models (LocalRetro and RetroKNN) in single-step retrosynthesis prediction. Unbelievable result just a month ago! It is the very first general-purpose model outperforming these stong specialists in our benchmark. Kudos to @gdb@joyjiao12 👏
In the #DDDBenchmark, the model ranks 3rd🥉overall, behind only 2 other specialist models including our model and MHNReact by @gklambauer@phseidl , while significantly outperforming all other frontier models (see in thread⬇️).
Btw, this synthesis visualization itself was also generated by GPT-6 Astra in just 15 minutes. We will see an explosion of 3D simulations of reactions thanks to models like Astra! 🧵1/2
#insilicoSOTAFM
Happy we released a new SOTA-level kinome model.
BTW, if you're interested in the design strategy behind rentosertib's selectivity, here's the write-up: https://t.co/Pct5VwAtPg
#insilicoSOTAFM
Built an app to show how Rentosertib actually works. 🧬✨
It recently showed an anti-aging effect in clinical trials. 👀
GPT-6 Astra again. The crazy part: ~40 minutes to build, plus an MD simulation 🤯
@Derya_tm@EvgenyKirilin, this really feels like a new level of structural analysis.
Rentosertib inhibits TNIK, a kinase and one of the molecular switches controlling signaling inside our cells. ⚙️ Kinase activity is also exactly the kind of property Insilico Medicine’s models are built to predict. 🤖🔬
#insilicoSOTAFM
It was a pleasure to work with #LFM by @liquidai@ramin_m_h to train our new #insilicoSOTAFM model for single-step retrosynthesis 🧪on ~46M reactions! It provides plenty of diverse reactions that only partially intersect with the reaction space by previous SOTA small models! 🔎👀
And the box isn’t a neat LEGO set — it’s 200,000 pieces where 99.999% belong to completely different models. No instructions, no picture on the lid, and half the bricks only snap together at 180 °C under argon.
That’s why chemical search space is the hard part, not the assembly
Sometimes, I get a bit frustrated when people in AI who are far removed from chemistry don’t understand why we benchmark AI’s chemical capabilities. So let me try to bridge that gap by simplifying and explaining what chemical synthesis 🧪 is all about.
How does a molecule go from an idea on a computer to the lab, and perhaps one day to the pharmacy 💊?
This is where chemical synthesis comes in 🧪: the science of building molecules piece by piece, a bit like working with microscopic LEGO 🧱.
Imagine receiving a complicated LEGO set without the instruction booklet. You have a collection of smaller pieces and a picture of what you want to build, but you have to figure out which pieces to connect, in what order, and how to connect them🤔.
Chemistry makes this much more challenging⚗️.
Unlike LEGO bricks, the pieces used to build molecules do not simply click together. Some connect only at certain positions or in a particular sequence. Even small changes, such as the temperature, the solvent, or whether air is present, can determine whether the process works.
Chemists therefore have to write their own instruction booklet: a careful, step-by-step plan for building the molecule. Without chemical synthesis planning and execution, promising molecules would remain pictures on a screen. By making them real, chemists allow researchers to study them, test them, and explore whether they could one day help treat disease.
What chemistry term should we simplify next?
Curious how well AI can write these molecular "instruction booklets"?
We gave Claude Opus the molecular LEGO challenge: "Here's the target molecule. Now figure out how to build it." See how it performed in our #DDDBench portal in comparison with other frontier models! ⬇️
@sumrexromanus That’s the real search problem in retrosynthesis: it’s not just “in what order do I connect things,” it’s “which few building blocks out of a catalogue of millions are the right ones at all.”
@sumrexromanus Great analogy — and I’d push it one step further.
Now imagine your LEGO Venator comes as a bin of 200,000 pieces, and 99.999% of them have nothing to do with a Venator. No instructions, and no way to tell which handful of bricks is even relevant.
Everyone is asking: how is new Qwen 3.8-Max in #URSAbench for synthesis planning ⁉️
1⃣ Well, it is much better than new Kimi K3
2⃣ Winter ❄️ (3.5) to summer, ☀️ Qwen's performance boosted +650% 🚀 , while Kimi is only +60%
3⃣ Qwen reached the level of Grok 4.5 (not bad), but beyond Opus (19%) and Sol 5.6 (26%).
Great job so far @Ali_TongyiLab@Alibaba_Qwen 👏
For more details, check our Benchmark Explorer at #DDDBench by @InSilicoMeds ⬇️
I am not a specialist in all of these disproved/proved conjectures and Erdős problems. But as a chemist I can see decisive improvements by the frontier models by @AnthropicAI to solve synthesis planning without ANY tools/orchestration.
Opus 5.0 has jumped 90% ! (10 to 19) on our hardest #URSAbench set from Opus 4.8! Amazing progress 🚀! Kudos to @mgdurrant@nc_frey@dagarfield@AlecTPhD 👏
Meet our #DDDBench for synthetic chemistry 🧪! Here are new results for Opus 5.0, Kimi K3, 5.6 Sol ☀️, Luna 🌒 and Terra🌎 ! The Insilico Index is heavily dependent on our #URSAbench with primary focus on multi-step synthesis planning. Kudos 🏆 to @sama@maggie_hott@joyjiao12
Thrilled to have contributed to #URSA! I worked on the codebase for this project, and I'm really proud to help bring a much-needed, realistic benchmark to the field of retrosynthesis. 💻🧪
🧪 Retrosynthesis #benchmarks are broken.
For ~9 years, they’ve rewarded patent matching or “reaching” purchasable building blocks, even through chemistry no chemist would trust.
The real test is simple: would a chemist take this route to the lab?
That’s why we built #URSA 🧵
Continue reading 👇