Top Tweets for #dddbench
Couldn't hold it
our InsilicoMMAI-4B is 1st🥇vs Chemeleon/etc on a separate 14-task ADMET experimental holdout from @InSilicoMeds programs.
🏆Additional test for generalizability under distribution shifts, where it shines the best!
On #DDDBench soon, stay tuned 👀
#insilicoSOTAFM
🧬 Rentosertib went from an AI-designed molecule to patients and recently showed shifts toward younger predicted biological age across 6 proteomic aging clocks.
The first experiment is done. Now you can try our new SOTA+ drug discovery models yourself.
InsilicoMMAI-Chem-ADMET achieves SOTA+ performance on 7 drug safety and metabolism prediction tasks. 🚀
It’s a 4B language model trained on 140K+ datapoints in our MMAI Gym for Science.
#insilicoSOTAFM
It was a pleasure to work with #LFM by @liquidai @ramin_m_h to train our new #insilicoSOTAFM model for single-step retrosynthesis 🧪on ~46M reactions! It provides plenty of diverse reactions that only partially intersect with the reaction space by previous SOTA small models! 🔎👀
2.6B parameters beating purpose-built models in single-step retrosynthesis! 🚀
High time to ensemble these fast SOTA LLMs with search tree algorithms to power the ultimate multi-step retrosynthesis engine. #DDDBench #ChemCensor
🪶 What can a 2.6B-parameter language model achieve in retrosynthesis?
With a leak-proof, broad-coverage drug discovery benchmark #DDDBench, it becomes easier to evolve general-purpose models into SOTA-level specialists.
#insilicoSOTAFM

We’ve updated the #DDDBench with our own models, including LFM2-MMAI-Chem-SSRS-2.0, our single-step retrosynthesis specialist and the current SOTA on #ChemCensor.

🪶 What can a 2.6B-parameter language model achieve in retrosynthesis?
With a leak-proof, broad-coverage drug discovery benchmark #DDDBench, it becomes easier to evolve general-purpose models into SOTA-level specialists.
#insilicoSOTAFM

�� 5 new drug-discovery specialist LLMs. SOTA-level performance across 70+ benchmark tasks.
Insilico Medicine has released a series of frontier Specialist Language Models for chemistry and biology, trained through MMAI Gym.
They cover drug safety, potency prediction, chemical synthesis, and biology.
🧬 From general-purpose models to scientific specialists.
See for yourself, everything's live - DDDBench scores, extended LinkedIn posts: 50 strict wins, 70+ SOTA-level, the examples, all of it if you want the numbers.
https://t.co/SWJTMJQ1Uv
https://t.co/TMMZ8VRvvA
https://t.co/l9bfo5XV5q
#DDDBench #insilicoSOTAFM

4/n
📄 Read the paper: https://t.co/JXVfb3PmCQ
🤗 Join the discussion on Hugging Face: https://t.co/3YqsGCX1aP
📊 Recent #ChemCensor benchmark results on #DDDBench: https://t.co/5MmJq51CjX
Kudos to team @sumrexromanus @chem_logos @chemcensor @mathieu_reymond @VladAladin @AlexAliper @biogerontology
One takeaway message is that Anthropic ramped up biology and medicine research. We already see this trend for Opus 5 on the #DDDBench https://t.co/SWJTMJPu4X
Will it be enough to actually make an impact on drug discovery and cure cancer? Only years can tell.
2/2 Second, on the messaging around AI. I do not agree that my messaging has been disproportionately negative. In fact it has been about equally balanced between risks and benefits: I’ve written one major essay about each, and even in interviews where I discuss the risks, I make sure to frequently mention the incredible benefits as well as proposing possible solutions to the risks (short clips from my interviews that end up on social media tend to be disproportionately negative, as that gets clicks). In fact, I wrote Machines of Loving Grace because I didn’t feel the AI industry was painting an inspiring enough picture of how the technology could radically transform the world for the better. The bulk of the essay is devoted to refuting skepticism of AI’s potential in health and biology, and showing why I think it will actually be possible to cure most human disease in ~5-10 years, as crazy as it may sound to ordinary people and frankly to biologists as well (I used to be one!). And, if you read my most recent essay (Policy on the AI Exponential), I discuss concrete proposals for how to streamline the FDA process to make sure the deluge of AI-accelerated drugs isn’t slowed down by the regulatory process. I feel the urgency here: I lost my father to Hepatitis C only a few years before the development of direct-acting antivirals (sofosbuvir), which cure 95% of patients and probably would have cured him.
I do agree that the public has a negative view of AI (and that this is a big problem), but I don’t think it is primarily caused by me or any other AI leader warning about AI’s risks. I think it is fundamentally a crisis of trust. I think that ordinary people don’t trust companies, governments, or the tech industry and always suspect that we are cooking up some new way to screw them over. The causes of this go back decades and AI is just the latest iteration of it. I don’t think that a glitzy marketing campaign with a positive spin (which some have advocated that Anthropic do) is the way to win back that trust — at this point, saying that AI will cure cancer is more a cliche than it is inspiring, and most people think it is deceptive. The thing that will work is *actually curing cancer*. I think by far the most accurate criticism of AI companies including Anthropic is that we haven’t yet delivered on our big promises to benefit the world. That is totally on us, and I think it’s the criticism you should be making, instead of all this stuff about messaging and marketing.
We are however doing our best to fix this: Anthropic is ramping up its efforts very quickly in biology and medicine, and we hope to have incredible results in the coming years and some early glimmers in the coming months. When we’ve actually accomplished something real, the whole world will hear about it, as loudly as possible, you have my word on that. But until then I don’t want to make empty promises, and in the meantime I feel compelled to speak honestly about the very real risks of AI and how to address them. Honesty is the right thing on the merits, and in terms of public credibility and trust it is no worse than, and may in fact be better than, an approach that ignores or distracts from risks which people instinctively understand are real.
If you think that structural biology is solved after AlphaFold, you are likely wrong. There is still a room for improvement, especially to make high quality explanations on amino acids mutations. Stay tuned and there will be new #DDDbench category very soon at our website.
Can frontier models reason beyond a protein's sequence?
At @InSilicoMeds, we tested Qwen3.8-Max on bovine rhodopsin.
✅ Correctly identified that E113→Q is disruptive because E113 is the counterion that stabilizes the protonated retinal Schiff base (reasoning in the comments.).
✅Correctly inferred that L112, particularly L112I, is relatively mutation-tolerant.
⁉️Missed the key structural insight: L112 faces the lipid bilayer rather than the retinal-binding pocket, providing the mechanistic explanation for its higher mutation tolerance.
Strong biochemical intuition (great job @Alibaba_Qwen), but still room to improve understanding of the 3D structural context that governs protein function.
#AI #StructuralBiology #GPCR #ProteinDesign #Bioinformatics #LLM #Qwen #DDDBench #StructBioBench

The "one model to rule them all" narrative doesn't hold up in drug discovery so far.
Our updated #DDDBench leaderboard shows a highly fragmented landscape where different frontier models dominate entirely different stages of the pipeline
🚀 How do frontier AI models differ across drug discovery?
We just updated our Drug Discovery & Development Benchmark portal with two new frontier models: GPT-5.6-Sol and Claude Opus-5.
The surprising result: there is no single winner across all categories.
🧪 Claude Opus-5 leads in clinical trial prediction @AnthropicAI
🧬 GPT-5.6-Sol dominates large therapeutic molecules @OpenAI
⏳ Grok 4.5 takes longevity @SpaceXAI
However, in many categories, frontier models have not yet surpassed specialist models.
Explore the benchmark, compare performance, and test your own models.

Meet our #DDDBench for synthetic chemistry 🧪! Here are new results for Opus 5.0, Kimi K3, 5.6 Sol ☀️, Luna 🌒 and Terra🌎 ! The Insilico Index is heavily dependent on our #URSAbench with primary focus on multi-step synthesis planning. Kudos 🏆 to @sama @maggie_hott @joyjiao12

Meet our #DDDBench for synthetic chemistry 🧪! Here are new results for Opus 5.0, Kimi K3, 5.6 Sol ☀️, Luna 🌒 and Terra🌎 ! The Insilico Index is heavily dependent on our #URSAbench with primary focus on multi-step synthesis planning. Kudos 🏆 to @sama @maggie_hott @joyjiao12

Last Seen Hashtags on Sotwe
Most Popular Users

Elon Musk 
@elonmusk
241.6M followers

Barack Obama 
@barackobama
119M followers

Cristiano Ronaldo 
@cristiano
114.2M followers

Donald J. Trump 
@realdonaldtrump
111.8M followers

Narendra Modi 
@narendramodi
107.2M followers

Rihanna 
@rihanna
98.7M followers

NASA 
@nasa
92.4M followers

Justin Bieber 
@justinbieber
91.8M followers

KATY PERRY 
@katyperry
89.9M followers

Taylor Swift 
@taylorswift13
83.8M followers

Lady Gaga 
@ladygaga
75.3M followers

Virat Kohli 
@imvkohli
73.2M followers

Kim Kardashian 
@kimkardashian
70.8M followers

YouTube 
@youtube
68.8M followers

Neymar Jr 
@neymarjr
66.2M followers

Bill Gates 
@billgates
65M followers

Selena Gomez 
@selenagomez
63M followers

The Ellen Show
@theellenshow
62.3M followers

CNN 
@cnn
61.8M followers

X 
@x
60.7M followers






