Promising results, but what about the most common drugs, small molecules?
#Bench3DFit from @InSilicoMeds challenges models to generate 3D molecule binders for proteins.
Opus-5 makes mistakes like clashes, small ligands and under-filling the pocket, but it’s come a long way!
Many drugs work by binding to a specific target in the body and blocking or changing what it does. An important first step in the drug development process is designing a molecule that can bind tightly to its target. Traditionally, that's meant weeks or months of expert work per target, sifting through a large number of candidates to identify the few that work.
We wanted to test if Claude could successfully design novel protein binders from scratch (also called de novo design). With a protein design prompt written by a human expert, Claude autonomously designed protein binders against 14 out of 15 targets.
We then worked with Adaptyv Bio and Twist Bioscience, who independently built and tested the proteins Claude designed.
GPT 5.6 Sol is another massive jump in 3D drug design ability!
Why does #Bench3DFit test text models on 3D output? Why not call a 3D generative tool?
It’s because understanding 3D interactions is needed for high-level DD goals and could even motivate more efficient tool calls!
🤔 GPT 5’s evolution on #Bench3DFit - what does it look like? We asked ourselves this after seeing Opus 5.0’s dramatic breakthrough.
🚀 The answer: constant major leaps - straight from zero to hero.
On the PLINDER test set, GPT-5.4 → 5.5 → 5.6 Sol shows a remarkably consistent trajectory:
🧪 Intramolecular pose validity rises from almost zero to a clear majority
🧩 More protein–ligand placements satisfy physical constraints
⚡️ UniDock scores shift into well-optimized territory
📈 This isn’t one isolated spike, but sustained progress.
Unlike Opus 5.0’s single dramatic jump, GPT-5 family shows sustained gains across releases. @OpenAI’s models are steadily improving physical-chemistry alignment and the reliability of realistic 3D ligand poses.
🤯 Excited to see what the next GPT version brings. Will it surpass diffusion models?
https://t.co/Y6x3xIFBSX
Same Target, Same Prompt, 4 Claude Opus models from @AnthropicAI#Bench3DFit from @InSilicoMeds Challenges LLMs to generate 3D molecules in a pocket using text.
Opus 5 reasons about connectivity, bond lengths, angles, protein interactions and more. A HUGE leap in 3D design!
Frontier LLMs are trained on basically the entire internet, so we need to be careful with how we benchmark them
The future is expert-curated benchmarks for socially and economically useful tasks on data they haven’t seen
Insilico is bringing this idea to Drug Discovery!
Think your foundation model is good at drug discovery?🧪
Prove it on clean data and earn your place on the leaderboard.
Every frontier model is already there. None of them wins everything.
Insilico DDD Benchmark. Submissions are open.
1/ 🧵
🤔 Does “add more reasoning” actually improve LLM-generated 3D ligands?
🧠 We tested #Opus 4.8 on the novel 3D-Fit benchmark using the PLINDER subset and comparing High, Medium, and Low reasoning levels.
🤯The result: more reasoning was not consistently better.
📉🔧"High" reasoning level produced more explicit geometry work — vector math, tetrahedral placement, clash checks — but this did not reliably lead to "raw" poses near the global optimum. And after UniDock local optimization, its advantage mostly vanished.
🧩The takeaway:
🧠⚠️The bottleneck isn’t just “add more reasoning.”
🎯❌Current LLMs can handle local geometric constraints, but still struggle with robust, global ligand placement in the binding pocket.
👇 Read more about 3D-Fit
Grok 4.5 just took the crown. 🚀
On DILI, the challenge of predicting drug-induced liver toxicity and one of the biggest causes of drug 💊failures, Grok 4.5 is now the #1 🥇model, outperforming every frontier flagship of the same LLM's generation.
Looking forward to seeing what Grok 4.6 can do in comparison to 5.6 Sol and Opus 5.0. @elonmusk
We're giving away 200 #LTXGoldSample bundles with our friends @LinusTech and @CoolerMaster!
Follow & RT to enter to win:
✅Intel Core i9-10900K "Golden Sample" up to 5.7Ghz
✅Cooler Master ML360 Sub-Zero powered by Intel Cryo Cooling Technology
Rules: https://t.co/qETkkjM3Ra
@zhaoxiaq Hi! I just read your tutorial on predicting molecular activity with Tensorflow. It was great! I was wondering if you could send me the code and supplementary info? The github link gives a 404. Thanks!