Senior Research Scientist @SandboxAQ, ex @GoodChemistryCo , Former Chem Graduate Student @ChemistryUIUC. Research on Drug Discovery and computational chemistry.
Free PDF of my new Optimization book:
https://t.co/2QQMQMJqr2
If you like it, check it out on Amazon or at Cambridge University Press!
Please leave a review and email me with any typos/corrections.
One can always trust FT to provide accurate representative images. In this case, they couldn't have done a better job of depicting IT employees in India.
We asked an unreleased research version of Claude to take a stab at the Riemann hypothesis.
It didn’t solve it, but it did make strides on a related problem: it increased the lower bound for the fraction of zeros of the Riemann zeta function that satisfy the hypothesis from 41.6% to 67.2%.
https://t.co/aZDvqqhHRi
After a multi-year hiatus I've written a new blog post. It tries to begin answering the question of what a meaningful human life will look like when thinking is best done by machines.
It doesn't directly have to do with proteins or bioML but it does speculate wildly about the biological future of humans. It's a bit absurd, may upset people, and may trigger ridicule. I wrote it because we are living through critical times and shouldn't be shy of considering ideas that may help shepherd us through these times just because they are weird or make us uncomfortable. It's my earnest belief that making space for unconventional conversations, and granting everyone a bit of grace, will increase our odds of not only surviving, but thriving, in the post-AGI era. Enjoy.
https://t.co/zYKNt56gYv
De novo design of a protein fold for small-molecule binding through aromatic π stacking
1 They computationally designed compact de novo proteins that bind the anticancer drug doxorubicin by anchoring the binding mode on a minimal aromatic π-stacking “Trp sandwich” motif, rather than starting from a known fold or repeat scaffold.
2 The workflow combined RFdiffusion All-Atom for backbone generation, LigandMPNN + Rosetta FastRelax for sequence/structure refinement, and structure-prediction-guided filtering (RaptorX-Single for apo folding; Chai-1 for holo geometry) to prioritize designs before experiments.
3 Starting from a 6-residue motif derived from an anthracycline-bound protein structure, they generated 500 backbones, filtered to 66, then selected 12 for experimental testing; 9 expressed solubly in E. coli, and one (DoxSand0) bound doxorubicin at low micromolar affinity (KD ~1.05 µM) without experimental screening campaigns.
4 The initial binder revealed a “jaw-like” 5-helix architecture (two helix-hairpins connected by a hinge helix) forming a buried cavity that positions doxorubicin between two opposing aromatics; early biophysics showed oligomerization and limited cooperative unfolding, motivating scaffold stabilization.
5 They improved scaffold quality by loop remodeling with MASTER (fragment-based loop replacement while preserving the binding geometry), shrinking the protein to 85 residues (DoxSand1) and increasing the monomeric fraction (29% → 51%) while maintaining similar affinity (KD ~0.78 µM).
6 Binding-site mutagenesis validated the designed recognition logic: the two aromatic residues are central for binding (W33 and W61 substitutions strongly reduced affinity when removed), while surrounding polar contacts (e.g., Asp65 near the amino sugar) contribute to stabilization.
7 A key optimization came from an unexpected pocket mutation: H73A improved binding into the nanomolar range (KD ~95 nM; globally fit KD ~58 nM), consistent with a subtle rearrangement that improved ligand burial and the aromatic interaction network.
8 Because H73A reduced monomeric fraction, they introduced a structure-guided V46I packing mutation to restore stability, yielding DoxSand2 with improved monomeric behavior (to ~61%) and strong binding (KD ~85 nM) in the same compact 85-aa scaffold.
9 A 1.56 Å X-ray crystal structure of the DoxSand2–doxorubicin complex closely matched the design (Cα RMSD ~0.57 Å vs Chai-1), confirming the intended π-π sandwich (W33/W61) and key contacts (including Asp65). The ligand pose was accurate overall (ligand RMSD ~1.34 Å) but showed small shifts that introduced additional interactions and an ordered water bridge, highlighting current limits in explicit solvent/pose modeling.
10 Foldseek searches against PDB100 and AlphaFold/UniProt50 found no statistically significant structural homologs, supporting that DoxSand2 adopts a previously unobserved compact 5-helix globular fold and suggesting small (<90 aa) functional fold space is not fully sampled by natural evolution.
11 In cells, the designed binders sequestered doxorubicin and reduced cytotoxicity in COV362 ovarian cancer cells: preincubation with 20 µM protein shifted apparent doxorubicin IC50 ~17-fold (DoxSand1) and ~140-fold (DoxSand2), consistent with affinity improvements translating to functional drug buffering.
12 Overall, the paper frames motif-guided generative design as a way to discover new compact folds for chemically complex small-molecule recognition, with iterative rational optimization improving stability and affinity without large-scale experimental affinity maturation.
📜Paper: https://t.co/hJwF8m0Ng0
#ProteinDesign #DeNovoProteinDesign #GenerativeAI #RFdiffusion #Rosetta #LigandMPNN #StructuralBiology #SmallMoleculeBinding #DrugSequestration #Doxorubicin #PiStacking #ComputationalBiology
De novo design of a protein fold for small-molecule binding through aromatic π stacking
1 They computationally designed compact de novo proteins that bind the anticancer drug doxorubicin by anchoring the binding mode on a minimal aromatic π-stacking “Trp sandwich” motif, rather than starting from a known fold or repeat scaffold.
2 The workflow combined RFdiffusion All-Atom for backbone generation, LigandMPNN + Rosetta FastRelax for sequence/structure refinement, and structure-prediction-guided filtering (RaptorX-Single for apo folding; Chai-1 for holo geometry) to prioritize designs before experiments.
3 Starting from a 6-residue motif derived from an anthracycline-bound protein structure, they generated 500 backbones, filtered to 66, then selected 12 for experimental testing; 9 expressed solubly in E. coli, and one (DoxSand0) bound doxorubicin at low micromolar affinity (KD ~1.05 µM) without experimental screening campaigns.
4 The initial binder revealed a “jaw-like” 5-helix architecture (two helix-hairpins connected by a hinge helix) forming a buried cavity that positions doxorubicin between two opposing aromatics; early biophysics showed oligomerization and limited cooperative unfolding, motivating scaffold stabilization.
5 They improved scaffold quality by loop remodeling with MASTER (fragment-based loop replacement while preserving the binding geometry), shrinking the protein to 85 residues (DoxSand1) and increasing the monomeric fraction (29% → 51%) while maintaining similar affinity (KD ~0.78 µM).
6 Binding-site mutagenesis validated the designed recognition logic: the two aromatic residues are central for binding (W33 and W61 substitutions strongly reduced affinity when removed), while surrounding polar contacts (e.g., Asp65 near the amino sugar) contribute to stabilization.
7 A key optimization came from an unexpected pocket mutation: H73A improved binding into the nanomolar range (KD ~95 nM; globally fit KD ~58 nM), consistent with a subtle rearrangement that improved ligand burial and the aromatic interaction network.
8 Because H73A reduced monomeric fraction, they introduced a structure-guided V46I packing mutation to restore stability, yielding DoxSand2 with improved monomeric behavior (to ~61%) and strong binding (KD ~85 nM) in the same compact 85-aa scaffold.
9 A 1.56 Å X-ray crystal structure of the DoxSand2–doxorubicin complex closely matched the design (Cα RMSD ~0.57 Å vs Chai-1), confirming the intended π-π sandwich (W33/W61) and key contacts (including Asp65). The ligand pose was accurate overall (ligand RMSD ~1.34 Å) but showed small shifts that introduced additional interactions and an ordered water bridge, highlighting current limits in explicit solvent/pose modeling.
10 Foldseek searches against PDB100 and AlphaFold/UniProt50 found no statistically significant structural homologs, supporting that DoxSand2 adopts a previously unobserved compact 5-helix globular fold and suggesting small (<90 aa) functional fold space is not fully sampled by natural evolution.
11 In cells, the designed binders sequestered doxorubicin and reduced cytotoxicity in COV362 ovarian cancer cells: preincubation with 20 µM protein shifted apparent doxorubicin IC50 ~17-fold (DoxSand1) and ~140-fold (DoxSand2), consistent with affinity improvements translating to functional drug buffering.
12 Overall, the paper frames motif-guided generative design as a way to discover new compact folds for chemically complex small-molecule recognition, with iterative rational optimization improving stability and affinity without large-scale experimental affinity maturation.
📜Paper: https://t.co/hJwF8m0Ng0
#ProteinDesign #DeNovoProteinDesign #GenerativeAI #RFdiffusion #Rosetta #LigandMPNN #StructuralBiology #SmallMoleculeBinding #DrugSequestration #Doxorubicin #PiStacking #ComputationalBiology
Inspired by this work and some recent discussions, I believe that current AlphaFold-class or ESMFold-class models are essentially composed of two main parts: a sequence encoder and a trunk + structural module. The front-end—whether based on MSA or a protein language model—borrows information from evolution. The back-end structural module fits a general physical background appropriate for common atoms and thereby supplies a great deal of additional information. Consequently, even though the information content (in bits) of the input protein sequence alone is insufficient to determine an atomic-resolution structure, the information obtained from these two components fills the gap and enables accurate structure prediction. This is a remarkably elegant design.
However, these models are characteristically top-heavy. Both MSA search time and the scale of protein language models are enormous, while the structural module remains difficult to scale—most likely because of the O(N^3)
complexity of the triangle operations.
As a result, the structural module has relatively few parameters yet consumes a disproportionately large amount of memory and compute.
In the post-AlphaFold 3 era our goal is a drug-discovery engine. We therefore want to recover the information intrinsic to the sequence itself: a sequence does not need to “know” its evolutionary ancestry in order to fold in solution, because in the actual physical process all the necessary information is supplied by this general physical background. At the same time we want to place small molecules, nucleic acids, and antibodies that are not under strong evolutionary constraint into a shared representational space for prediction and design.
I therefore suspect that future work will focus on removing the protein language model and MSA altogether and instead concentrating on fitting this physical background—scaling the latter half of the model as aggressively as possible, at least within the biomolecular setting.
@markkuilmari Työperäisen maahanmuuton hyödyt on noteerattu myös eurooppalaisessa tutkimuksessa, olen niihin viitannut. Euroopan kannalta ongelma kiteytyy siihen osaan maahanmuuttoa joka ei ole työperäistä.
EpiBench: Can LLMs Understand Epitopes for Antibody Drug Discovery?
1. The paper introduces EpiBench, a closed-book, sequence-only benchmark designed to test whether general-purpose LLMs can infer epitope information directly from antigen and antibody sequences, in a way that mirrors real antibody drug discovery decisions rather than isolated prediction tasks.
2. EpiBench is built around a connected 5-stage workflow: (T1) targetable region discovery, (T2) antibody-conditioned epitope identification, (T3) epitope binning, (T4) functional epitope assessment, and (T5) antibody escape assessment—covering discovery, characterization, and resistance risk.
3. The dataset contains 1,609 curated, automatically scorable samples grounded in experimental evidence from multiple sources: structural antibody–antigen contacts (AsEP, SAbDab), curated functional B-cell assays (IEDB), and deep mutational scanning escape measurements (DMS).
4. A key design choice is de-identification for “closed-book” evaluation: prompts remove PDB IDs, antibody/antigen names, accessions, and paper metadata, keeping only sequence-level evidence (e.g., VH/VL sequences and CDRs, numbered antigen residues, candidate epitope residue sets, and mutation descriptions).
5. The benchmark also includes shortcut-control strategies to reduce artifacts: balanced multiple-choice answer letters (Tasks 2/4), balanced binary labels (Tasks 3/5), candidate order randomization, antigen/antibody clustering to reduce redundancy, similarity-matched positive/negative pairs for binning, and mutation-level matching for escape tasks.
6. Nine general-purpose LLMs are evaluated zero-shot under a unified protocol (with explicit intermediate reasoning allowed). Results suggest models capture partial epitope-related signals but struggle with antibody-specific grounding, residue-level localization in long contexts, and biologically grounded functional/escape reasoning.
7. Task-level patterns: models show some ability to prioritize broad epitope-like regions in T1 (RegionRecall@50 improves over random), but AUROC stays near random, indicating weak residue-wise discrimination. T2 and T5 remain challenging because they require conditioning on a specific antibody to localize binding or predict mutation-specific escape.
8. Antigen length sensitivity is a major bottleneck in region discovery: average RegR@50 drops sharply as antigens get longer (reported from ~81 for <200 aa to ~13 for >800 aa), highlighting difficulties with long-context residue indexing and sparse-signal retrieval from large sequences.
9. Explicit reasoning (CoT) provides inconsistent gains: it helps some models notably on T1 (organizing open-ended residue selection) but does not reliably improve antibody-conditioned, functional, or escape decisions (T2–T5). Error-trace analysis shows frequent reliance on indirect proxies (motif fixation, generic physicochemical heuristics, similarity shortcuts, and “safe” guessing), plus concrete failures like coordinate misalignment in mutation mapping.
📜Paper: https://t.co/jnABLuN7pF
#AntibodyDiscovery #Epitopes #LLM #BioNLP #ComputationalBiology #ProteinScience #Benchmark #DrugDiscovery #AIforScience
Can frontier models reason beyond a protein's sequence?
At @InSilicoMeds, we tested Qwen3.8-Max on bovine rhodopsin.
✅ Correctly identified that E113→Q is disruptive because E113 is the counterion that stabilizes the protonated retinal Schiff base (reasoning in the comments.).
✅Correctly inferred that L112, particularly L112I, is relatively mutation-tolerant.
⁉️Missed the key structural insight: L112 faces the lipid bilayer rather than the retinal-binding pocket, providing the mechanistic explanation for its higher mutation tolerance.
Strong biochemical intuition (great job @Alibaba_Qwen), but still room to improve understanding of the 3D structural context that governs protein function.
#AI #StructuralBiology #GPCR #ProteinDesign #Bioinformatics #LLM #Qwen #DDDBench #StructBioBench
Downgrading to Opus 4.8... because Opus 5.0 is getting "too good" and beginning to act like a senior scientist in the group, overly confident, doesn't want to consult with others, thinks everything is a bad idea and not excited to try crazy new ideas. 🙃
We decided to open source the outputs of all the internal tests we do @try_litefold. From researching about cell line development, drug discovery, understanding targets, designing proteins, general research, everything is covered!
Checkout resources / use cases at https://t.co/nst2TBTadA
Spoke too soon; the Sabdab-2 annotations for shark VNARs mistakenly assume the existence of a CDR2, shown below for PDB 2I27 in a random loop at the bottom. These single domain antibodies have only CDRs 1 and 3. Sabdab-1 got this correct but is no longer accessible AFAIK
⚡ Co-folding models, a new feature now available in the #GenerativeBiologics platform.
The dedicated scoring function allows you to achieve a 70% hit rate within the top 10 predictions. See the preliminary results in this video here. 👇
#Biologics#GenerativeAI #ScoringFunction #PlatformUpdates