De novo protein design in an expanded chemical space
1 CheMoDesign is a generative framework that treats chemistry (PTMs, genetically encoded noncanonical amino acids, and post-expression chemical modifications) as a user-specified design variable, enabling de novo proteins whose binding and reactivity causally depend on the chosen chemical group.
2 Method core: it repurposes a pretrained all-atom structure-prediction model (Boltz-2) without retraining by weakening Pairformer pair-representation conditioning (zij → (1 − α)zij) to increase structural exploration, while preserving “chemical/functional anchors” that must remain fixed in the interface.
3 The workflow separates “dreaming” and “waking”: dreaming samples diverse target-associated backbones under weakened conditioning; waking performs sequence-conditioned truncated denoising to repair local geometry and reconcile the backbone with realizable sequences, improving sequence–structure compatibility across multiple targets.
4 Chemical anchors are implemented at atom-level: a parent residue representation is retained while an atom-level moiety is appended with correct covalent topology; all pairs involving the chemical moiety are protected during scaling/optimization, letting the model organize interfaces around the specified chemistry.
5 Retrospective validation: on native PTM-dependent complexes (acetyllysine, trimethyllysine, phosphotyrosine), CheMoDesign could recover native-like functional-group geometry (RMSD < 2 Å among passing designs), and “chemical ablation” back to canonical residues reduced predicted interface quality (ipTM drops of ~0.36–0.64), supporting modification dependence.
6 Prospective PTM design: sulfotyrosine (TYS) binders to thrombin were generated at multiple epitopes; 5/11 expressed designs all bound, with KD values spanning ~165–1193 nM, and Tyr substitution weakened binding across epitopes, indicating the designed sulfate group contributed to recognition.
7 Another PTM case: 3-nitrotyrosine (3-NT) binders to HER2 were designed without relying on an existing HER2–3-NT structural template; tested pairs showed 3-NT variants bound ~2.5–3.3× tighter than Tyr variants (e.g., KD ~236 nM vs ~780 nM), consistent with chemistry-guided interface formation.
8 Genetically encoded ncAAs enabled programmable behaviors beyond affinity: a tetrazine-bearing residue dominated IL-7Rα binding (KD ~8.4 nM; Ala/Phe substitutions >2 μM) and allowed post-production “rewriting” via tetrazine–TCO ligation, yielding payload-dependent affinity switching (including near-blocking by PEG-20K).
9 Reactivity as a designed output: pBpa-containing binders supported light-triggered covalent capture (rapid crosslinking upon 365 nm irradiation, including on-cell retention), while FSY-containing binders enabled proximity-driven SuFEx covalency to Trop-2 with crosslinking MS evidence for the intended linkage.
10 Post-expression chemistry broadened the accessible chemical space: installing a benzenesulfonamide modification on a designed Cys produced an isoform-selective CA IX binder/inhibitor (KD ~45 nM; weak CA XII binding; ~98% CA IX inhibition at 1 μM), and a photolabile 2-nitrobenzyl alcohol-derived modification enabled light-activated covalent PD-L1 capture with competitor-resistant cellular retention dependent on PD-L1 Lys75.
💻Code: https://t.co/t979TdZYAx
📜Paper: https://t.co/iTbjV17cvZ
#ProteinDesign #GenerativeAI #ComputationalBiology #ChemicalBiology #NoncanonicalAminoAcids #PostTranslationalModifications #CovalentInhibitors #BioorthogonalChemistry #Boltz2 #AlphaFold3
The results are finally in! 🏆💻🧬
I'm thrilled to announce that the manuscript for the Bits to Binders protein design competition is out on bioRxiv! Here's a summary of our findings, including some simple criteria that nearly *double* success rates when applied as a filter 🧵
Today, we’re announcing a collaboration with @GSK following its successful evaluation of Chai’s models.
@GSK tested designs from Chai's models in its own wet labs, generated zero-shot with no target-specific training, and found binders to all tested targets.
More ↓
RGI-Toolkit: Differentiable Restraints for Controllable Biomolecular Structure Prediction
1 RGI-Toolkit is a general-purpose Python library that brings restraint-guided inference (RGI) to diffusion-based structure predictors, enabling controllable structure generation without retraining model weights.
2 The key idea: at each reverse-diffusion step, the toolkit applies a differentiable restraint loss and locally optimizes (approximately “projects”) the denoiser’s predicted coordinates toward a restraint-satisfying set, then continues sampling with the corrected structure.
3 A major engineering contribution is portability: the toolkit centralizes restraint definition, atom selection, loss evaluation, autodiff gradients, and coordinate optimization into a shared, backend-agnostic engine—leaving each predictor to only a thin preprocessing adapter plus a single hook inserted into the denoising loop.
4 It currently supports six predictors: Boltz-2, Protenix-v2, AlphaFold3, Chai-1, OpenFold3, and ESMFold2, leveraging their similar EDM-style sampling to standardize where the RGI hook is applied.
5 Restraints are specified in a common YAML/JSON config (or Python dict for ESMFold2) with a selection DSL (chain/residue/atom + logical operators), so the same restraint recipe can be reused across different prediction backends.
6 Supported restraint types include: distance/angle/dihedral restraints over centroids of atom groups (harmonic or flat-bottom), reference RMSD restraints after Kabsch alignment (including multiple references and anchoring to external PDB/mmCIF coordinates), and custom restraints defined as user-written mathematical expressions over geometric quantities (enabling reaction-coordinate style control).
7 Ligand conformer restraints are a centerpiece: they enforce bond lengths, bond angles, chirality, cis/trans isomerism for acyclic double bonds, plus optional protein–ligand van der Waals repulsion; “ideal” conformers are generated via RDKit ETKDGv3 and optimized with UFF, with independently tunable term weights.
8 Benchmark 1 (209 protein–ligand complexes with cis/trans-distinguishable double bonds): across all six predictors, conformer restraints raised chirality and cis/trans agreement to 100% and improved local ligand geometry (lower bond/angle RMSDs), while adding vdw repulsion helped avoid steric clashes; consistent with prior work, pose accuracy was largely unchanged (geometry fixed without necessarily improving docking correctness).
9 Benchmark 2 (state control in QBP, ADK, and the MFS transporter DgoT): distance, angle, RMSD-to-reference, and custom reaction-coordinate restraints produced systematic, progressive shifts along target conformational coordinates across all models, often outperforming MSA subsampling in direct controllability (e.g., DgoT controlled via a custom ΔD = Din − Dout gate-distance difference).
10 The authors emphasize practical limitations: unreasonable or inconsistent restraints can induce strain or sampling failure; RMSD restraints can show residual deviation due to geometry constraints (e.g., triangle inequality with two references); iterative per-step optimization adds compute overhead; and ligand guidance quality depends on the correctness of the reference conformer.
💻Code: https://t.co/Mt0k3RVbgC
📜Paper: https://t.co/oK8NXPBwTs
#computationalbiology #structuralbiology #proteinfolding #diffusionmodels #AlphaFold3 #proteinligand #cheminformatics #RDKit #opensource #bioinformatics
"Simplifying in silico protein evolution with minimal screening by unZipro"
Qin Z ∙ Zhao S ∙ Deng Z ∙ Si X [..] Chen Z ∙ Song J ∙ Wang D ∙ Ji X. Mol Cell. 2026-09-24. https://t.co/CVkphDuG4b
unZipro (unsupervised Zero-shot inverse folding framework for protein evolution)
Testing fewer than ten nominated candidates (incl. AlphaFold3 model) identifies high-fitness variants
#CRISPR #BE #PE #Plant_breeding
Good afternoon, San Francisco!
Today, we're releasing Test-Time Structure-Space Search, a new OpenDDE inference algorithm for more effective co-folding when designing and selecting antibody drug candidates 🧬.
16% → 78%, that's top-1 success on antibody–antigen targets where OpenDDE usually fails, using Test-Time Structure-Space Search with just five correct contacts.
Frozen model. No retraining. No gradients through the network.
Test-time compute isn't only for LLMs. Rather than only picking the best of many samples, we search for better structures while the sampler is still denoising.
Same OpenDDE, same five contacts, same budget of 25 candidates per target:
• re-rank with the contacts: 33%
• steer the sampling with them: 78%
Re-ranking recovers only about a quarter of the gain. The rest comes from changing the sampling trajectory itself.
How it works:
• At scheduled denoising steps (about three dozen per sample), we check OpenDDE's predicted clean complex against the contacts.
• If they fail, a solver moves the antibody as a rigid body: Refine (nudge locally) → Search (broad pose search, only if still failing) → Refine (polish the best pose).
• A clash-aware energy refuses moves that create severe overlaps, and fused GPU kernels make thousands of pose evaluations per call practical.
• The corrected structure goes back to the sampler. If the contacts already hold, nothing moves.
Results:
805 held-out SAbDab targets where the base model usually fails (median DockQ < 0.23 over five samples).
Top-1 success by number of known contacts:
0 (unguided): 16% and 5: 78%.
With 5 contacts, mean top-1 DockQ goes 0.14 → 0.43 and the CAPRI-medium rate 11% → 39%.
Pocket mode (5 epitope residues, nothing on the antibody side): 17% → 33%, versus 26% for re-ranking. A smaller gain, same pattern.
Try it:
Put contact pairs or epitope residues in the "constraint" field of your input JSON and run with --use_tfg_guidance true. Checkpoints and output format are unchanged, and inputs without a constraint behave as before. With the optional trunk cache, inputs that differ only in their constraint share one trunk. Examples for Fab, Fv and VHH are in examples/tfg.
Fine print:
Constraints are simulated from reference structures, so they are correct by construction. This measures what accurate interface knowledge is worth, not robustness to experimental noise, and wrong contacts can push success below unguided. SAbDab only. Success = DockQ ≥ 0.23, 25 candidates per target, 95% bootstrap CIs. Validated on one GPU with 200 diffusion steps.
Code: https://t.co/dmvwAEZGj4
Tech Report: https://t.co/bKkjQNBFML
De novo design of monoclonal and bispecific antibodies with OFAntibody
1. OFAntibody is an all-atom generative framework that designs both monoclonal antibodies and bispecific antibodies by jointly modeling multiple antigen–antibody interfaces inside a single shared antibody architecture, rather than designing two binders independently and assembling them afterward.
2. The key bispecific innovation is arm-aware multi-hotspot routing: users can specify separate hotspot sets for different epitopes/targets, and the model explicitly couples each hotspot set to its intended binding arm, reducing ambiguity in multi-interface generation.
3. OFAntibody adds multi-component structural supervision by training on experimentally resolved ternary (“triplex”) complexes, enabling the model to learn coupled geometric constraints across multiple interacting interfaces that must coexist in bispecific formats.
4. It also scales antibody–antigen interaction learning via structural distillation: ~150,000 antibody–antigen pairs were generated from ASD entries using AF2-Multimer filtering (ipTM ≥ 0.7), expanding interface diversity and improving epitope-conditioned design.
5. The framework supports epitope-conditioned generation across monoclonal settings (e.g., VHH nanobodies, scFv-like contexts) and multiple bispecific formats, including tandem VHH, diabody, and CODV-like architectures, with format conditioning treated as part of the generation problem.
6. In a competitive nanobody benchmark (13 tasks across 12 targets, candidates from multiple methods ranked together), OFAntibody achieved Top-5 enrichment of 41.5%, a 5.39× improvement over RFantibody for competitive candidate ranking; it also improved overall prioritization (median rank 250 vs 300 for RFantibody).
7. Metric-level analysis indicated OFAntibody designs tended to have higher predicted interface confidence (higher iPTM, lower interface PAE) and more extensive polar and buried interfaces (more H-bonds/salt bridges and higher ΔSASA), while emphasizing that local contact metrics alone can be misleading without global geometric validity checks.
8. Distillation data materially contributed: compared to training without distillation, the full model improved Top-5/Top-10/Top-20 enrichment by +9.2/+6.9/+7.3 percentage points, suggesting broader learned interface geometry helps generalize across targets and epitopes.
9. In bispecific design tasks, OFAntibody achieved hotspot pass rates of 94–100% across evaluated cases; energy pass rates were 60% (diabody EGFR×CD3), 13% (tandem VHH EGFR 7D12×EgA1), and 8% (CODV-like CD28×CD3). The paper notes that CODV-like geometries may violate “canonical” CDR-interaction heuristics, motivating format-aware evaluation criteria.
10. Joint multi-interface generation was compared to independent binder design + assembly for a tandem VHH case: independent assembly produced severe VHH–VHH steric clashes in 80.2% of combinations and reduced epitope/hotspot compatibility, while joint generation avoided major spatial conflicts and improved CDR contact fraction and hotspot coverage (reported 98.6% hotspot coverage in the joint setting).
💻Code: https://t.co/ut2Gd8G8tz
📜Paper: https://t.co/0MJbvG6F3w
#AntibodyDesign #BispecificAntibodies #ProteinDesign #GenerativeAI #ComputationalBiology #StructuralBiology #Nanobody #AI4Science #DiffusionModels #Bioinformatics
1/5 New preprint!
1/5 🟢🔴 New preprint!
Meet eLACCO3 & R-eLACCO3: fast, Ca2+-independent green/red biosensors for extracellular lactate imaging. From protein engineering to multiplexed imaging in vivo.
https://t.co/D1ZqFtvdZb