Beyond Structure and Affinity: Context-Dependent Signals for de novo Binder Success
1. The study argues that de novo binder evaluation is still overly centered on structure confidence and affinity proxies, yet large public benchmarks show most designs fail experimentally and “good” in silico scores often do not predict in vivo success.
2. It re-analyzes two public datasets with very different deployment contexts: 11,984 CD20 binders used as CAR extracellular domains in primary human T cells (multi-gate: expression/recovery, enrichment, depletion) and 603 EGFR binders tested as standalone proteins (cell-free expression + BLI binding).
3. Instead of structural scores, it uses “biology-informed” sequence descriptors from ML models trained on natural proteins (Orbion Astra suite) to quantify disorder, amyloid/aggregation propensity, topology-like character, PTM-site patterns, and broad protein/functional label probabilities—treating outputs as compatibility signals rather than literal annotations.
4. A key contribution is multi-gate analysis: the same feature can help at one experimental stage and hurt at another. In CAR-T, several descriptors flip direction between the expression gate and the enrichment gate, implying that single-objective ranking can select candidates that pass one stage but fail later.
5. The most transferable cross-benchmark signal is lower aggregation/amyloid propensity: sequences with lower predicted amyloidogenicity are more likely to succeed in both CAR-T enrichment and EGFR binding, suggesting aggregation risk is a broadly useful pre-synthesis filter.
6. PTM-site density emerges as a recurring univariate correlate of success in both benchmarks (higher predicted PTM-site counts associate with enrichment/binding). However, in EGFR it is partly length-confounded due to variable sequence lengths (13–250 aa), so it is more robust in the fixed-length CAR-T setting.
7. Several signals are architecture-dependent (significant in both datasets but reversing direction), consistent with different requirements for membrane-displayed CAR domains versus standalone binders: topology-like character, disorder (especially C-terminal), and disulfide-related sequence character can indicate success in one context and failure in the other.
8. Context-specific signals also appear. In CAR-T, phosphorylation-site-related descriptors show a strong association with depletion (a potential failure mode signal), while in EGFR the dominant success signal is low disorder (large effect), consistent with the need for compact, independently folding binders.
9. Practical takeaway: stacked biology-informed filters can enrich hits. In CAR-T (after controlling for known simple predictors like cysteine and K+E fraction), adding filters for low amyloidogenicity, outside-topology-like character, and PTM sites ≥10 increases enrichment hit rate from 13.8% to 38.6% (2.8× lift) in a retrospective analysis, motivating context-aware pre-screening to reduce wasted synthesis/testing.
📜Paper: https://t.co/OaK4YGfaS4
#ProteinDesign #ComputationalBiology #Bioinformatics #MachineLearning #CAR-T #ProteinEngineering #DeNovoDesign #ProteinBinders #Benchmarking #Biophysics
AstraPTM2: A Context-Aware Transformer for Broad-Spectrum PTM Prediction
1. AstraPTM2 is a novel computational tool that predicts 39 distinct types of post-translational modifications (PTMs) on full-length protein sequences. This broad coverage is unprecedented and addresses a significant gap in the field, as most existing tools focus on a limited number of PTMs or truncate protein sequences, losing critical context.
2. The model integrates ESM-2 embeddings, AlphaFold2-derived structural features, and protein-level descriptors to capture both short-range motifs and long-range dependencies. This multi-scale approach allows AstraPTM2 to identify PTMs that are driven by complex structural and sequence contexts, enhancing its predictive power.
3. AstraPTM2 employs a three-stage curriculum learning strategy to balance the prediction of rare and common PTMs. By progressively introducing rarer labels during training, the model achieves a macro-F1 score of 59% and an AUROC of 0.99, demonstrating strong performance across a diverse set of PTMs, including those with limited training data.
4. The model’s outputs are calibrated using per-label affine calibration and optimized thresholds, ensuring that predictions are reliable and interpretable. This calibration process is crucial for generating actionable insights, especially for experimental biologists planning mutagenesis studies or designing protein constructs.
5. AstraPTM2 is deployed on the Orbion web platform, offering synchronized 2D and 3D visualizations of PTM predictions. The platform supports dual prediction modes—calibrated and exploratory—allowing users to distinguish high-confidence sites from lower-confidence hypotheses, facilitating both decision-making and hypothesis generation.
6. In hold-out tests, AstraPTM2 shows particularly strong performance on rare motif-driven PTMs such as O-linked glycosylation and sumoylation. This highlights the model’s ability to identify PTMs that are often overlooked by other tools, potentially revealing new biological insights and therapeutic targets.
7. The development of AstraPTM2 underscores the importance of balanced data curation, thoughtful curriculum scheduling, and motif-sensitive architectures. These principles are likely to guide the next generation of multi-PTM predictors, emphasizing that thoughtful design can be more impactful than simply increasing model size or dataset volume.
📜Paper: https://t.co/7eCzlocJ3T
#AstraPTM2 #PTMPrediction #ComputationalBiology #TransformerModel #ProteinStructure #MachineLearning #Bioinformatics
Would an organic Neuralink be more acceptable than the inorganic one? Is it the lack of acceptance happening because of the lack of control or something else?
This should have been an easy solution. It may not have been as sensational as these launches, but it could have been an amazing product and a good complementation to the smartphones. Probably an intermediate step in the evolution of smartphones. 🦎
I always imagined the #AI devices to be something like the book "Hitchhiker's Guide to the Galaxy" from the book "Hitchhiker's Guide to the Galaxy". IMO #humane came up with good ambition, but based on its potential, #rabbitr1 is closer to being one that'll serve this purpose.
- LLMs could have been run on the local device, but also provide a faster cloud alternative.
- Access everything on the smartphone so that every interaction can be fed to the LLM for personalization.
- The devices could have been much lighter and sold at a cheaper price.
I don't want to migrate to a new MacBook in the next six months. Making these switches a lot just gets annoying. Especially when trying to set up an environment for Unity.