5/5
We hope RNA-Scope becomes a useful resource for the RNA and AI4Science communities.
📄 Read the paper
🔁 Share it with biological foundation model researchers
💬 Feedback and collaboration are welcome!
#RNA#FoundationModels#ComputationalBiology
🧬 Excited to share RNA-Scope—accepted as a poster at the NeurIPS 2025 AI4Science Workshop!
It introduces a systematic benchmark for evaluating how well RNA language models understand structure, interactions, and function.
Paper: https://t.co/ZIC29w5y1x
4/5
Why does this matter?
Narrow benchmarks can provide an incomplete or biased picture of model capabilities. RNA-Scope enables broader, more biologically relevant, and reproducible comparisons—showing where current models succeed and where new methods are still needed.
3/5
The results reveal an important gap: existing RNA language models capture useful biological signals, but generalization remains challenging across RNA families, molecular targets, and environmental contexts. Complex structural dynamics and design tasks remain difficult.
4/4
This is a living, community-driven resource. If we missed a model or any metadata needs correction, issues and pull requests are welcome.
⭐ Star the repository
🔁 Share it with the genomics and RNA communities
🤝 Contributions welcome!
🧬 Excited to share Awesome Nucleotide Foundation Models—a curated, continuously maintained catalog of foundation models for DNA, RNA, and unified DNA+RNA sequence modeling.
Explore the landscape:
https://t.co/Zui0ThoCTB
#BioAI#Genomics#RNA
3/4
Beyond model lists, the repository includes a visual timeline, a conceptual model taxonomy, detailed notes, and categories spanning general, specialized, generative, and multimodal systems.
A practical starting point for exploring and comparing nucleotide foundation models.
In new work, we lay out a vision for a high-level programming language for generative biology, called Proto.
Proto composes generative and predictive models spanning DNA, RNA, proteins, ligands, and their interactions, which we use to design complex biological functions. 1/n
Limits of Deep-learning-based RNA Prediction Methods
1. The study benchmarks 10 deep-learning RNA structure predictors on an independent, non-redundant dataset designed to test real generalization: 79 monomeric RNAs and 158 RNA-containing complexes, filtered by sequence identity (CD-HIT-EST 80%) and structural redundancy (RNA-align TM-scoreRNA; removing pairs with TM-score > 0.7).
2. The central finding: current RNA predictors perform well mainly when the target fold resembles structures already present in training data, and performance drops sharply for structurally divergent RNAs—suggesting the models often recognize recurring motifs rather than learn broadly transferable RNA-folding principles.
3. For monomeric RNA, AlphaFold3 (AF3) and Boltz-1 are consistently among the top methods across metrics, but the “best” method depends on which accuracy measure is used and whether targets are “easy” (training-set-similar) or “hard” (training-set-dissimilar).
4. Easy monomeric targets (max similarity to training set TM-score ≥ 0.5): NuFold has the highest mean TM-score (0.482) and GDT-TS (0.610), while AF3 leads on local/contact-oriented measures with the best INF (0.777) and lDDT (0.693). Several top methods are statistically indistinguishable on TM-score/GDT-TS, but AF3 is notably strong on lDDT/INF.
5. Hard monomeric targets (max similarity TM-score < 0.5): absolute performance is low across the board. Boltz-1 has the highest mean TM-score (0.247), GDT-TS (0.437), and INF (0.673), while AF3 has the best lDDT (0.553). HF3 is statistically indistinguishable from the top methods on these hard targets, emphasizing how small the gaps become when generalization is truly tested.
6. RNA type matters: well-defined, recurring folds (e.g., L-shaped tRNA-like RNAs and simpler loop/helical motifs) are predicted more accurately by most methods, while structurally challenging categories—especially small G-quadruplex RNAs—are frequently among the worst predictions, even when the model “recognizes” the fold.
7. The paper highlights a practical evaluation pitfall: single metrics can overstate success. When aggregating all predictions across all methods, only 12 models pass success thresholds simultaneously for TM-score, GDT-TS, INF, and lDDT; many more pass only one metric, reinforcing the need for multi-metric assessment of RNA 3D predictions.
8. MSA depth still matters for RNA: using a uniform MSA pipeline across methods, the authors observe a positive correlation between MSA depth and TM-score for MSA-using methods, with the strongest improvements for AF3 and NuFold—though gains are most apparent at very high coverage and remain moderate per-target.
9. For RNA complexes (AF3, Boltz-1, HF3, RF2NA), AF3 and Boltz-1 clearly outperform HF3 and RF2NA on both global topology (TM-score) and interface quality (DockQ). AF3 vs Boltz-1 differences are generally not statistically significant on most interface measures, indicating similar top-tier performance in complex prediction under this benchmark.
10. A key limitation for complexes: high global TM-score does not guarantee correct binding. Many cases show accurate individual chain folds but mislocalized RNA relative to its partner, yielding low DockQ/IF-DockQ despite decent TM-score. Binding-site F1 analysis further separates “right site but wrong pose” from “wrong site,” and shows RNA–protein interfaces remain a pronounced weak point.
💻Code: https://t.co/bommQGiHRL
📜Paper: https://t.co/fIoJYVglUW
#RNA #StructurePrediction #DeepLearning #AlphaFold3 #Benchmarking #ComputationalBiology #Bioinformatics #RNAProteinInteractions #MSA #PDB
Excited to share that Spatial Hi-C-RNA is now out in @CellCellPress! It brings the 3D genome into spatial multi-omics by co-mapping chromatin architecture and gene expression in the same tissue section. Congrats to the whole team! https://t.co/ov9hQkKNOA
Breakthroughs in AI for biology are largely catalyzed by the diverse datasets deposited by scientists.
To make experimental data on biomolecular dynamics from NMR as easily accessible as the PDB, @HWaymentSteele and I are releasing makeshift, installable via PyPi
Preprint: https://t.co/OhxZV8OFrt
Docs: https://t.co/4htuCxpPGQ
Today, we are incredibly excited to present LiteMol-1, our very first foundation model from LiteFold.
Structure-based models like BoltzGen, O-Design, and RFdiffusion have become the de facto standard for designing biomolecules. However, they are expensive to run at scale. In many campaigns, we have to generate tens of thousands of designs and rigorously filter them down to a handful of top candidates.
More importantly, most molecule design systems are primarily optimized around binding. But what about everything else that makes a molecule a useful therapeutic: ADME, toxicity, drug-likeness, membrane permeability, synthesizability, selectivity, and more?
That’s where we introduce LiteMol-1, our first Multi-Molecule Foundation Diffusion Language Model, pre-trained from scratch. With a single set of weights, LiteMol-1 can conditionally generate small molecules, peptides, cyclic peptides, depsipeptides, peptides with ncAAs, macrocycles, and PROTACs.
You can generate molecules unconditionally or condition generation on a protein target sequence. You can in-paint molecules, preserve parts of an existing scaffold, generate non-canonical peptides, cyclize designs, and continuously edit and regenerate them.
We also introduce a Monte Carlo Tree Search framework for multi-objective molecular design. Instead of trying to make one molecule perfect at everything, the search keeps a set of promising molecules, each with different trade-offs across properties. This matters because molecular design is fundamentally not a single-objective optimization problem.
Finally, and probably the part we are most excited about: this is a model for Agents. LLMs are not particularly efficient interfaces for repeatedly reasoning over thousands of large PDB/CIF files inside an AutoResearch loop. Sequence space is different. SMILES and molecular sequences are compact, editable, and much easier for an agent to inspect, compare, modify, and reason over.
So we treat LiteMol-1 as an infinite molecular canvas.
The model generates possibilities. The agent takes inspiration from them, evaluates them, edits them, optimizes them, and generates again. The AutoResearch loop continues until it finds candidates that satisfy the given design objectives; while bringing in expensive structure prediction, docking, or simulation only when they are actually needed.
Our evaluations across peptide and small-molecule generation show that LiteMol-1 is competitive with, and in several settings on-par with or better than, frontier structure-based and sequence-based models, while operating at a fraction of the generation cost and time.
And this is only the first step. Check out our technical research blog post in the comments to learn more about LiteMol-1.
Morpheus-3D: Structural Diversity-Guided Detection and Localization of Protein Fold Switching
1 Morpheus-3D is a sequence-based framework that detects fold-switching (metamorphic) proteins and pinpoints the specific sequence regions responsible for conformational transitions, addressing a key gap in prior predictors that largely stop at protein-level classification.
2 The central idea is to quantify residue-level tertiary structural diversity using entropy over Foldseek’s 3Di structural alphabet (20 states capturing nearest-neighbor tertiary interaction geometry), enabling detection of fold switching even when secondary structure composition is largely preserved.
3 Pipeline overview: split a query sequence into overlapping 7-mers; retrieve identical 7-mer matches from a large structure-informed database; extract the 3Di state and DSSP secondary structure for each hit; compute Shannon entropy of 3Di-state distributions (H3Di) and secondary-structure ambiguity metrics; summarize with rolling windows to produce features for classification and tracks for localization.
4 Database scale and multimodality: ~200k redundant PDB structures plus ~1.2M AlphaFold2 models are encoded into 3Di sequences, paired with DSSP secondary structure; AlphaFold pLDDT is used to filter low-confidence fragments (pLDDT < 70), and redundancy is controlled via clustering to reduce bias from overrepresented sequences.
5 Training/evaluation: 94 experimentally verified metamorphic proteins vs 94 monomorphic controls (balanced). Models compared included Logistic Regression, SVM, Random Forest, Gradient Boosting; nested cross-validation selected an RBF-kernel SVM as the final classifier.
6 Performance and what drives it: Morpheus-3D improves over secondary-structure-based baselines (Chen et al. and Morpheus-1), with the largest gains attributed to tertiary-interaction diversity. Permutation importance highlights mean 3Di entropy as the dominant feature, consistent with fold switching that rewires tertiary contacts without major secondary-structure changes.
7 Residue-level localization is demonstrated on classic metamorphic proteins (e.g., RfaH, Lysenin, KaiB, PimA): elevated H3Di and secondary-structure entropy cluster in experimentally validated switching segments, and mapping H3Di onto 3D structures shows spatial clustering of high-entropy residues within the switching region.
8 Generalization beyond training: Morpheus-3D correctly classifies and localizes switching regions in recently characterized natural fold-switchers and engineered systems absent from training, including PopP2 (InsP6-induced helix-to-strand transition), an ancestral DPBB/DZBB fold-switcher, viral nsP4, and designed two-state hinge/biosensor proteins where secondary structure changes are minimal but tertiary rearrangements are substantial.
9 Proteome-wide screening across 57 representative proteomes (plus a viral proteome) suggests fold-switching potential is widespread but uneven (reported range after disorder filtering: ~1.2% to 26.5%). Enrichment is observed in regulatory/pathogenic/adaptive contexts (notably bacterial pathogens and viruses), while vertebrates show higher predicted fractions than invertebrates, with functional enrichment in signaling and immune regulation.
10 Evolutionary analysis integrates ancestral sequence reconstruction (ASR) with entropy profiling for XCL1/CCL20, KaiB, and RfaH/NusG. Results support lineage-specific emergence of conformational plasticity (e.g., increasing entropy along XCL1 but not CCL20), localization of substitutions to switching regions (KaiB), and polyphyletic acquisition patterns in RfaH/NusG.
💻Code: https://t.co/LtXj6kMMnN
📜Paper: https://t.co/vhAStLqsVJ
#computationalbiology #bioinformatics #proteinstructure #metamorphicproteins #foldswitching #proteomics #evolution #machinelearning #structuralbiology #Foldseek #AlphaFold
Task- and dataset-specific information in protein language models
1 They systematically test a common PLM habit: using last-layer embeddings by default. Across 13 protein language models and 15 downstream tasks, the deepest layer is best only ~18% of the time; performance often peaks in mid layers (roughly 10–90% of depth) and then drops near the end.
2 The study runs layer-wise “probing” (linear probes, plus kNN probes) on embeddings from every layer, spanning 11 datasets and both protein-level and residue-level tasks. This turns “which PLM should I use?” into “which layer should I use for my dataset/task?”
3 A key pattern: residue-level tasks (e.g., secondary structure prediction, binding-residue classification) usually improve monotonically with depth, consistent with pretraining objectives that are residue/token-focused (MLM or next-token prediction).
4 Protein-level tasks behave differently: for many datasets, performance rises quickly in early layers, peaks in intermediate layers, and declines in the deepest layers. The authors connect this to “piecewise training”: PLMs are pretrained for token prediction, not end-to-end for the downstream objective, so later layers may become specialized for the pretraining head rather than general downstream separability.
5 They support this with latent-space geometry across layers: intrinsic dimension (TwoNN), neighborhood overlap between adjacent layers, and PCA variance@10. Shallow and deepest layers look more similar to each other than to mid layers, consistent with the idea that MLM-trained models “return” toward token-prediction-friendly representations late in the network.
6 Fine-tuning changes the story: when ESM-2 150M is fine-tuned on downstream tasks, layer-wise performance becomes much more monotonic (each deeper layer tends to improve on the fine-tuning objective). But this comes with reduced performance on the original pretraining objective (MLM loss gets worse), illustrating a trade-off and objective shift.
7 Dataset structure strongly predicts where useful information sits. Deep mutational scanning (DMS) datasets centered on a single protein (e.g., GFP fluorescence, GB1) tend to prefer shallow-layer embeddings; diverse multi-protein datasets (e.g., DeepLoc2.0 localization, DeepSol solubility, SCOPe40 fold/superfamily) tend to benefit from deeper layers (until the final drop).
8 They show that within the same dataset, changing the downstream objective (classification vs regression; binary vs multiclass; 3-class vs 8-class SSP) barely changes the layer-performance “shape”. This suggests the dataset’s composition can dominate over task formulation when deciding which layer is optimal.
9 Practical takeaway: identifying a near-optimal layer does not require full training data. Using only ~15–20% of downstream training data was often enough to find a layer reaching ≥95% of best-layer performance; larger PLMs were more robust in this sparse setting.
10 A caution for protein design: PLMs perform notably worse on artificial proteins in tested cases (e.g., designed proteins in Rocklin stability; ProGen-generated lysozymes with measured activity). The results suggest PLMs may encode “natural evolutionary sequence space” more than generalizable whole-protein function for novel designed sequences, even when those sequences are experimentally functional.
📜Paper: https://t.co/sGtfAHkGN0
#ProteinLanguageModels #PLM #ESM2 #ProtT5 #ProGen2 #RepresentationLearning #Probing #Bioinformatics #ComputationalBiology #ProteinDesign
Distribution-Constrained Optimization for Reliable ML-Guided 5′UTR Sequence Design
1. The paper studies a common failure mode in ML-guided 5′UTR optimization: genetic algorithms can “reward hack” a translation-efficiency predictor by drifting out of the model’s validated training distribution, producing sequences with high predicted mean ribosome load (MRL) but elevated risk of misprediction.
2. Their core proposal is a trust-region style constrained optimization: keep candidate 5′UTRs inside the predictor’s training distribution (where accuracy has been validated). “Reliable” is defined as in-distribution with respect to training data, not as a guarantee of wet-lab performance.
3. To define out-of-distribution (OOD) for nucleotide sequences, they compare two scores computed from UTR-Insight’s internals: pseudo-perplexity (PPPL) vs k-nearest-neighbor (KNN) distance in the model’s embedding space. PPPL fails to separate in- vs out-of-distribution in this setting (the 4-letter alphabet makes PPPL broadly permissive), while embedding KNN distance provides a usable OOD signal.
4. They show OOD drift is not hypothetical: under unconstrained GA optimization, 72–96% of final-generation candidates exceed the self-KNN p95 threshold (across seeds), and the median KNN distance increases ~4.7-fold from early generations—consistent with optimization pushing into extrapolative regions.
5. Using KNN distance as a hard GA constraint (feasibility boundary set by percentiles of the training data’s self-KNN distribution) keeps all candidates within the trust region while maintaining predicted MRL at essentially the unconstrained level (median predicted MRL ~9.8–10.1 across seeds under KNN-Pr).
6. A practical payoff: compared with post-hoc filtering of unconstrained outputs under the same compute budget, the hard KNN constraint yields ~4.3× more “selectable” low-risk candidates (i.e., high predicted MRL without leaving the trust region).
7. They examine which reference distribution to use for KNN: (Pr) unlabeled native 5′UTRs from pretraining corpora vs (Sv) the supervised MPRA library used for predictor training. Distances correlate strongly (Pearson ~0.94), but the native-reference trust region is effectively stricter: KNN-Pr constraint tends to satisfy both Pr and Sv trust regions, while KNN-Sv may not satisfy Pr consistently across seeds.
8. They compare other search-space controls: an output extrapolation guard (constraining predicted MRL to be within the supervised training MRL range, e.g., p95=8.28) strongly suppresses OOD as a side effect, but it also caps predicted MRL and concentrates solutions near the ceiling—trading performance headroom for conservatism.
9. They test additional objectives/constraints often used in 5′UTR design: adding RNA secondary-structure accessibility near the start codon (RNAplfold) as a secondary objective broadens exploration and creates a Pareto trade-off with predicted MRL, but does not suppress OOD risk; reference-sequence similarity suppresses OOD when used as a constraint (at the cost of lower MRL), while using similarity as an objective enables fine-grained “distance-from-reference” control but can still include many OOD candidates.
📜Paper: https://t.co/Drj63MuszW
#ComputationalBiology #SyntheticBiology #mRNA #UTR #MachineLearning #OutOfDistribution #GeneticAlgorithms #SequenceDesign #TrustRegion #Bioinformatics
still using SAEs to explain llms? way too slow! check out these 4 much faster methods for llm interpretability:
1. ICA Lens: Interpreting Language Models Without Training Another Dictionary (ours). ICA is an underrated efficient approach. non-gaussianity is a signal worth paying attention to https://t.co/4WTixvUY7T
2. svd as a fast interpretability method for transformers
https://t.co/uw3kjvjFst
3. sparse weight decomposition for efficient circuit extraction https://t.co/Jw9QirVN7b
4. language model circuits are sparse in the neuron basis https://t.co/1mYvxQEnI2
overall it seems building a new dictionary is actually pretty easy. saes are slow and mediocre; a lot of these newer dictionaries are almost free with very low generation cost. we should prioritize them.
labeling the items inside the dictionary is still the real pain though. most people just use llms for auto-annotation and the results are only so-so, that’s where the next real breakthrough needs to happen.
CENO: A Genome-Scale World Model for Evolutionary Sequence Interpretation and Programmable Regulatory Design
1 CENO is presented as a “genomic world model”: one autoregressive generative system that keeps nucleotide-resolution state over very long contexts, can score counterfactual mutations via likelihood deltas, can condition on homologous evidence (MSA), and can generate sequences for downstream structural/functional objectives.
2 The core architecture is a hybrid causal backbone that interleaves Mamba-2 sequence-mixing blocks, sparse attention blocks, and mixture-of-experts capacity. Models are trained at 300M/600M/1B total parameters (about 100M/200M/400M active parameters per token).
3 A staged context curriculum separates model scale, context length, and data composition: Stage I/II at 8k tokens on broad cross-domain genomes; Stage III continues at 131k tokens; Stage IV continues at 1M tokens using long windows from complete eukaryotic genomes. This is used to test whether whole-genome continuation improves gene-scale reconstruction and long-range reasoning (not just accepting longer inputs).
4 Practical long-context inference is a key claim: in long-prefill and sustained decoding benchmarks, CENO maintains high tokens/s/GPU at 131k context and avoids OOM failures seen in some large long-context baselines under matched setups.
5 Long-range “use of context” is probed directly. In a synthetic DNA needle-in-a-haystack retrieval assay, distant inserted sequences measurably shift terminal base predictions across 32k–1M contexts, with retrieval generally improving with model size and later long-context training.
6 Without any task-specific fine-tuning, long-context continuation produces chromatin-scale signals: attention maps show higher within-TAD vs across-boundary attention at 131k vs 8k, and this pattern generalizes across 5 human cell types and 5 mouse cell/tissue settings (strongest for the 600M and 1B models).
7 Frozen-state representations also encode boundary information: linear probes trained on hidden states (with the backbone frozen) distinguish TAD boundary bins from matched random bins, with AUROC rising markedly when using 131k vs 8k contexts in both human and mouse settings.
8 For interpretation, CENO is evaluated with a unified zero-shot variant-effect paradigm: score variants by reference–alternate log-likelihood deltas across a broad benchmark spanning splicing, expression/eQTL, enhancer–gene links, structural variants, ClinVar pathogenicity, BRCA, Mendelian/complex traits, and fitness assays. Aggregate rankings place long-context CENO-1B checkpoints near the top overall, while task-level leadership varies by dataset and context.
9 Long context yields the clearest functional gains where distance matters: on TraitGym, moving from 8k to 131k matched-input evaluation improves AUROC more strongly for distal variant–gene pairs (e.g., beyond 100 kb), consistent with long-range regulatory dependency.
10 For generation, CENO is tested on (a) partial-gene continuation across eukaryotic, bacterial, and archaeal species, where recovery improves with scale and later whole-genome long-context training, and (b) structure-guided long insertion design at boundary-deletion and neo-domain loci. Candidates are selected with one structure oracle (Akita) and evaluated with a held-out model (AlphaGenome), showing cross-model transfer for boundary restoration at specific loci and partial attenuation of ectopic contacts in a duplication setting.
11 Evolutionary conditioning is added via post-training on packed real MSA contexts (CENO-P): row-local layers preserve within-sequence modeling while fusion layers allow homologous rows to condition target representations. Variant scores remain simple likelihood deltas under the same MSA context, and post-training substantially improves BRCA1/2 AUPRC versus the matched pretrained backbones.
12 On cross-species conservation/enrichment benchmarks, CENO-P ranks first across multiple species panels with strong odds ratios, suggesting the MSA-conditioned likelihood interface captures evolutionary constraint signals broadly across clades when homologs are available.
13 For programmable regulatory design, CENO is used as the backbone of a mouse cortex cell-type-specific enhancer workflow: a CENO-based accessibility oracle is trained on scATAC-seq peak regions, interpretable motif grammar is extracted via in silico mutagenesis, and a conditional generator is optimized via supervised fine-tuning plus GRPO-based reinforcement learning to increase predicted target-cell accessibility while penalizing off-target activity.
📜Paper: https://t.co/URqrsYNuTW
#ComputationalBiology #Genomics #DNAFoundationModels #LongContext #GenerativeModels #VariantEffectPrediction #MSA #Evolution #RegulatoryGenomics #EnhancerDesign