Cool eval. Simply ask an LLM “Land or Water?” and give it a latitude and longitude coordinate as text. Ask 16,200 times, plot as image. The models know. From compressing the internet.
Genomes are full of dark matter of unknown functions.
We present Minerva, a new method for discovery guided by genome language models. Minerva reveals the interactions hidden in non-coding DNA, pointing to hundreds of new putative RNAs and repetitive elements per bacterial genome. Minerva allows us to find and study elements invisible to traditional methods at orders of magnitude greater scale than before.
Our group discovered that reasoning models produce fractals when asked to solve hard problems. We can use nonlinear dynamics to probe the thinking processes of recurrent depth models on Sudoku, mathematics, and even ARC-AGI (1/N)
https://t.co/Q3u8OqylZf
Accelerating Molecular Dynamics Simulations with Foundation Neural Network Models using Multiple Time-Step and Distillation
1. A novel approach to accelerate molecular dynamics (MD) simulations has been introduced by Cattin et al., leveraging foundation neural network models and a dual-level neural network multi-time-step (MTS) strategy. This method pairs a highly accurate potential with a faster, distilled model, achieving significant speedups while maintaining accuracy.
2. The core innovation lies in the application of a RESPA-like MTS scheme tailored for neural network potentials. By integrating fast-changing forces with a small time step and slower forces with a larger time step, the method reduces the computational cost without compromising the fidelity of the simulations.
3. The study demonstrates substantial performance gains, with up to 4-fold speedups in homogeneous systems like bulk water and 2.3-fold speedups in large solvated proteins. This is achieved by evaluating the computationally expensive model only every 3 to 6 femtoseconds, depending on the system.
4. Two complementary distillation strategies are presented: an on-the-fly system-specific model and a generic model. The on-the-fly model is trained for each specific system, offering high accuracy, while the generic model provides broader applicability and faster deployment across different systems.
5. The robustness and accuracy of the MTS scheme are validated through various numerical experiments, including stability tests in bulk water, hydration free energy calculations for small molecules, and simulations of protein-ligand complexes. The results show that the MTS scheme preserves key physical observables with minimal loss of accuracy.
6. The implementation of this MTS scheme is integrated into the FeNNol library, a neural network framework for molecular simulations. This integration allows for seamless coupling with the scalable Deep-HP machine learning interface and the GPU-accelerated Tinker-HP molecular dynamics package.
7. Future work will focus on expanding the applicability of this approach through randomized time stepping and fragment-based active learning strategies to further improve stability and efficiency for more complex chemical systems.
📜Paper: https://t.co/QK01HOxSs4
#MolecularDynamics #NeuralNetworks #SimulationAcceleration #ComputationalBiology #MachineLearning #MTS #FoundationModels
Skala is now available to everyone!
Why are we releasing it? Because we’re not just aiming to publish a cool paper — we’re on a mission to bring DFT to chemical accuracy using deep learning. And to make real progress, we need the community’s feedback. #compchem
Generation of protein dynamics by machine learning
1. Machine learning, particularly generative models, is revolutionizing the prediction of protein dynamics by enabling the generation of structural ensembles beyond traditional simulations. This review highlights emerging approaches that capture protein dynamics in various forms, including PDB-like ensembles and acceleration of molecular simulations.
2. One significant innovation is the development of deep generative models (GMs) based on AlphaFold2, such as AlphaFlow and UFConf, which can generate multiple conformations from a single protein sequence. These models outperform traditional sampling methods in capturing PDB-like conformations.
3. The review emphasizes the importance of hybrid models that integrate experimental and simulation data. BioEmu, a diffusion model, demonstrates unprecedented performance in modeling both PDB and MD ensembles by leveraging a hybrid training strategy. This approach captures large and biologically significant conformational changes.
4. For non-globular proteins, especially intrinsically disordered regions (IDRs), ML methods are crucial for generating ensembles. Models like IDPFold and BioEmu show promise in capturing experimental observables of IDRs, such as chemical shifts and radius-of-gyration, using a combination of PDB structures and simulations.
5. The integration of experimental data directly into the generative process is another key advancement. Methods like DynamICE and DEERFold incorporate NMR and other experimental data during training, enhancing the accuracy of generated ensembles. This approach is essential for guiding ML models towards biologically relevant conformations.
6. Despite these advancements, challenges remain, including the transferability of models beyond training data and the generation of states with correct relative probabilities. The scarcity of long MD simulation datasets and the need for larger, more diverse training sets are also highlighted as critical areas for future work.
📜Paper: https://t.co/RlcASWEmx6
#MachineLearning #ProteinDynamics #StructuralBiology #GenerativeModels #Bioinformatics
DynaRepo: The repository of macromolecular conformational dynamics
1. DynaRepo is a novel repository that addresses the critical gap in studying the dynamic behavior of macromolecules, which is essential for understanding interactions like antibody–antigen recognition and protein–nucleic acid binding. This repository provides a comprehensive dataset of macromolecular conformational dynamics, offering over 1100 µs of molecular dynamics data from simulations of around 450 complexes and 270 single-chain proteins.
2. The repository includes a diverse range of macromolecular complexes and single-chain proteins sourced from PDBbind, the Structural Antibody Database (SAbDab), and benchmark sets. Each complex was simulated in triplicate for 500 ns, ensuring robustness and reliability of the data. The extensive precalculated analyses provided with each entry make it a valuable resource for researchers.
3. DynaRepo is designed to support dynamics-aware deep learning frameworks, which is a significant step forward in the field of computational biology. Unlike most existing repositories that focus on isolated proteins or small molecules, DynaRepo emphasizes macromolecular complexes, including protein–nucleic acid assemblies. This focus is crucial for capturing the dynamic heterogeneity essential for biological function.
4. The data selection process for DynaRepo is meticulous, involving multiple filtering steps to ensure high-quality and representative structures. For example, protein–protein complexes from PDBbind were filtered based on resolution, structural gaps, and clustering to select 405 representative structures for simulation. The transient benchmark set and antigen set were also carefully curated to include diverse and relevant entries.
5. The molecular dynamics simulations were performed using state-of-the-art protocols and force fields. For proteins, the GROMACS 2024.2 software with the CHARMM36m force field was used, ensuring physiological conditions and accurate modeling. For protein-nucleic acid complexes, simulations were performed using Amber20 or NAMD3 with appropriate force field parameters. The detailed simulation protocols ensure that the data generated are of high quality and relevant for various applications.
6. DynaRepo follows the standard workflow of MDDB for data analysis, metadata preparation, and visualization, integrating a suite of biomolecular analysis tools into a reproducible pipeline. The database is FAIR-compliant, promoting open science and enabling the reuse and exploration of MD trajectories by the broader research community. Users can access detailed post-simulation analyses through interactive visualizations, making the data easily interpretable and usable.
7. The primary goal of DynaRepo is to enable the study of macromolecular dynamics at scale, which is essential for advancing our understanding of biological processes and developing therapeutic strategies. The data in DynaRepo can be used for various applications, including binding site prediction, functional annotation, interface detection, and binding affinity estimation. The repository's extensive analyses and interactive visualizations make it a powerful tool for researchers in biology, biophysics, and therapeutic discovery.
📜Paper: https://t.co/gvWldTHYz8
#DynaRepo #MacromolecularDynamics #MolecularDynamics #DeepLearning #ComputationalBiology #OpenScience
This is so true! Documentation is SO valuable! Good documentation allows me to use the tool. Great documentation teaches me the tool - the whys and hows of its working.
1/ I’ve reviewed hundreds of bioinformatics GitHub repos in my career.
Here’s the brutal truth: most tool documentation fails the people it’s meant to help.
And it’s not because the algorithms are bad.
Are frontier AI models really capable of “PhD-level” reasoning? To answer this question, we introduce FormulaOne, a new reasoning benchmark of expert-level Dynamic Programming problems. We have curated a benchmark consisting of three tiers, in increasing complexity, which we call ‘shallow’, ‘deeper’, ‘deepest’.
The results are remarkable:
- On the ‘shallow’ tier, top models reach performance of 50%-70%, indicating that the models are familiar with the subject matter.
- On ‘deeper’, Grok 4, Gemini-Pro, o3-Pro, Opus-4 all solve at most 1/100 problems. GPT-5 Pro is significantly better, but still solves only 4/100 problems.
- On ‘deepest’, all models collapse to 0% success rate.
🧵
Megalodon 🦈 grounds genAI for chemistry in physics: an E(3)‑equivariant transformer considers rotational symmetry while QM‑energy checks keep outputs within ~3 kcal/mol of xTB minima so that nearly every molecule is designed thermodynamically realistic. 🔥⚛️
Try it out today.👇
#AIForScience #DrugDiscovery
Just found a guide to writing a GUI from scratch in Wayland! This is from the same dude as yesterday, who made a GUI in Assembly, now doing it in Wayland with C! This dude is next level, you can learn so much from this
🔥HOW MUCH DATA DO YOU NEED to train an #AI model to generate realistic photos, videos or medical images, without memorizing* the training samples??
*which is not great for privacy/copyright/lawsuit reasons Answer below💡🧵👇
We present Thera🔥: The new SOTA arbitrary-scale super-resolution method with built-in anti-aliasing. Our approach introduces Neural Heat Fields, which guarantee exact Gaussian filtering at any scale, enabling continuous image reconstruction without extra computational cost.
@GeneSmi96946389@draparente I'm new to this stuff. Could you tell me what DNA foundation models try to do, and what is the "variance" and "additive effects" in your context? Ofc, I don't mind looking it up, but I thought I'd get a first hand perspective!
Protein engineering mfers will be like “see it works like this” and then show you the most broken slinky ass looking pile of random squiggles ever to appear in a scientific figure
Introducing ESM Cambrian.
Unsupervised learning can invert biology at scale to reveal the hidden structure of the natural world.
We’ve scaled up compute and data to train a new generation of protein language models. ESM C defines a new state of the art for protein representation learning.
"Beyond Human-Like Processing: Large Language Models Perform Equivalently on Forward and Backward Scientific Text" Our take is that large language models (LLMs) are neither stochastic parrots nor faithful models of human language processing. https://t.co/pbw3Mmou4G 1/2
If you cluster language model features (from SAEs) into two groups based on whether they tend to fire together in the same document, you find two "lobes" of features that also turn out to be geometrically distinct! Math vs. prose features separate in this t-SNE plot...