Excited to share our latest survey paper: "Artificial Intelligence for Science in Quantum, Atomistic, and Continuum Systems"! 🚀 ArXiv: https://t.co/lr97U1VjCz. Feedback is warmly welcomed! Some highlights are as follows:
(1/n)
AI+Science book is now up,
After years of work, our book on
Artificial Intelligence for Science in Quantum, Atomistic, and Continuum Systems
is available at Foundations and Trends® in Machine Learning.
We cover the history of works in the intersection of AI and Science, to novel breakthroughs and advance in foundation of AI.
From neural networks to neural operators,
From PINNs to PINO and so on,
From computer vision to quantum, graph data, grid data, point cloud data, and GenAI in sciences,
From classical reduced order methods to learned ones.
Etc.
With an extraordinary list of co-authors
Xuan Zhang, Limei Wang, Jacob Helwig, Youzhi Luo, Cong Fu, Yaochen Xie, Meng Liu, Yuchao Lin, Zhao Xu, Keqiang Yan, Keir Adams, Maurice Weiler, Xiner Li, Tianfan Fu, Yucheng Wang, Alex Strasser, Haiyang Yu, YuQing Xie, Xiang Fu, Shenglong Xu, Yi Liu, Yuanqi Du, Alexandra Saxton, Hongyi Ling, Hannah Lawrence, Hannes Stärk, Shurui Gui, Carl Edwards, Nicholas Gao, Adriana Ladera, Tailin Wu, Elyssa F. Hofgard, Aria Mansouri Tehrani, Rui Wang, Ameya Daigavane, Montgomery Bohde, Jerry Kurtin, Qian Huang, Tuong Phung, @MinkaiX , Chaitanya K. Joshi, Simon V. Mathis, Kamyar Azizzadenesheli, Ada Fang, Alán Aspuru-Guzik, Erik Bekkers, Michael Bronstein, Marinka Zitnik, @AnimaAnandkumar , @StefanoErmon , Pietro Liò, @yuqirose , Stephan Günnemann, @jure, Heng Ji, Jimeng Sun, Regina Barzilay, Tommi Jaakkola, Connor W. Coley, Xiaoning Qian, Xiaofeng Qian, Tess Smidt, @ShuiwangJi
Our 500+ page AI4Science paper is finally published:
Artificial Intelligence for Science in Quantum, Atomistic, and Continuum Systems. Foundations and Trends® in Machine Learning, Vol. 18, No. 4, 385–912, 2025
https://t.co/RzxYTDOJwx
A Benchmark for Quantum Chemistry Relaxations via Machine Learning Interatomic Potentials
1.PubChemQCR is the largest publicly available dataset of DFT-based molecular relaxation trajectories, with 3.5 million molecules and over 300 million conformations, including 105 million computed with DFT. Each conformation includes total energy and atomic force labels.
2.The dataset captures full geometry optimization trajectories, not just final structures—addressing a key gap in previous datasets. This enables machine learning interatomic potentials (MLIPs) to learn from both stable and non-equilibrium geometries.
3.PubChemQCR offers broad chemical diversity, spanning 25 elements and a wide range of molecular sizes and conformational complexities. It was built from PubChemQC’s raw optimization outputs, spanning PM3, Hartree–Fock, and DFT stages.
4.Compared to existing datasets like QM9, GEOM, or ANI-1x, PubChemQCR provides significantly more conformational data, better element coverage, and crucial force labels at high-accuracy DFT level—making it uniquely suited for training MLIPs.
5.A curated subset, PubChemQCR-S, contains \~41K DFT relaxation trajectories for efficient model benchmarking. This subset supports rapid prototyping, ablation studies, and hyperparameter tuning.
6.The authors benchmarked 9 MLIP models (SchNet, PaiNN, NequIP, FAENet, Equiformer, etc.) on energy and force prediction tasks using PubChemQCR-S. Equiformer achieved the best overall performance on both energy and force metrics.
7.In geometry optimization tasks, Equiformer outperformed all other models, achieving 70.15% average energy minimization, 23.81% chemical accuracy success rate, and a 19.85% force convergence rate. Most other models struggled, especially with force convergence.
8.The dataset supports supervised pretraining of 3D molecular models with physically grounded energy and force labels—potentially benefiting downstream property prediction tasks in drug discovery and materials science.
9.It also enables training of generative models for 3D molecular structures. These models can learn to generate low-energy conformations directly from the data, bypassing costly DFT optimization.
10.Limitations include the dataset's near-equilibrium bias (due to DFT relaxation) and inconsistent label quality across optimization stages. Also, chemical element coverage is capped at 25 due to DFT method constraints.
11.Despite limitations, PubChemQCR is a foundational resource for building accurate, transferable, and data-efficient MLIPs. It can accelerate atomistic simulations, geometry optimization, and generative modeling in quantum chemistry.
💻Code: https://t.co/8P56b6Yh7h
📜Paper: https://t.co/SQTZYZnugP
#QuantumChemistry #ML4Science #DFT #GraphNeuralNetworks #MolecularSimulation #MachineLearning #OpenScience #MolecularModeling
How to become expert at thing:
1 iteratively take on concrete projects and accomplish them depth wise, learning “on demand” (ie don’t learn bottom up breadth wise)
2 teach/summarize everything you learn in your own words
3 only compare yourself to younger you, never to others
(5/n) More details can be found in the preprints:
PubChemQCR: https://t.co/6dWI6jzsDB
TDN: https://t.co/dDtyWl7qHD
MLIP Foundation Model: https://t.co/zkMnSBhyvg
(4/n) 🧠 MLIP Foundation Model: Existing MLIP models pre-trained on PubChemQCR can serve as foundation models that can either generate low-energy geometries or be fine-tuned for downstream tasks to boost molecular property prediction. 🎯
Excited and proud mentor moment 💫! A new materials foundation model with two of my undergraduate mentees Montgomery Bohde and Andrii Kryvenko as core authors.
For Materials Foundation Models, Invariance V.S. Equivariance? Invariance + Equivariance!
Scientists @TAMU tackle the challenges of structure-based drug design with a new #AI approach. Their #Frag2Seq model applies language models to generate drug molecules fragment by fragment, showing promise in creating more effective target-specific drugs.
https://t.co/MaEKL2cjqE
Interested in learning about latent diffusion models and how they can be used for efficient protein structure generation?
Read the latest blog by @Cong_Fu_ and @KeqiangY to learn more about LatentDiff: https://t.co/LHr2hnOBcF
Advances in artificial intelligence (AI) are fueling a new paradigm of discoveries in natural sciences. Excited about our latest paper about this incredible transformation. It was a great collaboration led by @ShuiwangJi.
Excited to share our latest survey paper: "Artificial Intelligence for Science in Quantum, Atomistic, and Continuum Systems"! 🚀 ArXiv: https://t.co/lr97U1VjCz. Feedback is warmly welcomed! Some highlights are as follows:
(1/n)
(3/n)🔬 An in-depth yet intuitive discussion on symmetry, as well as explainability, out-of-distribution generalization, large language models, and uncertainty.
📖 Access categorized lists of resources to enhance learning and education.
Check out our #ICML2023 paper, “Group Equivariant Fourier Neural Operators for Partial Differential Equations” (https://t.co/K3zrfrHVJN). We solve PDEs by encoding symmetries in Fourier convolutions (1/3)