Research in ML for drug discovery at SandboxAQ. Previously worked in ML research @ DreamFold & MILA. Before that, ML4Astro @ Sussex, UCL. Ph.D. from Cambridge
We're so excited to release SAIR, the biggest public dataset of synthetic protein-ligand structures. Looking forward to seeing the exciting science that will come from it!
Today we’re releasing SAIR, the Structurally Augmented IC50 Repository.
SAIR is the Largest Open-Sourced Binding Affinity Dataset with Cofolded 3D Structures. It includes more than 5 million protein-ligand structures, generated using our Large Quantitative Models and labeled with binding affinity data.
By providing this unprecedented scale of structure-activity data, we aim to enable researchers to train and evaluate new AI models for drug discovery, bridging the historical gap between molecular structures and drug potency prediction.
The SAIR dataset was created using the @nvidia DGX Cloud and is now publicly available on the @Google Cloud Platform.
Access and build with SAIR today!
📰Read the Press Release: https://t.co/0yr53BFCYT
📥Learn More and Download the Dataset at https://t.co/tQC4ECJm5b
#DrugDiscovery #LQMs #SAIR #AIforScience #SandboxAQ
We will be presenting our work on PQMass, a method to quantify how good generative models are, that is statistically principled, scalable and versatile. Come check out our poster on Thursday afternoon, and message me if you want to know more!
https://t.co/dOmtX5T9z4
I am heading to Singapore for #ICLR2025 ! Send me a message if you would like to talk about AI4Science, or if you’re an old friend and just want to grab a drink!
I am happy to announce that I am returning to Madrid, after more than 11 years living abroad, and that I am starting a new position as a Staff Research Scientist in Machine Learning & Biopharma at @SandboxAQ . Excited to join this amazing team!
You heard all about AI accelerating simulations (maybe from me?), but do you know...
How can AI tell you what is in the Universe?
Our new series of work from #SimBIG team led by @changhoon_hahn published recently by @NatureAstronomy did just that!
Interesting things we did:
👉 This is the first time one simulates the Universe observed via a spectroscopic telescope (@sdssurveys) well enough to compare it to the actual Universe!
We simulated 20,000 of these universes!
👉 For each simulated universe, it gives you summary statistics (x) and the fundamental properties of the simulation (y). We then train an AI using these 20,000 pairs of (x,y) to calculate the posterior P (y | x ).
👉 Now you can give this AI observed summary statistics (x') from the real observed Universe and here come the fundamental parameters of the Universe with appropriate errors💫
By extracting non-Gaussian cosmological information on galaxy clustering at non-linear scales, a framework for cosmic inference (SimBIG) provides precise constraints for testing cosmological models. @ChanghoonHahn@cosmo_shirley@DavidSpergel et al.: https://t.co/lXiOSKc8T8
Can we perform unbiased bayesian posterior inference with a diffusion model prior? We propose Relative Trajectory Balance (RTB) which allows us to directly optimize for this posterior model. We apply this to several tasks in image, language and control!🧵https://t.co/h7elRgivcC
Do you want to work on creating the largest collection of virtual universes to date? Would you like to explore these with state-of-the-art deep-learning techniques? Do you want to use these simulations to unveil the mysteries of the Universe? Consider applying to our 4.5 months
I am thrilled to announce that I have officially started my new job as an AI Research Scientist at @DreamFoldAI. I feel grateful to have met many wonderful people during my 8.5 years in astrophysics. I am excited to use generative AI to revolutionize the way we treat diseases!
Very proud to announce...
LtU-ILI, an all-in-one framework for ML parameter inference in astro/cosmo!💻🔭🌌 https://t.co/eYBF9ZWMNq
We unified the neural nets, training methods, and validation metrics used at the forefront of ML+astro in one accessible package... (1/9)