Very excited to launch BroadBox alongside our new science sandboxes preprint and especially to share MelanomaBox, which we’ve been building with @jacksonweir4, @SandeepKambham2, @kbryanhsu, @shantanuXsingh, @PardisSabeti , and @insitubiology.
Bryan and I had been developing ways for coding agents such as Claude Code to operate lab robots. Together with Jackson, Sandeep, and Shantanu we connected that automation to a real melanoma drug-combination experiment designed to run on a weekly cadence with minimal human labor.
Bryan and I were also part of the preprint team developing the broader science sandboxes framework. Arya and I saw how that framework and our experimental infrastructure could come together. The efforts converged, and we decided to open these sandboxes to the community as shared challenges for AI scientists. From that synthesis came BroadBox.
If AI is to take part in discovery at all, we must give AI scientists experiments they can actually run, and then measure how evidence forces them to revise or abandon a conjecture. We must also give them a benchmark that is not saturated: nature.
We’re excited to maintain and grow BroadBox across biology. Bring your agent or bring an experiment that should become a sandbox!
https://t.co/UpRZqy99YS
Last week, we released CATv1, which is a collection of ~7,500 Cherimoya models trained to predict chromatin accessibility across diverse human cell types and tissues. Now, we've integrated this collection into Chorus, the conversational genomics project from @lucapinello
Thanks to everyone that participated in the @arcinstitute-@scverse_team workshop on modeling cellular perturbation data!
The turnout+engagement was amazing and we plan to organize similar meetings that discuss ML but are grounded in real lab experiments.
Highlights 🧵(1/6)
TissueMosaic, our method to study how changes in tissue structure across conditions affect cell-intrinsic function, is now out @CellSystemsCP!
https://t.co/ctFE33aFPP
7 years ago, I met a junior fellow named @JD_Buenrostro who blew me away with a vision of futuristic genomic technologies
Today, we (@ajaylabade31, @carolinecomenho) are excited to share our first steps into that future: Expansion in situ genome sequencing
1/
Stellar work by @_michellemli. Nothing exists in a vacuum in Biology. Protein representation learning has focused on learning on sequence or structure but has not considered cell-type specific protein interactions. Including this context via PINNACLE can improve performance!
Defended my PhD from @harvardmed yesterday! What an incredible 4 years. Could not be more grateful for the opportunity to work with and learn from so many amazing people.
Introducing the SPECTRA python package!
https://t.co/a1dKDRmtGo
This package implements the spectral framework for model evaluation. All you need to get started is (1) a model, (2) a dataset, and (3) a definition of sample to sample similarity!
It is no secret there exists a generalizability problem in AI for biology. Despite all the advances in ML methods, the way we evaluate generalizability in ML models has not changed. How well does your ML model generalize across the entire spectrum of possible dataset splits? [1/9]
1/6 🍍🍕🧬 What do a photo of Italians savoring pineapple pizza and a synthetic DNA sequence have in common? Both can be intricately crafted by generative AI! Introducing DNA-Diffusion, our AI model set to advance #SyntheticBiology and #GeneRegulation: https://t.co/rK4MzRnvw7
To all who like this - what if I told you this book is not about the linear deterministic kind of SSM used by Mamba, S4, etc. but instead is about Bayesian inference in stochastic nonlinear dynamical models - would you still be interested?
Very excited to share our study on the benchmarking of methods to detect spatially variable genes (SVGs) and peaks (SVPs), a collaborative work with @zafateniac, @SongDongyuan, @GuanaoYan, @jsb_ucla, and @lucapinello. 🧵 (1/n)
1/7 How do 1000s of small genetic changes cause complex diseases? The core gene model helps to explain and we developed a novel method base on graph neural networks to identify core genes for 5 groups of complex diseases. https://t.co/EWmHyrpclF
Excited to unveil today CRISPR-CLEAR at #ASHG23, our new CRISPR screen modality to decode genotype-phenotype relationships at nucleotide & variant-level resolution. Join us at Conv Ctr/Ballroom A/Level 3, 2:15pm-2:30pm. #Genomics#CRISPR
Aneuploidy is a defining feature of cancer cells, but is it also present in healthy normal tissues? In our paper out today on @NatureGenet, we report hundreds of mosaic chromosomal alterations (mCAs) found in diverse tissues from #GTEx. A thread (1/7)
https://t.co/JtzP7d1oRL