🧬 Training and evaluating AI scientists on drug target discovery is hard.
We introduce DrugTargetWorld: Synthetic Biomedical Worlds for Training and Benchmarking AI Scientists
The environment turns end-to-end drug target discovery into a problem where AI scientists can be evaluated against causal truth.
🎯 We’ve been able to hill climb on problems like Chess, Go, StarCraft, and even now computer science and mathematics because they have a verifiable reward.
We want to bring that to drug discovery.
Built with @BGomes_1905@euanashley and an incredible team at Stanford and beyond.
Until now, inserting a gene size change into the genome only worked in dividing cells. Now, you can write a whole gene into a chosen spot in the genome of a cell that will never divide again (e.g., neurons, etc).
https://t.co/Q1ni1rYP3V
CRISPR screens are great at telling you what a gene is doing inside your cells. But what if you're interested in what's happening to its neighbors?
Today, I'm excited to share match-seq, a new method that deciphers cell-cell interactions using barcoded RNA transfer.
1/9
Exciting new insights on CpG islands (CGIs) regulation by transcription factors (TFs)! CGIs drive most transcription initiation with unclear regulation. We find that chromatin-opening TFs are key players—following a surprisingly simple rule.
https://t.co/mXkDodTALR
1/9
(1/n) Super excited to share that our preprint is out today in @NatureSMB with a new name "Integrated MINFLUX tracking reveals two distinct chromatin dynamics classes across cell types" and more than 2x more data: https://t.co/M87IqgI8mk
MIT NEWS: https://t.co/Y2i8BagHUY
Traditional single cell transcriptomics doesn't capture most non-coding RNAs. In this work, led by @alinaisakovaSci and @StephenQuake, we introduce TotalX, a @10xGenomics-compatible pipeline that captures both coding and non-coding transcripts. Out now in @NatureBiotech! 1/7
Which mutations rewire function of regulatory DNA?
Excited to share SEAM: Systematic Explanation of Attribtuion-based Mechanisms. SEAM is an explainable AI method that dissects cis-regulatory mechanisms learned by seq2fun genomic deep learning models.
Led by @EESetiz
1/N 🧵👇
How do protein language models (PLM) think about proteins?🧬 We answer this w/ #InterPLM, just published in @naturemethods!
Using sparse autoencoders + LLM agent, we identify 1000s of interpretable concepts learned by PLMs, pointing to new biology 🧵
Dive into the intricate connections in the mouse brain! 🧠✨
This video shows reconstructed neurons projecting from the thalamus to various regions of the cortex. Each neuron was fluorescently labeled and imaged with our cutting-edge ExA-SPIM microscope.
I reverse engineered the San Francisco parking ticket system. I can see every ticket seconds after it's written
So I made a website. Find My Friends? AVOID THE PARKING COPS.
How do language models actually develop their capabilities during pre-training? We need mechanistic insights into what's happening inside!
We used crosscoders to track linearly interpretable features across 32 training snapshots, revealing a surprising two-phase learning process.
Is euchromatin really “open”? 🧬Our new study @bioRxiv suggests otherwise. Using super-resolution imaging🔬 @shiori_iida@MasaAShimazoe reveals: Euchromatin forms condensed domains in live cells. Cohesin constrains them and prevents domain mixing. 🔗https://t.co/2iISQjxjgh (1/3)
Attention is all you need - but how does it work? In our new paper, we take a big step towards understanding it. We developed a way to integrate attention into our previous circuit-tracing framework (attribution graphs), and it's already turning up fascinating stuff! 🧵
Fantastic view of the 50 largest bilaterally symmetrical cells in the fly brain by @quorumetrix
View these cells in Codex. Click 3D view if you dare! https://t.co/lxnZPTrb9w
New #Rstats blog post: Plot Data Along a Genome with karyoploteR. Demos: genome density, per-base coverage, structural variation, GWAS Manhattan plots, combine multiple data types, gene expression results from DESeq2, epigenetic regulation from ENCODE https://t.co/lH7ifcclHE
More than 2700 3′UTRs are highly conserved.
These 3′UTRs are essential components in mRNA templates, as their deletion decreases protein activity without changing protein abundance.
Highly conserved 3′UTRs help the folding of proteins with long IDRs.
https://t.co/b6hd4AlIoX
How do non-coding variants in enhancers cause human disease?
Here, in my main PhD work with @evgenykvon, we uncover a surprising mechanism, with generalizable implications for human genomics. https://t.co/UHMAjxs6Qj n/