Why haven't organoids solved all of drug discovery?
(5.8k words, 26 minutes)
here, i try to answer a question that i imagine has been on peoples minds for awhile. first essay of a three part series over organoids, because this subject is complicated
https://t.co/b25zgVlt2M
Read the NEO-unify blog post from SenseNova about encoder-free unified multimodal models and got nerd-sniped. Built a toy-scale version from scratch with MLX on synthetic 16x16 images to see how the Mixture-of-Transformers idea actually behaves. Went through 5 phases from VQ-VAE baselines to full MoT with CFG and latent-space flow matching.
Sharing the 2nd NEO-series work: NEO-unify✨
We revisit first principles:
can a model interact directly with pixels 🖼️ and words ✍️?
Yes! No VE! No VAE!
Three months exploring an idea we’ve long believed in:
multimodal may not need encoders at all.🚀
Thanks for all partners!
I’m excited to share one of the last pieces of my PhD dissertation: CELLECTION, a framework for predicting emergent phenotypes from single-cell populations.
https://t.co/f6zhZpnwgd
Excited to share #AlphaGenome, a start of our AlphaGenome named journey to decipher the regulatory genome! The model matches or exceeds top-performing external models on 24 out of 26 variant evaluations, across a wide range of biological modalities.1/6
🔥 Super cool research in longevity by Shift Bioscience!
A new preprint ("A single factor for safer cellular rejuvenation") from Shift Bioscience reports a striking finding: a single gene, SB000, achieves robust cellular rejuvenation across two germ layers—fibroblasts and keratinocytes—without inducing pluripotency. This could mark a meaningful step toward safer partial reprogramming.
Using transcriptomic aging clocks, the authors show that SB000 reverses age-associated gene expression, reduces senescence markers, and maintains cell identity—outperforming OSK across dozens of epigenetic clocks, including Horvath and DunedinPACE.
This matters. While OSK-based reprogramming has dominated the field, its translational barriers—tumorigenicity, dedifferentiation, and delivery—remain high. A factor that reverses biological age without disrupting cell function could change the game.
But perhaps the deeper story here is how they got there.
This discovery wasn’t driven by a one-at-a-time gene screen. It relied on a scalable, high-throughput computational pipeline—a blueprint for how AI-first biology can unlock novel rejuvenation strategies.
This is exactly the promise of foundation models like scGPT: large-scale pretraining on single-cell atlases to learn the latent rules of gene regulation, cell state, and age progression. These models can predict perturbation effects, identify rejuvenation signatures, and even suggest interventions de novo.
The convergence is clear:
--AI enables global search over the gene regulatory space. A one-at-a-time gene screen didn’t drive this discovery
--Foundation standards contextualize interventions across tissues, age, and health status.
--Experimental biology validates, refines, and feeds back into the loop.
Shift's SB000 is an exciting advance. But it’s also a case study of where aging research is going: from hypothesis-driven, to model-guided; from single-gene overexpression, to generative design grounded in cellular language.
The future of aging is being written in code—both genetic and computational.
Congratulations to the whole shift bio team! @lucascamillomd@dives86@BrendanMSwain
#AIinBiology #Longevity #Reprogramming #AgingResearch #scGPT #FoundationModels
I’m excited to share that our paper, "sciLaMA: A Single-Cell Representation Learning Framework to Leverage Prior Knowledge from Large Language Models," has been accepted to #ICML2025
https://t.co/VbdDAMSHDG
🔥 Unveiling the Future of Genomics with Genome Language Models (gLMs)! 🔥
Our comprehensive review, "Transformers and genome language models," is finally published in Nature Machine Intelligence!
Link: https://t.co/hCk6EzLKDB
Key Highlights:
🔬 The Challenges Addressed by gLMs: gLMs tackle the intricate task of interpreting vast genomic sequences, enabling predictions about gene regulation, variant effects, and more.
🧠 Transformers in Genomics: Discover how transformer architectures, renowned for their success in natural language processing, are adept at capturing long-range dependencies in genomic data, leading to more accurate models.
🚀 Beyond Transformers—Introducing HyenaDNA: Explore innovative architectures like HyenaDNA, which offer efficient long-range genomic sequence modeling at single nucleotide resolution, pushing the boundaries of genomic research.
📊 Comparative Analysis of Models: We delve into the evolution from sequence-to-function models like DeepSEA and Enformer to sequence-to-sequence models such as DNABERT and Evo, highlighting their respective strengths and applications.
⚡ Strengths, Limitations, & Future Directions: Gain insights into the current capabilities of genomic AI, its limitations, and the promising avenues for future research and application.
This pivotal work is the result of a collaborative effort led by Micaela E. Consens (@micaelanonsense ), with contributions from Cameron Dufault, Michael Wainberg (@michaelwainberg ), Duncan Forster, Mehran Karimzadeh, Hani Goodarzi (@genophoria ), Fabian J. Theis (@fabian_theis ), Alan Moses.
@UHNAIHUB@UHN@VectorInst @uoftoront
#Genomics #AI #MachineLearning #Transformers #HyenaDNA #DeepLearning #Bioinformatics #GenomeResearch
1/9 🚨 New Paper Alert: Cross-Entropy Loss is NOT What You Need! 🚨
We introduce harmonic loss as alternative to the standard CE loss for training neural networks and LLMs! Harmonic loss achieves 🛠️significantly better interpretability, ⚡faster convergence, and ⏳less grokking!
Authors introduce a practical training framework to improve protein language model representations by integrating biological features and prior information through contrastive learning. #BiotechNatureComms
https://t.co/AS5QV9IC6H
Happy Chinese New Year! I'm glad to share our new work on using biological knowledge to guide the design of interpretable AI models, for prioritizing potential driving regulators for cell state transitions. https://t.co/LnqY5HgQdN
Reading group session on Monday: @ValentinDeBort1 joins us for his paper "Accelerated Diffusion Models via Speculative Sampling" https://t.co/rnx54RictU
On zoom at 9am PT / 12pm ET / 6pm CET: https://t.co/R8d1EHxdMZ
As you can probably tell from our work, we're now all in on diffusion and flow matching models for protein/peptide sequence generation! 🌟 These models are just way better for tightly controlling generation (vs. say autoregressive models, which we've found are poorly set up for biological tasks). As such, while, we have recently moved to discrete diffusion architectures to operate purely in sequence-space (i.e. PepTune, MeMDLM, etc.), you may remember that we got our start early in 2024 by doing continuous diffusion on pLM latents with our AMP-Diffusion paper (https://t.co/RHMufepRYX) to generate antimicrobial peptides! 🦠 It was beautiful work by my (now graduated!) Masters student, @LeoTZ03! 🎓
AMP-Diffusion was the first example of performing latent diffusion for protein/peptide sequence generation, and we got pretty good quality peptides with AMP-like sequence composition and physicochemical properties and strong predicted inhibitory potential! While we saw this as proving to ourselves that diffusion could work (we are definitely convinced!), we didn't think to go much further. 🙃
I then met @delafuentelab (who specializes at screening AMPs!) at @pennbioeng who offered to experimentally test out our peptides! 🤗 Through a really fun collaboration over the past year with, they've shown that AMP-Diffusion-generated AMPs demonstrate bacterial killing in vitro, and the peptides showed favorable physicochemical profiles. 🦠 In preclinical mouse models of infection, our lead peptides reduced bacterial burdens, and showed really strong efficacy comparable to polymyxin B and levofloxacin, with no detectable adverse effects! 🐁
Take a read of our collaborative preprint, where we describe these exciting results!! 🙌 I hope our work convinces you that diffusion for protein/peptide sequence generation can have meaningful translational impact! ⚕️
📜: https://t.co/ImVCpRzAXl