Congratulations to Stanford Bioengineering PhD candidate Yixin Wang on being named a 2027 #SiebelScholar! 🎉 Yixin has made significant contributions to AIMI as a Research Mentor and Program Technical Lead. We’re proud to see her recognized! 🔗 https://t.co/xeJhNw4cXQ
“AI agents will outperform humans at almost all jobs by 2026–2027.” - The forecast is everywhere.
So we built the exam to test that claim, on real labor-market aligned work. On the hardest tier, top agents pass 2.6%.
Meet Agents' Last Exam (ALE), a rolling benchmark measuring whether agents can actually do real jobs. 🧵👇
⛏️Overcomes slice-to-volume limitations via 3D blockface as intermediate.
🔍 Brewster’s-adjacent optics → distortion-free blockface volumes 📏 TIRL MIND-based 3D & 2D registration → Dice ≈ 90–95% 🧪 Validated in human & pig brain tissue
📖 Read more: https://t.co/7mAOTjDs7q
🧠 Can we bridge MRI & histology down to the micron?
Our new study in Imaging Neuroscience introduces Brewster’s Blockface Quantification (BBQ) enables accurate mapping of MRI biomarkers to cellular pathology in #Alzheimers … 💻 Code: https://t.co/Yreq2UwHu0
New paper in Imaging Neuroscience by Yixin Wang, Michael Zeineh, et al:
Precise MRI-histology coregistration of paraffin-embedded tissue with blockface imaging
https://t.co/JjlTLniape
🚨 New to single-cell analysis? Still stuck in a trial-and-error loop?
🚨 Already experienced? Your “optimal” preprocessing pipeline might not actually be optimal.
🚨 Relying on a BioAgent? It likely runs standard procedures that are broadly accepted — but not truly optimal for your dataset or downstream task.
No matter where you are — 💡DANCE 2.0 is here for you. We turn single-cell preprocessing from a trial-and-error guessing game into a systematic, data-driven, and interpretable workflow.
🔍 Highlights of DANCE 2.0:
1. Overview
DANCE 2.0 automatically searches for optimal preprocessing pipelines tailored to your new methods and datasets.
It’s not just a recommendation engine — we open the black box of preprocessing and reveal which patterns hurt or help your downstream task performance. (See Figure 1)
2. Method-Aware Preprocessing (MAP)
Have a new downstream task method? Just plug it into DANCE 2.0.
We’ll run a systematic pipeline search to find what preprocessing works best for it.
Our large-scale experiments show DANCE 2.0 significantly outperforms original study pipelines across tasks like annotation, clustering, imputation, joint embedding, and cell type deconvolution. (See Figure 2)
3. Dataset-Aware Preprocessing (DAP)
Need to preprocess a new dataset?
DANCE 2.0 uses similarity-based atlas matching to recommend optimal pipelines on the fly.
We built a preprocessing atlas where each dataset is paired with its optimal pipeline (pre-computed via DANCE 2.0).
Your query dataset is matched to similar ones, and inherits the best preprocessing accordingly. (See Figure 4)
4. Interpretable Insights: Unlocking the Preprocessing Black Box
Through extensive experiments, we found that even within the same task, the optimal preprocessing varies across datasets and methods — depending on dataset characteristics like tissue type, sparsity, or technical noise.
DANCE 2.0 exposes and explains these hidden dependencies. (See Figure 5)
5. A Scientific Asset
We executed over 325,000 pipeline searches across 6 major downstream tasks.
Each run is labeled with task-specific performance, offering an unparalleled resource for benchmarking and insight discovery.
🔗 Paper: https://t.co/haPcpab7Wi
💻 Github: https://t.co/7s70qvXu4j (will be open-sourced very soon!)
🌐 Web platform coming soon! You'll be able to upload your new dataset and get on-the-fly recommendations for optimal preprocessing — powered by DANCE 2.0.
DANCE 2.0 will remain an open-source initiative — and we warmly invite the community to contribute, extend, and build upon it.
With amazing collaborators @Stanford@CMUCompBio@StanfordMed
Zhongyu Xing, @YixinxinWang@RemyLau3@ShengLiu_@ZhiHuangPhD@WenzhuoTang@Yuyingxie@james_y_zou@Xiaojie_Qiu@jmuiuc Guoxian Yu @tangjiliang
🚀Extremely excited to announce our new Foundation Model for single-cell, Tabula -- a privacy-preserving foundation model of single-cell transcriptomics
with federated learning and tabular modeling.
Pre-print available: https://t.co/XaF4MVcBPD
Open-source code and models: https://t.co/gP5TLuu9e2
Highlights of Tabula include:
💡Federated Learning: Tabula enables collaborative training across distributed clients (e.g., hospitals) without sharing raw data, preserving data privacy.
🔬Tabular Modeling: Tabula introduces a novel self-supervised pretraining objective that explicitly models the tabular structure of single-cell data.
🏪 Tissue-Specific Embedders: Tabula generates tissue-specific embedders that capture unique features of individual tissues, especially in zero-shot settings.
⚒️ Downstream Task Benchmarking: Tabula outperforms state-of-the-art Foundation Models in various downstream tasks below with only half the data for pretraining.
Gene-level:
- Genetic perturbation and reverse-perturbation predictions
- Gene imputation
Cell-level:
- Cell type annotation
- Multi-batch integration
- Multi-omics integration
🎸 Biological Verification: Tabula accurately reveals pairwise and even combinatorial regulatory logic across diverse biological systems, including hematopoiesis, pancreatic endogenesis, neurogenesis, and cardiogenesis using one model without any fine-tuning.
Kudos to team members Jianhui (@JilinJJ) and Shiyu (@shiyu_jiang23 ) for their incredible dedication to this project and the guidance and support from Xiaojie (@Xiaojie_Qiu ), Jiliang (@tangjiliang ), and Min Li.
We are thrilled to share our new single-cell foundation model, Tabula (preprint: https://t.co/hxGu7X8p4z; package: https://t.co/auxfUPZwhS)—a privacy-preserving predictive foundation model for single-cell transcriptomics, leveraging federated learning and tabular modeling.
Over the past year, we’ve seen a surge in foundation models for single-cell genomics, where genes are often arbitrarily ordered to mimic NLP paradigms. Furthermore, as we start to train on large datasets comprising thousands of individuals, the ethical and privacy concerns arise as well. To address these challenges, we introduce Tabula: a federated-learning-based, privacy-preserving foundation model that explicitly represents single-cell data using tabular modeling. Tabula demonstrates excellent performance across diverse tasks, including cell type annotation, multi-omics and multi-batch integration, gene imputation, denoising, and both gene perturbation and reverse perturbation predictions. Very excitingly, as one of the first examples of a truly predictive foundation model, Tabula accurately uncovers pairwise and even combinatorial regulatory logic across diverse biological systems, including hematopoiesis, pancreatic endogenesis, neurogenesis, and cardiogenesis, all of which have very well validated regulatory networks. For more details, please see @JiayuanDing 's excellent post here: https://t.co/db1qeUgic5
This is really an amazing collaboration with three brilliant young trainees, including @JiayuanDing Jianhui (@JilinJJ) and Shiyu (@shiyu_jiang23) and two other labs Jiliang (@tangjiliang), and Min Li..
Kudos to @JiayuanDing , the incredibly talented PhD student in my lab who led this project! Jianyuan is currently on the faculty job market and would be an outstanding addition to any institution. Please consider him and feel free to reach out to Jianyuan or me for more information.
Additionally, Shiyun @shiyu_jiang23 , who contributed to innovative approaches for pairwise and combinatorial perturbation prediction, is applying to PhD programs. Please consider this rising star to join your PhD program as well!
This work represents another important advance in my lab’s long time vision to establish a predictive “virtual embryo” model for human health. We are currently extending this approach to 3D Spatial Transcriptomics data as we previously reported in Spateo (https://t.co/fSexZQeejN) . If you are a developmental biologist, technology developer, or a machine learning expert, please consider joining us on this exciting journey, please reach out regarding potential positions at all levels in my lab (https://t.co/qpOqV6MRvg. Email: [email protected]). We also highly welcome graduate students from Stanford
@StanfordEng@DevBioStanford@ChemSysBio@StanfordData@StanfordAILab
for rotations in my new lab. I am excited about many collaboration opportunities within Stanford and the broader Bay Area as well.
3/n:MMedAgent answers 3 questions: (1)What’s the performance of MMedAgent in addressing diverse medical tasks across various modalities(2)Does MMedAgent exhibit superior performance in open-ended biomedical dialogue?(3)What’s the efficiency of MMedAgent in incorporating new tools
🚀 Our recent paper MMedAgent has been accepted by EMNLP! We built the first LLM-based Multimodal
, Multitask Medical AI Agent 🤖 #emnlp#MedicalAgent
📖 Paper link: https://t.co/3r6ICVqlZd
👩💻 Github: https://t.co/whFhYxvGKD
2/n: We build the first open-source instruction tuning dataset for general-purpose multi-modal medical agents. All our codes/data are open source: https://t.co/whFhYxvGKD
Exciting News! Our DANCE version 1, "DANCE: a deep learning library and benchmark platform for single-cell analysis" is now finally published in Genome Biology (@GenomeBiology ) 🎉 !!!
DANCE has impacted the field, and got 290+ GitHub stars 🌟 before its official publication!
2023 Recap In AI
2023 will go down in history as the first year when AI innovation started inflecting... It was also a seminal year for open-source AI. Here is a quick re-cap
January
- ChatGPT becomes the fastest-growing app in history
February
- Meta launches Llama-1, under a research license
- The first big usable open-source model sparks the 1st round of research and innovation
- Runway launches Gen 1, 1st gen AI video synthesis model based on Stable Diffusion
March
- GPT-4 was launched, and no one has yet been able to outperform this LLM!!
- Google launches Bard
- LLM sys chatbot arena to compare multiple LLMs.
April
- Drake, The Weeknd - Heart on My Sleeve (AI cover by Ghostwriter) - AI-generated song gets 20M+ views
May
- DPO paper publication. New technique to fine-tune LLMS that is much simpler than RLHF
- QLoRA paper publication. Efficient Finetuning of Quantized LLMs
July
- The 1st usable open-source model, Llama-2, with. a commercial license, was launched! Several open-source companies get started!
August
- Several open-source fine-tune based on Llama-2
- Llama-2 70B deployed in real-world production scenarios
September
- DALL-3 integrated into ChatGPT. LLMs and vision models can communicate with each other!
October
- Several companies, including Abacus AI, offer open-source-based fine-tunes, inference, and retrieval APIs
- Page on LLMs being world models. Language Models Represent Space and Time
- AI executive order dropped 😢
November
- Grok from xAI is the first uncensored LLM!
- OpenAI launches GPT-4v and turbo, reduces GPT-4 prices
- Stable Diffusion Video launch
- Orca paper - Teaching small models how to reason.
- Emu from Meta is a text-to-video model that can generate entire videos from text prompts.
December
- Mistral MoE open-source drop! Several GPT-3.5 class models are open-source.
- Midjourney v 6.0 can handle text in images and can create stunning photorealistic images
- Google announces Gemini Ultra and performance benchmarks comparable to GPT-4
2023 is just the beginning; there will be more acceleration in 2024! Saying Goodbye to 2023 and thankful for the speed, the spirit, and the passion in the open-source community! ��🙏
Welcome 2024, Can't-Wait to Keep Building!🚀🚀🚀🚀
Releasing Cellpose 2.0 today! You can now train your own state-of-the-art models in <1h, all from the GUI. Massive improvements for some images!
paper: https://t.co/BlG8mcwZNu
code: https://t.co/O7DgqGs1ax
Full story below 👇. #cellpose with @computingnature