This morning, mathematicians at OpenAI announced that a group of 10,000 autonomous AI agents under their direction proved that the Navier-Stokes equations can “blow up,” thus solving one of the six remaining Millennium Prize Problems. But this massive result is not without controversy. https://t.co/I0f1hwvETo
This feels like one of those moments when the future arrives all at once.
OpenAI has published a solution to the Navier–Stokes existence and smoothness problem. This is one of seven Millennium Prize Problems selected by the Clay Mathematics Institute in 2000, carrying a $1 million prize and more importantly representing one of mathematics’ deepest unresolved questions.
The Navier–Stokes equations describe how fluids move: air over an aircraft wing, ocean currents, weather systems, and blood flowing through our bodies. For almost 90 years, mathematicians have asked: if a three-dimensional fluid begins in a smooth state, must it remain smooth forever, or can the equations “blow up in finite time"?
OpenAI’s answer is that such a breakdown can happen. Its AI system constructed a finite-time singularity and produced both an analytical proof and a formalization in Lean (a mathematical programming language if you want, to make sure proofs are correct). Obviously independent experts must still confirm that the formalization precisely matches the official problem. The mathematical significance would be enormous, even though the construction itself is highly engineered.
But this goes far beyond an unchecked AI-generated argument. The impact is primarily foundational, ie doesnt suddenly better aircraft or weather forecasts. But it does/would settle a central question about the limits of one of physics’ most important models while demonstrating that AI can produce genuinely new, formally verified research mathematics.
I believe also Google DeepMind has pursued AI-assisted mathematical reasoning and related work on fluid-dynamics singularities, but who would have imagined, when AI mainly meant prediction, classification, and data analysis, that pure mathematics would become such an extraordinary playground?
Perhaps a broader pattern is emerging: AI appears particularly powerful at searching vast spaces for strange constructions that overturn universal claims. Earlier this year, an OpenAI model disproved Erdős’s 80-year-old unit-distance conjecture by finding an unexpected infinite family; more recently, an AI-assisted counterexample refuted the 87-year-old Jacobian conjecture in dimensions three and higher.
We may be entering an era in which mathematicians and AI explore the deepest frontiers of knowledge together. Crazy & incredibly exciting times!
[OpenAI announcement](https://t.co/En4JM2zVSg)
[Navier–Stokes proof paper](https://t.co/IxAeVDgc3I)
[Related Euler proof paper](https://t.co/6b1bveCBBf)
#AIforScience #Math #NavierStokes
I'm currently returning to Toronto from a summit on the future of mathematics, at OpenAI. @SebastienBubeck asked me to talk a bit about the future we'd all like to avoid, where humans are mathematically disempowered. @Jacob_Tsimerman advised us to try to prioritize detail over correctness, and I have no doubt that I succeeded in deprioritizing correctness.
I tried to find a title that wasn't too bombastic:
Happy to share our latest work, now published in Nature!
Ligand-enabled remote C–H activation converts abundant saturated fatty acids into desaturated lactones.
https://t.co/cCpMO9idcM
https://t.co/Yl6ERXsAPG
Thanks to @NandanNilekani Foundation, @ANRFIndia@iitbombay@dm_lab
MIT is offering Books on AI & ML (ABSOLUTELY FREE):
1. Foundations of Machine Learning
https://t.co/78p57EBbL8
2. Understanding Deep Learning
https://t.co/D2oyRrXqcE
3. Introduction to Machine Learning Systems
https://t.co/hkaYi0dd1k
4. Algorithms for ML
https://t.co/lntuD4Q19H
5. Deep Learning
https://t.co/vCHVIZQYTI
6. Reinforcement Learning
https://t.co/JNWhFCuCkH
7. Distributional Reinforcement Learning
https://t.co/GXpkV4BDZi
8. Multi Agent Reinforcement Learning
https://t.co/T8zVmQVutO
9. Agents in the Long Game of AI
https://t.co/HeD3Nsm5zz
10. Fairness and Machine Learning
https://t.co/csAjhdf7Lb
11. Probabilistic Machine Learning
❯ Part 1 : https://t.co/5Leef9ypGj
❯ Part 2 : https://t.co/vRbF0rEIuh
I’m thrilled to introduce you to the NeuMap! our latest work in @Nature. A global, comprehensive, single-cell transcriptional atlas of neutrophils across 47 biological conditions in human and mice. A real tour-de-force 🗺️ @AndrsHidalgo16 https://t.co/W8X6eOUsKc
I'm recruiting multiple PhD students this cycle to join me at Harvard University and the Kempner Institute! My interests span vision and intelligence, including 3D/4D, active perception, memory, representation learning, and anything you're excited to explore!
Deadline: Dec 15th.
Understanding generative AI output with embedding models
Deep neural networks don’t just spit out answers—they quietly build geometric representations of the world. Every text, image, or audio sample passing through a foundation model is turned into a high-dimensional vector, an embedding, that encodes structure we rarely see directly. The question is: can we mine those embeddings to understand what generative AI is actually doing—without retraining or opening the model?
Max Vargas and coauthors show that we can. They treat the embedding layers of large models as measurement devices and then apply simple, classical tools like principal component analysis (PCA) and linear discriminant analysis (LDA) on top. Across text and image domains, they find that embeddings naturally separate data by intuitive factors: language, topic, even whether a text is a human-written news article or a machine translation. Most strikingly, they show that embeddings from foundation models exhibit intrinsic separability between real data and content generated by AI systems.
The authors go further and introduce the idea of “generative DNA” (gDNA): statistical fingerprints that generative models leave in the embedding space. Using off-the-shelf embedding models, they can distinguish not only real vs AI-generated content, but also outputs from different LLMs, diffusion models, or even different prompts to the same model—and do so with simple linear methods and outlier detectors. In contaminated corpora (e.g., human answers mixed with a small fraction of LLM-produced replies), projecting embeddings and running isolation forests is enough to surface synthetic samples that look perfectly fine to human readers.
The implication is powerful and a bit sobering. Even when generative models produce outputs that seem realistic, their latent geometry reveals systematic biases and shifts from true data distributions. Embedding-based analysis offers a lightweight, model-agnostic way to audit generative systems: for deepfake detection, for monitoring AI-written scientific text, for comparing models, or for probing what our LLMs are really optimizing for. As AI-generated content floods the web, tools that “look inside” embedding spaces may become a core part of how we keep generative AI robust, interpretable, and trustworthy.
Paper: https://t.co/Prh4D4bF88
We are looking for PhDs and Postdocs!
So proud of my students on achieving so many amazing things during their "very first year".
I have been asked many times how I like being faculty, especially with funding cuts. My answer is always "it is the prefect job for me"! Still deep in the honeymoon phase.
The only reason is the students are so amazing, making my transition so much easier. One year in, they already collected paper awards, orals, spotlights, etc
What makes me proudest is they are vividly alive: curious, playful, confident in their own weird way, light up when talking about ideas, and never afraid to explore "the thing might fail".
Everyone is just… themselves. And somehow, that version of themselves keeps shipping amazing work.
In today's anxious academic world, this kind of aliveness is what I will try best to protect.
Maybe the best part of being an advisor is that every student is so different and unique lol
Interestingly, coming to second year, they've got their own passions, I can't just plug my ideas into their heads. So when I get excited about sth new, my first thought is: "Okay, time to find some fresh first-years who will be thrilled about this!"
MLL lab is 1 year old, we started right in Oct 2024. We are growing and looking for more phds to join us!
1. Why our lab? (1/2)
2. Why @northwesterncs? (2/2)
In 2025 alone: NU has 7 faculty as Sloan Fellows, plus a Nobel winner! Check more below
Deep generative classification of blood cell morphology
Blood smears under the microscope are still a cornerstone of haematology. Yet automating what experts do by eye is hard: subtle shape changes, noisy imaging conditions and rare, atypical cells all conspire to break standard deep learning models that only learn a decision boundary between classes.
Simon Deltadahl and coauthors introduce CytoDiffusion, a diffusion-based generative classifier that learns the full distribution of how each blood cell type can look, rather than just where to draw the line between them. Trained on real clinical data, it generates synthetic blood cell images that expert haematologists can’t reliably distinguish from real ones (performance close to random guessing), a strong signal that the model has captured true morphology instead of dataset shortcuts. On top of that, it beats state-of-the-art discriminative models in anomaly detection (AUC 0.990 vs 0.916), robustness to domain shifts across labs and scanners (0.854 vs 0.738 accuracy) and performance in low-data regimes (0.962 vs 0.924 balanced accuracy).
Because CytoDiffusion models the underlying distribution, it also gets a handle on uncertainty. Using tools from psychophysics, the authors show that its confidence scores track actual correctness more reliably than those of human experts, and are better calibrated than transformer-based baselines. And instead of opaque saliency maps, the model provides counterfactual heat maps—highlighting which parts of a cell would need to change for the image to be classified as a different type, aligning closely with how haematologists describe morphology.
Generative diffusion models like this point to clinical AI systems that are not only accurate, but also robust to real-world variability, transparent in their reasoning and explicit about when they are unsure—key ingredients if AI is going to become a trusted partner in diagnostic workflows.
Paper: https://t.co/gmgyGHSPdE
After 9 months, 5 rounds of chemo, and getting to ring the cancer-free bell, we got to come home today from @StJude. Definitely counting our blessings over the holidays and so happy to be home.
STORIES: Learning cell fate landscapes with spatially aware optimal transport
Understanding how cells differentiate over time has often been described using the metaphor of an “epigenetic landscape,” where cells move downhill from stem-like states toward committed identities. But in real tissues, this journey unfolds not only in gene expression space, but in physical space as cells reorganize, migrate, and interact with their environment.
Spatial transcriptomics now allows us to capture both dimensions simultaneously, yet most computational trajectory methods still treat gene expression alone, discarding geometry. As a result, they may reconstruct differentiation timelines while missing how the structure of the tissue shapes fate decisions.
Huizing and coauthors introduce STORIES, a machine learning approach that learns a differentiation potential directly from spatiotemporal transcriptomic data. The method trains a neural network to represent this potential, and interprets differentiation as a Wasserstein gradient flow, so that the gradient of the learned potential defines the direction in which a cell’s gene expression should evolve over time. The central challenge in spatial data is that tissues can change shape from one time point to another: slices may rotate, expand, or reorganize. Simply using raw coordinates would make the model sensitive to these incidental transformations.
To overcome this, STORIES trains the neural network using Fused Gromov–Wasserstein optimal transport, a formulation that compares cell populations in a way that is invariant to rotation, translation, and deformation. The approach jointly considers gene expression similarity and the relative geometry of cells in space, allowing the model to infer how differentiation unfolds without manually aligning tissue slices. By learning from spatial relationships rather than fixed coordinates, STORIES captures how the microenvironment influences fate decisions.
On large-scale datasets spanning mouse development, zebrafish embryogenesis, and axolotl neural regeneration, STORIES achieves greater biological coherence than existing optimal transport–based trajectory inference models. It not only predicts future gene expression states more accurately but does so while preserving the spatial organization of the tissue. The resulting landscapes recover known differentiation pathways and highlight new regulatory genes that may drive transitions, offering mechanistic hypotheses grounded in both molecular and spatial context.
This study suggests a direction toward a new generation of computational models where cell fate is not inferred in isolation, but understood as a process embedded in the architecture of living tissues.
Paper: https://t.co/3sCvMkfzTH
Squidiff: Generative diffusion models for predicting cell fate and perturbation responses
Single-cell sequencing lets us observe how individual cells differ, evolve, and respond to their environment. But while we can measure these transcriptomic states, predicting how a cell will change under a new stimulus — a gene knockout, a drug, or even something as complex as radiation exposure — remains extremely difficult. Experiments are slow, expensive, and often impossible to run at the scale needed for mechanistic insight.
Squidiff proposes a different path. This framework combines a semantic encoder with a conditional diffusion model, enabling the generation of new transcriptomic profiles by iteratively denoising from latent space. The key idea is that cellular identity and environmental cues can be captured as smooth, manipulable vectors in a shared latent representation. By shifting these vectors, Squidiff can navigate cellular state trajectories: from pluripotent stem cells into differentiated lineages, across gene perturbations, or along drug-response gradients.
What makes this interesting is not just that Squidiff generates realistic single-cell data — many generative models attempt that — but that it captures transient cell states and developmental trajectories that are often inaccessible experimentally, including intermediate stages and nonlinear responses. The diffusion process encodes the underlying stochasticity, while the semantic space carries the structured biological signal, enabling controlled interpolation across time and condition.
The authors demonstrate this on multiple fronts: predicting iPSC differentiation into germ layer lineages; modeling non-additive gene perturbations without graph priors; reconstructing cell-type-specific drug responses; and, remarkably, predicting the effects of neutron irradiation and the protective effects of G-CSF in blood vessel organoids. In the organoid system, Squidiff recovered not only cell-type-specific damage signatures but also the dynamic progression of vascular disruption, and how G-CSF shifts these trajectories toward recovery — all from sparse experimental sampling.
This suggests something important. Instead of treating single-cell sequencing as a static snapshot technology, we may be moving toward generative, predictive models of cellular development, where experiments guide the model, and the model guides the next experiment. Squidiff does not replace data — it amplifies the value of each dataset, enabling in silico hypothesis generation and perturbation screening before wet-lab validation.
Paper: https://t.co/JOf6A1RT3I
🧬 Excited to share Nicheformer out now in Nature Methods!
A transformer foundation model linking single-cell & spatial omics, learning spatial context from gene expression to map tissue organization.
Led by Ale Tejada & Anna Schaar 👏
👉 https://t.co/ba9DX7h2xg
Terrific work from @KyleFerchen, @nsalomonis, @LeeGrimesLab, and colleagues that advances our understanding of hematopoiesis by integrating multiple single-cell data types in @NatImmunol: https://t.co/XhZN6SlzGy