Standing in the transformer era, looking back at U-Net's greatness in segmentation is fascinating 🧠.
Before transformers, we were bottlenecked by data, so we hacked U-Net in endless ways to fit each specific task. Then nnU-Net was a wake-up call: new architecture? No❌, one architecture + the right config is enough.
Once transformers revealed the power of scale, few-shot learning + transformers took over. As data grew, one model + different prompts could adapt across tasks and domains. 🚀
Data ⇄ architecture ⇄ training methods, all interlocked, all reinforcing each other. And we're nowhere near done: new problems and needs bring new data. Will we need new architectures? New metrics? What counts as effective training by then? Can't wait! 🔥
It's hard to believe it's been over a decade since #UNet was introduced in 2015. I still remember reading that paper during my first year of my PhD and realizing how groundbreaking it was.
Now, more than a decade later, we're excited to look back and look ahead.
🔬 A decade of biomedical image segmentation, in one map (2015–2025).
https://t.co/5Mf6aY6cy6
From task-specific U-Nets to universal promptable foundation models, we summarize the key ideas, milestones, and emerging directions shaping the next generation of biomedical AI.
10 years of breakthroughs, distilled into a single timeline:
🩺 Task-specific U-Nets
➡️ Self-supervised learning
➡️ Multimodal foundation models
➡️ Universal promptable models
The field has evolved from "one model per task" to "one model for many tasks."
What I define AGI:
1. Understand the physical world and be able to interact for some good(instead of damaging something).
2. Learn OOD task with small samples.
SpaceX makes the future exciting! Maybe people in 22 century look back at us as we look back to Qing Dynasty. Can't wait to see the imagination comes true!
SpaceX a clôturé son premier jour de cotation à 2 100 milliards de dollars, +19%. Tout le monde regarde le chiffre. Personne ne regarde ce qu'il price réellement.
Laissez-moi vous dire ce que le marché vient d'acheter, et pourquoi je pense que cette boîte vaudra 30 à 50 trillions d'ici 5 ans.
D'abord, le symbole. Cette IPO est un référendum. D'un côté, 20 ans de discours sur la décroissance, la sobriété, la redistribution, la fin de l'histoire gérée par des comités. De l'autre, un homme qui a dit "je vais rendre l'humanité multiplanétaire", que tout le monde a traité de clown, et qui vient de créer la plus grosse entreprise cotée de l'histoire en partant d'un entrepôt à El Segundo. Le marché a voté. Le wokisme avait des départements RH, SpaceX avait des fusées. Les fusées ont gagné.
Ensuite, la mécanique économique, parce que c'est là que tout le monde se trompe. Les analystes valorisent SpaceX comme une entreprise de lancement plus Starlink. C'est comme valoriser Internet en 1995 sur le marché du fax. Starship ne réduit pas le coût du kilo en orbite de 20%, il le divise par 100. Et chaque fois dans l'histoire qu'un coût d'infrastructure est divisé par 100, ce n'est pas le marché existant qui grossit, ce sont des industries entières qui naissent. Le coût du calcul divisé par 100 a donné Internet, le smartphone, l'IA. Le coût de l'orbite divisé par 100 va donner une économie spatiale complète.
Faisons la liste de ce qui devient rentable quand le kilo en orbite coûte le prix d'un billet d'avion. Les data centers orbitaux, avec énergie solaire continue et refroidissement gratuit, au moment exact où l'IA fait exploser la demande énergétique terrestre. La fabrication en microgravité de semi-conducteurs, de fibres optiques, d'organes imprimés impossibles à produire sous gravité. Le tourisme orbital de masse, puis les hôtels lunaires, qui passeront du fantasme au business plan exactement comme la croisière de luxe au 20ème siècle. Le transport point à point terrestre, Paris-Tokyo en 40 minutes. L'industrie minière des astéroïdes, dont un seul corps de classe M contient plus de métaux que tout ce que l'humanité a extrait depuis le néolithique. Et Mars en ligne de mire, pas comme destination touristique, mais comme le plus grand projet d'infrastructure jamais entrepris, avec tout ce que ça implique de demande en énergie, matériaux, robotique, IA.
SpaceX ne participera pas à ces marchés. SpaceX possède le péage d'entrée de tous ces marchés. C'est AWS, mais pour la civilisation. Apple vaut 3 500 milliards en vendant des rectangles de verre sur une seule planète. Le premier monopole d'accès à une frontière infinie à 30 ou 50 trillions dans 5 ans, ce n'est pas de l'exubérance, c'est une simple règle de trois sur l'expansion du marché adressable.
Et maintenant, la partie que je préfère. Ce futur n'a pas besoin de bureaucrates. Il n'y a pas de comité consultatif en orbite. Pas de commission Théodule sur Mars. Chaque dollar de cette nouvelle économie sera créé par des ingénieurs, des techniciens, des soudeurs, des pilotes, des entrepreneurs. Les diplômés en gestion de la norme vont devoir apprendre un métier utile, et franchement, c'est une excellente nouvelle pour eux aussi : construire est infiniment plus fun que contrôler.
Parce que c'est ça, le vrai signal d'aujourd'hui. Pendant 50 ans on nous a vendu un futur rétréci : moins d'énergie, moins d'enfants, moins d'ambition, gérer le déclin proprement. Et là, d'un coup, le plus gros actif financier du monde est un pari sur l'abondance, l'expansion et l'aventure. Le pessimisme vient de passer en position vendeuse sur lui-même.
Le futur sera méga fun. Il y aura des hôtels avec vue sur la Terre, des honeymoons en orbite, des gamins qui diront "papa, c'était comment avant les fusées réutilisables" comme on dit "c'était comment avant Internet". Et quelque part dans les années 2030, un humain marchera sur Mars en livestream devant 5 milliards de personnes, et ce jour-là plus personne ne se souviendra du nom d'un seul de ses détracteurs.
Achetez de l'optimisme. C'est encore sous-valorisé.
Check out our new survey — we trace 10 years of biomedical image segmentation, from task-specific experts to one universal promptable foundation model. That roadmap figure says it all 👇
😋From WAM to WPAM: World-Action Models should NOT stop at pixels!
We release PointAction, lifting World Models from RGB to RGB+XYZ and using dynamic pointmaps as universal action representations for robot control.
Page: https://t.co/dDItqZwYfm
Paper: https://t.co/EB6BXbiqvm
🧐Q: Why not pixels only?
Pixels tell us what changes, but not always how a robot should move in 3D. Learning this mapping from RGB alone often requires massive paired action data, while raw motor commands are embodiment-specific and less transferable across robots.
Our intuition: World-Action Models should model the physical world in the same space where actions take effect — and that space is 3D.
Instead of predicting raw motor commands directly from video, PointAction uses 3D point dynamics as a richer and more robust bridge 🌉:
- they make metric motion, spatial constraints, and contact-relevant geometry explicit;
- they are less tied to a specific robot’s motors;
- they can be extracted from much broader robot video data.
PointAction first learns a general diffusion-based 4D world-action backbone in RGB+XYZ space, predicting robot-centric 3D point dynamics, then decodes them into embodiment-specific controls with lightweight action heads.
---
This project is led by @TongMutianTMT (incoming PhD student at @PennCIS@GRASPlab), and @hanjiang00 (talented undergrad who visited my lab last year). Huge congrats to all coauthors @WindStyle1459 and @LingjieLiu1! 🎉
🚀 Excited to share PointAction:
A new Video-Point-Action Model that uses dynamic 3D pointmaps as a universal, geometry-grounded action representation for robot control.
VLA → VAM → ? Lift RGB to RGB+XYZ, then decode robot-specific actions.
https://t.co/z6sB8aUQaS
[1/6]
"Latent Reasoning with Normalizing Flows"
NF-CoT makes latent reasoning feel native to LLMs. So instead of forcing every intermediate thought through verbose CoT text, it learns compact continuous thoughts with a normalizing flow inside the causal LLM stream.
The key move is that latent thoughts become sampleable, scoreable, and RL-trainable like tokens, with exact likelihoods and KV-cache friendly decoding.
This beats explicit CoT and prior latent methods, while using 64 latent tokens to compress roughly 385 CoT tokens and running much faster than diffusion-based latent reasoning.
Did you know? For many native signers, written text is actually a barrier to accessing information.
That’s why our successful launch in Singapore's MRT is just the start. We’re on a mission to bring sign language everywhere, ensuring the world's 70 million people with hearing loss are finally included. Let's make the unheard, heard.
#TechForGood #Inclusion
@thoma_gu Latent reasoning is what we need! NF-CoT brings normalizing flows into LLM's causal stream, boosting Qwen3-8B-Base by +13% ⬆️on code generation while running ~2× ⚡️faster than diffusion-based method. Hugely proud of our team! 🚀
🏆 Think your AI agent can do basic medical research end-to-end? Submit to our leaderboard (https://t.co/tXO9ioUPiN) and find out!
We're launching #AutoMedBench, the first benchmark for evaluating Medical #AutoResearch agents across the entire research workflow—not just the final answer.
📄 Paper: https://t.co/X2G66j2Qx6
🌐 Project: https://t.co/bYKEVcfqEU
💻 Code: https://t.co/z37BfBBNm5
📝 Plan → ⚙️ Setup → 🔍 Validate → 🚀 Inference → 📦 Submit
As AI agents move from answering medical questions to conducting end-to-end medical AI research, we need to measure where they succeed—and where they break.
What we benchmark:
• 24 tasks across segmentation, image enhancement, VQA, report generation, and lesion detection
• 48 task-tier combinations spanning Lite and Standard settings
• 6 frontier AI agents under a unified interface
• Thousands of runs with detailed logs of stage performance, costs, tokens, wall time, and failure modes
📊 Current leaderboard:
🥇 #Opus 4.6: 66.5
🥈 #GLM-5: 61.6
🥉 #Gemini 3.1 Pro: 59.0
4️⃣ #ChatGPT-5.4: 55.3
5️⃣ #MiniMax-M2.5: 51.6
6️⃣ #Qwen3.5: 51.2
🔎 Key findings:
⚠️ Agents are better at completing workflows than producing high-quality scientific outputs.
⚠️ Validation is the weakest stage; Setup is the strongest.
⚠️ More scaffolding is not always better—some frontier agents actually perform worse with additional guidance.
⚠️ The dominant failures are verification and submission, not task understanding.
💡 Takeaway:
The next frontier for research agents isn't just more medical knowledge—it's better workflow control, validation, error recovery, and artifact-level reasoning.
#AIAgents #MedicalAI #AgenticAI #LLM #MultimodalAI #HealthcareAI #Benchmark
🤔Can LLMs reason by sampling continuous thoughts — not just tokens?
Introducing NF-CoT: Latent Reasoning with Normalizing Flows. It samples continuous chain-of-thoughts directly in the stream of LLM with exact likelihood -- powered by STARFlow.
🌐Page: https://t.co/yTJ8vATfJZ
Language Models Need Sleep
"Transformer-based large language models are increasingly used for long-horizon tasks; however, their attention mechanism scales poorly with context length. To handle this, we study a sleep-like consolidation mechanism in which a model periodically converts recent context into persistent fast weights before clearing its key-value cache."
"increasing sleep duration N for our models improves performance, with the largest gains on examples that require deeper reasoning."
Can fast generative models still be likelihood-based?
Excited to share our new work @Apple MLR --Normalizing Trajectory Models
a step toward high-quality few-step generation with exact trajectory likelihood, powered by normalizing flows.
Paper: https://t.co/4VjJZpW4pC
[1/9]
As a big fan of Claude Code, I tried Codex even on a free account, and asked it to fix a bug on my personal website. It gave me a screenshot after it was done — god, compared with how lazy CC has been recently, I'm totally convinced by Codex now...
People talk, listen, watch, think, and collaborate at the same time, in real time. We've designed an AI that works with people the same way.
We share our approach, early results, and a quick look at our model in action.
https://t.co/AFJZ5kH7Ku
Excited to share STARFlow2 from Apple MLR :
🥨Bridging Language Models and Normalizing Flows for Unified Multimodal Generation.
One model to understand, reason, and generate continuous images with a single unified autoregressive mechanism?
Paper: https://t.co/IA1pJ5AtOX
1/9
Codex grew programmatic policies with no neural nets: max score on Breakout, and SOTA-level scores on MuJoCo.
Maybe heuristics were not too weak. Maybe they were just too expensive to maintain. Maybe it's the next paradigm.
https://t.co/1ZaIneleuW