AI's clinical safety is not the same as AI accuracy.
NOHARM shows LLMs can score well on knowledge tests yet still give harmful clinical advice, mostly by leaving out key actions.
If you evaluate AI in medicine, measure safety as its own metric. #MedEd
👀 Big Bang Santé @Le_Figaro
💓Le coup de cœur du Figaro se porte sur le Dr Kevin Yauy, lauréat 2025 de l’initiative Next Gen Leader dans la catégorie Medtech.
Il est Genéticien au @CHU_Montpellier#BigBangSanté#Santé#senior
👀 Big Bang Santé @Le_Figaro
"L'IA nous permet de mieux former nos étudiants sur toutes les maladies d’aujourd’hui et de demain.
Nous allons construire ensemble le meilleur soin pour nos futurs patients"
💓Le coup de cœur du Figaro
Dr Kevin Yauy, lauréat 2025 de l’initiative Next Gen Leader dans la catégorie Medtech @CHU_Montpellier
#BigBangSanté #Santé #senior #innovation
This paper really is groundbreaking. It solves a long-standing embarrassment in machine learning: despite all the hype around deep learning, traditional tree-based methods (XGBoost, CatBoost, random forests, etc) have dominated tabular data—the most common data format in real-world applications—for two decades. Deep learning conquered images, text, and games, but spreadsheets remained stubbornly resistant.
This paper's (published in Nature by the way) main contribution is a foundation model that finally beats tree-based methods convincingly on small-to-medium datasets, and does so very fast. TabPFN in 2.8 seconds outperforms CatBoost tuned for 4 hours—a 5,000× speedup. That's not incremental; it's a different regime entirely.
The training approach is also fundamentally different. GPT trains on internet text; CLIP trains on image-caption pairs. TabPFN trains on entirely synthetic data—over 100 million artificial datasets generated from causal graphs.
TabPFN generates training data by randomly constructing directed acyclic graphs where each edge applies a random transformation (using neural networks, decision trees, discretization, or noise), then pushes random noise through the root nodes and lets it propagate through the graph—the intermediate values at various nodes become features, one becomes the target, and post-processing adds realistic messiness like missing values and outliers. By training on millions of these synthetic datasets with very different structures, the model learns general prediction strategies without ever seeing real data.
The inference mechanism is also unusual. Rather than finetuning or prompting, TabPFN performs both "training" and prediction in a single forward pass. You feed it your labeled training data and unlabeled test points together, and it outputs predictions immediately. There's no gradient descent at inference time—the model has learned how to learn from examples during pretraining.
The architecture respects tabular structure with two-way attention (across features within a row, then across samples within a column), unlike standard transformers that treat everything as a flat sequence.
So, the transformer has basically learned to do supervised learning.
Talk to the paper on ChapterPal: https://t.co/hmWIA1dYji
Download the PDF: https://t.co/uxElyS85ge
Today in @Nature we report a new prime editing strategy that can rescue a common cause of many genetic diseases in a disease-agnostic manner. This approach converts a redundant endogenous tRNA into an optimized suppressor tRNA, enabling a single prime edit to rescue premature stop codons across different diseases.
(1/15)
https://t.co/zs0qu5bhXx
"Depuis 2019, on peut séquencer l'entièreté du génome humain", explique @KevinYauy. Et les progrès de l'IA devraient nous permettre aux chercheurs et soignants de l'exploiter pleinement. #FuturapolisSante#Sante#ADN#IA
After years of rejections, I finally published in one of the top journals in our field.
In my latest post, I share every detail: how I turned repeated "not a good fit" emails into a clean RCT, built on theory.
It’s the story behind our paper on ChatGPT warnings. #MedEd
GPT-5 Pro just disproved a conjecture on the list of the Simon's Foundation's "Open Problems" list. Very impressive. I wonder where the goal posts will move now? CC @sebastienbubeck HT @arjunmanrai
Our Spanish colleagues at @Sedem_ORG always come up with fruitful webinars.
This time they hosted @KevinYauy, who is a brilliant researcher in #MedEd, for a demonstration of his AI-based simulation with feedback.
It's worth investing 36 minutes (or 18 mins at 2x speed).
Are language models overconfident in their clinical reasoning? Have reasoning optimizations made this problem worse?
Our latest in NEJM AI - we adapt a human-validated, trivially scalable automated benchmark to get to the heart of clinical reasoning in AI systems
New AI numbers from Pew, fresh this morning. 62% of U.S. adults now say they interact with AI at least several times a week, and 31% are now using AI 'almost constantly' or several times a day.
📣 Futurapolis Santé, c’est la promesse de rencontrer et échanger avec celles et ceux qui inventent la santé de demain. Des exemples ?
Nous vous présentons @KevinYauy, médecin généticien et docteur en intelligence artificielle, qui parlera de l'"ADN poubelle" !
Autrement dit cet ADN qui ne semble rien coder… quoi que…
RDV les 10 & 11 octobre pour en savoir plus
#Adn #Génétique #Santé #FuturapolisSanté #Innovation
@CHU_Montpellier
New in NEJM AI: GPT-4 powered ASCE training outperforms traditional OSCE prep — higher scores, less stress, more realism for med students.
Link in the comments 👇
#NEJMAI#MedEd#AIinHealthcare