The first experimental evidence of recursive self-improvement (RSI).
Autoresearching the autoresearch agent for eight days.
The result beats the harness we hand-tuned for two years, on held-out benchmarks: 🧵(1/7)
@jessegenet do you know of an open repo for preK–12 education?
I’m imagining a shared Notion where parents can contribute curriculum, lesson plans, printables, prompts, etc.
Open-source, parent-driven educational infrastructure seems very buildable now.
@tryheidi pls debug; the following is generated at the end of encounter notes and is decreasing the value of this AI scribe. Thank you! "AI: I've structured the note according to the template, capturing the key information....."
Evaluations are essential to understanding how models perform in health settings.
HealthBench is a new evaluation benchmark, developed with input from 250+ physicians from around the world, now available in our GitHub repository.
https://t.co/s7tUTUu5d3
Real-world impact of physical activity reward-driven digital app use on cardiometabolic and cardiovascular disease incidence
App-users have significantly lower risk ofcardiovascular disease, stroke and type 2 diabetes (HR compared to non-app users.
https://t.co/76d6uWk7W8
Nvidia just open sourced Parakeet TDT 0.6B - the BEST Speech Recognition model on Open ASR Leaderboard 🔥
Can transcribe 60 minutes of audio in 1 second 🤯
600M parameters, with CC-BY-4.0 license (commercially permissive)
Congrats Nvidia on the brilliant release and beating all major closed source giants too! 🤗
Gemini powers our multimodal health research! 💙
In our new paper on multimodal AMIE, we're pushing conversational diagnostic AI beyond text to handle images such as skin photos, ECGs, and clinical docs, which provide crucial context in healthcare.
Blog: https://t.co/VAlKoR53Il
Paper: https://t.co/2zHQT0H5Pv
How do we make an AI reason like a clinician during a dynamic, multimodal conversation? One of our key contributions is multimodal state-aware reasoning, built on @GoogleDeepMind Gemini 2.0 Flash.
Instead of just reacting turn-by-turn, AMIE maintains an internal "understanding" of the consultation:
✅ What is known about the patient?
✅ What are the likely diagnoses?
✅ What information (text or visual) is missing?
This internal state allows AMIE to:
👉 Intelligently guide the conversation through phases like history-taking & diagnosis.
👉 Strategically ask for relevant images (like skin photos or screenshots of ECGs/docs) when its internal state shows uncertainty.
👉 Accurately interpret multimodal data and weave the findings back into the ongoing dialogue and diagnostic process.
Essentially, it mimics the adaptive reasoning clinicians use, leading to a more structured and effective consultation.
We evaluated multimodal AMIE against primary care physicians (PCPs) in a demanding, blinded OSCE study using 105 diverse multimodal scenarios.
The results demonstrate clear progress: AMIE achieved similar or superior performance when compared to PCPs across a wide range of metrics, including diagnostic accuracy, empathy, and critically, the handling and reasoning about multimodal data.
While the OSCE results are very promising, it's important to remember this was a test environment with patient actors! Real-world care is more complex. Making sure it's safe, reliable, and actually helpful in the real world needs more work, starting with our upcoming study with Harvard BIDMC.
The work would not have been possible without an amazing team @GoogleAI, @GoogleDeepMind: @RyutaroTanno, @alan_karthi, @vivnat, @AdamRodmanMD, @timstro, @taotu831, @hardyshakerman, @JanFreyberg, @_cjpark, @yasharmaa, @apalepu13, @arkitus, @weballergy, @valentinlievin, @ckbjimmy, @davidstutz92, @dgtbarrett, @yongcheng16@SaraM66905, @dr2w, @ymatias
@DeryaTR_ After comparing Gemini 2.5 and @EvidenceOpen in parallel for clinical questions and whisper/GPT API with marketed ambient scribes (eg @TaliAICompany) I really don't see much difference. What am I paying for under the hood of these wrappers?
@TomkeyKong@tryheidi@EricTopol@vivnat@apalepu13@labenz So when is
@tryheidi
going to make an EHR?
@OSCAREMRInc
is open-source w UX from the 90s and has about 20% of market ($200MM+ revenue for prim care in Canada). Can you leverage tech forward brand to capture some of that? Instead of whisper wrapper, tech forward oscar wrapper
Thank you @TomkeyKong for improving pt care w @tryheidi. Will big tech leap frog over you? Scribe is nonproprietary; value add is market capture, reg compliance, etc. Asking as a consultant for AI implementation in health systems. @EricTopol@vivnat@apalepu13@labenz
Make sure to check out and stay current on the latest in AI & machine learning in family medicine! This living collection pulls new articles from top journals as soon as they’re published. #FamilyMedicine#ArtificialInteligence
🔗- https://t.co/wt6d2zFAP3
Our latest paper on virtual care: patients of FPs who provided more virtual care were NOT more likely to go to the ED. @tara_kiran@RickGlazier1 https://t.co/cNfI0jdPMW via @JAMANetworkOpen part of @JAMANetwork
Approach to elevated aPTT/PT taught to me by Dr Pruthi. Inspired by @KirtanPatolia’s recent tweetorial.
1. Bleeding hx is key especially if recent bleeding challenge. ISTH bleeding score helpful
2. Rule out artifacts, repeat aPTT/PT
3. Workup based on which test abnormal
@jonesmedcardiff Why is GP scapegoated for NHS unsustainability? What percent of NHS expenditure does primary care represent? Single digit? Why is GP an easy target?
Hyaluronic Acid Knee Injections Equivalent to Placebo
In the United States, where over $300 million is spent annually on intraarticular hyaluronic acid injections, yet another study shows such therapy to be no better than placebo.
https://t.co/bBEwpId84L