A VLM Can Read One Medical Screen.
A Medical AI Agent Must Finish 24 Screens in a Row.
๐ง๐ต๐ถ๐ ๐ถ๐ ๐บ๐ ๐ณ-๏ฟฝ๏ฟฝ๐๐ฒ๐ฝ ๐ฟ๐ผ๐ฎ๐ฑ๐บ๐ฎ๐ฝ ๐ณ๐ฟ๐ผ๐บ ๐ฆ๐ฐ๐ฟ๐ฒ๐ฒ๐ป-๐ฎ๐-๐ฆ๐๐ฎ๐๐ฒ ๐๐ผ ๐ฎ ๐ฝ๐ฟ๐ผ๐ฑ๐๐ฐ๐๐ถ๐ผ๐ป ๐ ๐ฒ๐ฑ๐ถ๐ฐ๐ฎ๐น ๐๐ ๐๐ด๐ฒ๐ป๐.
Real medical agent must move through many screens.
It must click, zoom, type, segment, check, and finish.
If it forgets one step, the whole task can fail.
My 7-Step roadmap based on this recent paper: https://t.co/v0GhbutrR0
ใ๐ฆ๐๐ฒ๐ฝ ๐ญ: Start with the Real Goal
โธ Tell the agent the full medical task.
โธ Define what โdoneโ means.
โธ Do not give a vague screen task.
โ Example: โOpen the CT scan, zoom to the liver, draw ROI, export statistics.โ
ใ๐ฆ๐๐ฒ๐ฝ ๐ฎ: Treat Each Screen as State
โธ Every screen tells the agent where it is.
โธ Track open panels, active tools, and patient fields.
โธ Do not treat screens as random images.
โ Example: Is the agent in view mode, zoom mode, or annotation mode?
ใ๐ฆ๐๐ฒ๐ฝ ๐ฏ: Ground the Screen with Tools
โธ Use OCR to read labels.
โธ Use object detection to find buttons.
โธ Use zoom or crop for tiny UI parts.
โ Example: Find the โExportโ button before clicking.
ใ๐ฆ๐๐ฒ๐ฝ ๐ฐ: Use an Actor Agent
โธ The Actor chooses the next action.
โธ One step at a time.
โธ Do not let it jump to the final answer.
โ Example: CLICK, SCROLL, ZOOM, TEXT, SEGMENT, or COMPLETE.
ใ๐ฆ๐๐ฒ๐ฝ ๐ฑ: Add a Critic Agent
โธ The Critic checks the Actorโs action.
โธ It asks: โDoes this action fit this screen?โ
โธ It blocks unsafe or early actions.
โ Example: Do not allow COMPLETE before the result is saved.
ใ๐ฆ๐๐ฒ๐ฝ ๐ฒ: Give the Agent Memory
โธ Short-term memory remembers the last step.
โธ Long-term memory remembers the whole path.
โธ Do not reset the agentโs brain on every screen.
โ Example: Remember that the zoom tool was already selected.
ใ๐ฆ๐๐ฒ๐ฝ ๐ณ: Build the Runtime Agent
โธ Use the Critic during testing and training.
โธ Keep only corrected successful paths.
โธ Deploy a faster agent that learned from feedback.
โ Example: The runtime agent acts faster, but still follows the learned checks.
๐ฃ๐น๐ฒ๐ฎ๐๐ฒ ๐ฅ๐ฒ๐บ๐ฒ๐บ๐ฏ๐ฒ๐ฟ:
Medical AI agents do not need only better vision.
They need state, tools, memory, and checks.
A screen reader is not a workflow finisher.
--
๐ฅ Skip the trial-and-error of building production AI agents.
Watch my free 30-min training + get 88 pages of production guides.
Join 46,000+ AI engineers, architects, and directors.
๐ https://t.co/t9xyLtyvG1
Present-day LLMs like ChatGPT and Claude can write poetry and solve difficult algebra problems with astounding speed and precision. Some researchers have referred to this phenomenon of AI acquiring startlingly human-like skills as "emergence." But not everyone agrees with this terminology.
A new paper by SFI's David Krakauer, Melanie Mitchell, and John Krakauer proposes a framework grounded in complexity science to clarify whether a system's capabilities are truly emergent.
https://t.co/yyENFiN4Om
650+ clinical AI models now run on your iPhone.
30-40x faster than CPU. Fully private. The PHI never leaves your device.
Today I shipped 410+ new MLX medical models: biomedical NER, disease, drug, and PII de-id. All open weights, Apache 2.0.
Same model, same entities. Watch:
Doctors 25% more likely to miss a diagnosis on a patient within 3 months after Ai is implemented. The results indicate Ai has an immediate negative effect on professionals skillsets vs prolonged reliance.
https://t.co/ciPB3DdR2J
Today @Nature published 2 new AI medical agentic models that take capabilities to a new level, from end-to-end care after presentation to the emergency department and longitudinally through 3 out-patient visits. There's a lot to unpack.
Summarized in a new Ground Truths post.
https://t.co/NRzuMe23DC
For medical information, general AI frontier models (Google, OpenAI, Anthropic) outperformed specialized @openevidence and @UpToDate as assessed by 12 US clinicians, randomized and blinded to which model and extensive testing/benchmarks. This was not anticipated. @NatureMedicine
https://t.co/KCH1ADfQWz
Today I'm publishing a new essay, Policy on the AI Exponential. AI is progressing extremely fastโmuch faster than the policy process was built to handle. The essay lays out where I think the technology is now, and the action needed to close the gap: https://t.co/Lh6PWae178
I must say - Iโve anecdotally noticed a MASSIVE improvement in @openevidence results on this over the past month or few - specifically, much better grounding to the clinical context without swinging the other way into subtle confirmation bias.
Feels like the reasoning focused on the clinical intent is much improved, without overoptimizing to an โagreeableโ reward function.
The burnout framing is real, but it stops short of the harder problem. Passive formats aren't just a design flaw you can fix by switching to better content. The passivity is structurally incentivized.
Here's what I mean: the ACCME's accreditation fee model is funded by complexity. More credit hour requirements, more compliance tracking, more consulting services to navigate standards. An organization financially dependent on that apparatus has no incentive to endorse formats that are actually efficient, because efficiency shrinks the billable surface area. Physicians sitting through lecture-based modules for 20 to 50 hours annually isn't an oversight. It's load-bearing infrastructure for a system generating over $5 billion in annual spending.
The pharma angle compounds this. When companies structure CME funding as unrestricted educational grants, they don't need to dictate content directly. The requirement that CME address "unmet medical needs" already channels education toward newer, expensive drugs. Passive formats are actually useful here because a physician watching a sponsored lecture absorbs framing without resistance the way a problem-based simulation wouldn't permit.
AI-driven personalized learning genuinely threatens this. Not because it's better pedagogy, though it is. Because it eliminates the accreditation and compliance layer that extracts rent from every credit hour. That's why the incumbent ecosystem is structurally motivated to resist it regardless of outcome evidence.
Calling this a burnout problem suggests the fix is ergonomic. The fix is architectural.
https://t.co/lFUC9Ena84
Anthropic vient de lancer des agents IA pour les avocats et juristes.
Couplรฉ avec le MCP Pappers, les rรฉsultats sont impressionnants.
On a testรฉ l'agent claim sur un contentieux fictif, gรฉnรฉrรฉ en 2 minutes...
Le rรฉsultat est disponible en dessous
HealthBench Professional is now live on Medical Sphere ๐ฅ
Built by @OpenAI, this benchmark contains 525 real clinician chat cases across care consults, medical documentation, and medical research.
https://t.co/iI2cy3oQqD
@DGlaucomflecken A word about Vinayโs statement about Dr Glaucomflecken not being a health policy expert. Do we need to be health policy experts to know that our patients are being denied. When was the last time he was a clinic writing appeal letters for patients?
@DGlaucomflecken Your brand of humor is gentle and intelligent. We doctors need your voice. You are our truth teller.
Any fair person knows that you have nothing to do with this senseless event .
So many podcasts to listen to this #NephMadness 2024! Great way to start to help you decide your picks! The PodCrawl https://t.co/RyROzzTEld via @AJKDonline
mRNA as a medicine in nephrology: the future is now!
๐What are the key strengths of RNA-based therapies?
๐What are the potential future directions & challenges of this promising technology in nephrology?
๐Find out more here๐๐๐
๐https://t.co/qqEhC4y6Pu