Beautiful paper from Google DeepMind.
Explains the pathways from AGI to ASI, and why that jump could happen through several routes.
The authors frame the AGI-to-ASI transition around 4 technical pathways:
- continued scaling of compute, model size, data, and test-time inference;
- algorithmic paradigm shifts beyond today’s transformer-based foundation-model stack;
- recursive self-improvement, where AI accelerates AI R&D and improves future systems; and
- multi-agent collective intelligence, where large populations of specialized agents coordinate into a superhuman group agent.
Scaling may work for a while, but it could hit limits in data, compute, energy, or weaker returns from making systems larger.
Recursive improvement is the most uncertain path, because AI could speed up AI research, but that loop may also slow if hard research problems need real-world testing, scarce hardware, or new ideas.
Multi-agent collectives may be the most underappreciated path, because a society of competent digital workers could outperform a brilliant individual model through specialization, speed, and coordination.
The big point is that ASI may not arrive as 1 sudden event, but as a chain of faster changes as AI helps create better AI and stronger scientific tools.
----
Link – arxiv. org/abs/2606.12683
Title: "From AGI to ASI"
What stood out most:
Errors were not random hallucinations; they were clinically plausible differentials.
The path forward looks hybrid: local models for scale, frontier models for complex cases.
Proud to be a co-first author on this work. Our team at @UHN@UHNAIHUB asked a simple question: can on-device models actually work in clinical settings? Here's what we found 🧵
Just updated our paper on on-device LLMs for clinical decision support.
Paper: https://t.co/c4O4KQFocz
Here's why I think this matters:
We've been asking the wrong question. The debate around LLMs in medicine has been "how accurate are
they?", but the harder problem is deployment. Patient data can't leave the hospital. Most clinics don't have the bandwidth or budget for cloud inference at scale. The real question is: can a model that runs locally, on modest hardware, actually be trusted for clinical decisions?
After benchmarking 188 models across general disease diagnosis, ophthalmology, and clinical judgment simulation — the answer is yes.
Gemma 4 31B (@googlegemma) hits 86.5% on general diagnosis, beats GPT-5-mini, scores 100% on uroradiology and breast imaging, and runs at 18 GB. Qwen3.5-27B (@Alibaba_Qwen ) at 16 GB matches DeepSeek-R1 at 671B, that is one-twenty-third the memory, same clinical accuracy. Fine-tune Qwen3.5-35B with domain-specific reasoning traces and it reaches 87.9%, approaching GPT-5.1 (89.4%). No extra memory. No cloud call. No PHI leaving the building.
One thing that surprised me: 87.2% of errors across all models were clinically plausible differentials. the model picked a reasonable diagnosis, just not the right one. Above ~31B parameters, hallucination rate drops to zero. Errors start looking like the kind a careful clinician makes on a hard case, not the kind that would make you distrust the system.
There's also a pass@3 upper bound of 93.2% for fine-tuned Qwen3.5-35B. The model already "knows" the right answer in most cases. That's a verifier problem, not a model-size problem.
Gemma 4 and Qwen3.5 are the first generation where the local deployment story actually holds up under rigorous clinical benchmarking. That's a real milestone.
Huge shoutout to the team who made this happen: Alif Munim (@alifmunim ), Omar Ibrahim, Alhusain Abdalla, Jun Ma @JunMa_AI4Health (all equal contributors), Meng Wei, Shuolin Yin, and Leo Chen from @UHN AI hub. Proud of what this group built 🔥🔥
We tested how far local/on-device models can go.
Gemma 4 31B reached 86.5% on general diagnosis and outperformed GPT-5-mini at ~18 GB.
Qwen3.5-27B (~16 GB) matched DeepSeek-R1, a 671B model, at a fraction of the memory.
That efficiency gap matters for real deployment.
We have fixed a major inference bug in https://t.co/QJ8km69UXC, significantly improving the quality of reasoning
Give BioReason-Pro another try! And please keep the feedback coming
You can also find a guide on setting up the model locally at https://t.co/SBINSVnu7d
Sakana AI's AI Scientist just landed in @Nature. Not a tool that helps you write papers. A system that does science — generates ideas, writes code, runs experiments, drafts the manuscript. And v2 already passed human peer review. The new finding: there's a scaling law for AI-generated science itself. Better models → better papers. Automatically. We've had scaling laws for language, reasoning, and coding. Now apparently science too.
Wild time to be a researcher.
We @arcinstitute, @UHN, and @VectorInst recently released out BioReason-Pro, a multimodal reasoning LLM for protein function prediction, trained via SFT on synthetic reasoning traces and subsequent RL.
I had a chance to interview @BoWang87 and @genophoria on their vision for the work and what comes next. Was fun to pick their brains on the bio!
Check out the interview: https://t.co/YWslLYKFxf
I just used BioReason-Pro on a gene I am subcloning and was quite impressed. The processing time is reasonable, and the results appear accurate. That said, the functional summary could be expanded to provide more depth and context. In addition, the GO-GPT predictions section would benefit from clearer guidance and more informative explanations.
Still, amazing work! Congratulations to @BoWang87@genophoria@arcinstitute. I plan to use more in my future research.
Big day — BioReason-Pro is live.
Proud to share this as co-first author on a joint release from @arcinstitute and @UHN.
The first multimodal reasoning LLM for protein function prediction. It reasons step-by-step like a biologist — not just classifies.
2026 may be the year AI starts to truly reason about biology.
AlphaFold helped close the sequence → structure gap.
The next frontier is sequence → functions.
Today, together with @genophoria and the team at @arcinstitute , we’re releasing BioReason-Pro — the first multimodal reasoning model for protein function prediction.
Big day — BioReason-Pro is live.
Proud to share this as co-first author on a joint release from @ArcInstitute and @UHN, with a team across 13 institutions.
🌐 https://t.co/yfHiF5jPia
📄 https://t.co/T5lk5CMRZZ
BioReason-Pro is the first multimodal reasoning LLM for protein function prediction.
It doesn’t just classify proteins — it reasons step by step like a biologist.
Excited to see this work shared. A key takeaway: smaller models are becoming strong enough for real clinical AI, even in on-device settings. Thanks to @BoWang87 for the guidance, the amazing team for the collaboration, and @Alibaba_Qwen for the great models.
Two major AI releases this week:
• Qwen3.5 — new open-source small models
• GPT-5.4 — newest frontier closed model
Most benchmarks compare math and coding.
But the real test for frontier AI should be biology and healthcare.
That’s where mistakes actually matter.
So our team at @UHN ran them on EURORAD — 207 expert-validated radiology differential diagnosis cases.
Results:
GPT-5.4: 92.2%
Qwen3.5-27B: 85%
Gemini 3.1 Pro: ~79%
A 27B open model that runs on a laptop is only 7 points behind the most powerful AI model on earth — and already beating Gemini on this benchmark.
That gap is much smaller than people expected.
And it matters.
For years hospitals faced an impossible tradeoff:
Frontier models → patient data leaves the hospital
Local models → not good enough
That tradeoff may finally be ending.
Qwen3.5-27B runs fully local.
No API. No cloud. No patient data leaving the building.
HIPAA / PHIPA compliance becomes architecture, not paperwork.
Interesting detail: 27B and 122B score almost identically here.
Scaling bigger didn’t help much.
One caveat: with web-scale training, it’s hard to completely rule out that frontier models like GPT-5.4 may have seen parts of evaluation datasets.
Still, the signal is clear:
Small models are getting good enough for real clinical AI.
And if we want to measure real AI progress,
biology and healthcare should be the benchmark.
Huge credit to the team
@alifmunim@AlhusainAbdalla@JunMa_AI4Health@Omar_Ibr12@oliviaamwei