Excellent question. We are currently exploring longer-term agentic tasks in real-world medical scenarios. We believe that the model's medical knowledge reasoning and deep retrieval of evidence-based information form the foundation. Currently, in the medical field, what needs to be addressed more is how to define long-term medical tasks.
@TheAiGirl0@AntLingAGI This version of the AFUMED-Drug evaluation dataset covers medication dosage and contraindications, and we will continue to update the dataset to include more content.
Sante is now available on @OpenRouter and @vercel_dev. Healthcare professionals, researchers, and developers can try Sante free for one month through the OpenRouter API and explore its capabilities on real-world healthcare tasks.
Try it here:
OpenRouter: https://t.co/pY2P5lRIT2
Vercel: https://t.co/nMxvcRme75
A special thank-you to @novita_labs for supporting this launch.
This is our team's capability-balanced, size-friendly, and highly efficient new-generation medical health model, built on the foundation of Ling-3.0-Flash. It focuses on medical knowledge reasoning, clinical diagnosis alignment, medical safety ethics, deep evidence-based analysis, general reasoning, and agentic capabilities. We look forward to your usage and feedback, as we will iterate even faster!
Today, we’re introducing Ling-3.0-flash-Sante — an MoE model enhanced for health and medicine, built on Ling-3.0-flash.
Inspired by the French word “santé,” meaning “health,” Sante is built for real-world healthcare tasks spanning medical reasoning, professional healthcare tasks, deep research, and evidence-based retrieval.
Sante shows competitive results across MedXpertQA-Text, DiagnosisArena-MCQ, AFUMED-Drug, HealthBench Professional, and BrowseComp, with leading performance among open-source models and results competitive with flagship models.
Here’s to better health — Santé! More below ↓
Today, we’re introducing Ling-3.0-flash-Sante — an MoE model enhanced for health and medicine, built on Ling-3.0-flash.
Inspired by the French word “santé,” meaning “health,” Sante is built for real-world healthcare tasks spanning medical reasoning, professional healthcare tasks, deep research, and evidence-based retrieval.
Sante shows competitive results across MedXpertQA-Text, DiagnosisArena-MCQ, AFUMED-Drug, HealthBench Professional, and BrowseComp, with leading performance among open-source models and results competitive with flagship models.
Here’s to better health — Santé! More below ↓
Financial work depends on trustworthy sources, consistent definitions, accurate calculations and auditable outputs.
Introducing Ling-3.0-flash-Fin, a finance-enhanced version of Ling-3.0-flash, developed with financial institutions and domain experts.
With 124B total and 5.1B active parameters, it supports information retrieval, research, valuation modeling and report preparation across long reports, research materials and complex workbooks.
The model showed competitive results across FinFIRST, FinSearchComp Verified, FinCRAFT, FinanceAgent v1.1/v2, APEX-Agents, SpreadsheetBench v1/v2 and τ³-Banking.
We will open-source the model weights next week.
Proud to announce our first open-source medical LLM: AntAngelMed .
6.1B-active MoE ≈ 40B dense, 128K context, zh/en, 3-stage medical training.
Use it & tell us what to improve → https://t.co/fMcPhzRQJd
AntAngelMed 🏥 open medical LLM from Ant group @ant_oss
https://t.co/Kt9suwO6RY
✨ Efficient MoE: 6.1B active params ≈ 40B dense model performance
✨ 128K long context
✨ Chinese/English support
✨ Three-stage medical training for strong reasoning, safety, and clinical accuracy
Is a 0.5% performance gain worth paying 6x the price?
As a Medical AI researcher aiming to provide accessible health services globally, I analyzed the new GPT-5.2 release vs. Gemini 3 Pro. Here is my take on ROI and Model Capability:
1. The Efficiency Calculation
When we look at fundamental reasoning (GPQA, Math, Code):
• GPT-5.2 Thinking: 92.4%
• Gemini 3 Pro: 91.9%
The difference is a negligible 0.5%.
However, looking at the ARC-AGI-2 cost chart:
• Gemini 3 Pro costs ~$0.80 per task.
• To reach GPT-5.2's "High/X-High" performance zone, the cost skyrockets to $4.00 - $5.00+.
Conclusion: GPT-5.2 charges 5-6x more for a marginal <1% gain on GPQA. For high-concurrency medical consulting products, this cost gap determines product viability. Gemini 3 Pro is the current ROI King.
2. Insight: Where AGI is actually heading
Looking closely at the benchmarks, I found two fascinating trends:
• Math Competitions are "Solved": GPT-5.2 hit 100% on AIME 2025. Optimizing for this is now diminishing returns.
• Gemini's "Hidden" Strength: On FrontierMath (Tier 4)—which represents novel, complex research problems—Gemini 3 Pro (18.8%) actually outperforms GPT-5.2 (14.6%)!
This suggests that while GPT is stronger at "exams" with known patterns, Gemini may possess deeper reasoning potential for exploring unknown, complex problems—exactly what we need in medical diagnostics.
A year ago, we verified a preview of an unreleased version of @OpenAI o3 (High) that scored 88% on ARC-AGI-1 at est. $4.5k/task
Today, we’ve verified a new GPT-5.2 Pro (X-High) SOTA score of 90.5% at $11.64/task
This represents a ~390X efficiency improvement in one year
A year ago, we verified a preview of an unreleased version of @OpenAI o3 (High) that scored 88% on ARC-AGI-1 at est. $4.5k/task
Today, we’ve verified a new GPT-5.2 Pro (X-High) SOTA score of 90.5% at $11.64/task
This represents a ~390X efficiency improvement in one year
Awesome! From the HLE/BrowseComp results and 256K context, it looks like K2 did this carefully. I think three pillars matter:
(1) verifiable math for step-wise logic, (2) true long-context (train+eval beyond recall), (3) agent/task diversity with deep tool-use loops.
🚀 Hello, Kimi K2 Thinking!
The Open-Source Thinking Agent Model is here.
🔹 SOTA on HLE (44.9%) and BrowseComp (60.2%)
🔹 Executes up to 200 – 300 sequential tool calls without human interference
🔹 Excels in reasoning, agentic search, and coding
🔹 256K context window
Built as a thinking agent, K2 Thinking marks our latest efforts in test-time scaling — scaling both thinking tokens and tool-calling turns.
K2 Thinking is now live on https://t.co/YutVbwktG0 in chat mode, with full agentic mode coming soon. It is also accessible via API.
🔌 API is live: https://t.co/EOZkbOwCN4
🔗 Tech blog: https://t.co/n7xxaszqzF
🔗 Weights & code: https://t.co/4ukcXB0iP6
[M2M-03] Single-variable ablation result:
Adding 7k curated math competition problems into the reasoning RL mix (with medical reasoning items) improved 15 medical evals (knowledge + reasoning) by +0.66 avg (69.42→70.08).
👏Every leaderboard went up.
🥳Hard medical reasoning sets saw ≥+2.0 on average: NEJMQA, GPQA-Med, SuperGPQA-Med, DiagnosisArena.
Next I’ll detail the math-set construction. Core idea (inspired by DAPO/AIME): convert diverse math answers to integer targets to enable precise rule-based rewards and avoid parser errors.
Our pipeline = web/official sources → filtering → answer transformation to integers → final 7k set used for RL.
#MedAI #LLMs #RL
🚀We officially release Ring-1T, the open-source trillion-parameter thinking model built on the Ling 2.0 architecture.
Ring-1T achieves silver-level IMO reasoning through pure natural language reasoning.
→ 1 T total / 50 B active params · 128 K context window
→ Reinforced by Icepop RL + ASystem (Trillion-Scale RL Engine)
→ Open-source SOTA in natural language reasoning — AIME 25 / HMMT 25 / ARC-AGI-1 / CodeForce
Deep thinking · Open weights · FP8 version available
We're still training and pushing to do more. Stay tuned!
1/5