I solved 6 open Erdős problems in 5 days, using @OpenAI GPT-5.6 Sol.
I have a math background, but the Codex workflow I used does not require deep mathematical knowledge.
Here’s exactly how I approached it, including my prompts 🧵
Introducing Kimi K3: Open Frontier Intelligence
🔹 2.8 Trillion Parameters, 1 Million Context, Native Multimodal
🔹 Kimi Delta Attention enables up to 6.3x faster decoding in million-token contexts
🔹 Attention Residuals deliver ~25% higher training efficiency at <2% additional cost
🔹 Built for long-horizon agentic coding and self-evolving workflows
Kimi K3 is now live on on https://t.co/zrk6zZxZUo, Kimi Work, Kimi Code, and the Kimi API.
Open Weights by July 27, 2026.
🔗 API: https://t.co/XCrgjXAqMw
🔗 Tech blog: https://t.co/YTfiMSNM1f
Imagine a web design team at your beck and call. Now imagine that team already knows you from hundreds of working sessions over the years. Your career, your voice, your risk tolerance, what you refuse to sound like. A human agency bills weeks for discovery. This team finished discovery before the project started. That is why the site sounds like me and not like a template. The model that knows you builds differently than the model that just met you.
🔥Today, we are releasing one of the first visual reasoning benchmarks for autonomous AI diagnosis in healthcare!
🚀Introducing Radiology’s Last Exam 2.0 (RadLE 2.0) from @CRASHLabAI, an uncertainty-aware benchmark for autonomous diagnosis in radiology!
✅In the last few days, the AI frontier has moved significantly. @OpenAI launched GPT-5.6 Sol. @Meta launched Muse Spark 1.1. @xAI dropped Grok 4.5.
🙌We’ve benchmarked all frontier, open-source and medical VLMs in RadLE2.0 and the leaderboard is now LIVE!
🚨 Before AI models are handed autonomy, one question matters more than any accuracy score: Do they know when to STOP and hand over to a human?
⚠️ A confident wrong diagnosis is far more dangerous than an honest “I don’t know.” Yet most models are bad at admitting the latter!
🚀 We release five RadLE 2.0 Scores: Confidence Weighted, Reliability, Accuracy, Safety and Handover Readiness and we find that models from @OpenAI@AnthropicAI@MetaAI@GoogleDeepMind@xAI@nvidia@Alibaba_Qwen@MistralAI@MiniMax_AI all score very differently as they optimize for different metrics!
🚨But most importantly, NONE of the Models have been able to reach the average human expert baseline!
⚡️A thread on what we found and which models aced our metrics! Link to the leaderboard and technical report at the end of the thread!
Yesterday, we made GPT-5.6 Sol Ultra generally available. Today, we're sharing that it produced a proof of the 50-year-old Cycle Double Cover Conjecture using 64 subagents in just under one hour. We're sharing the prompt and proof below. We're excited to see what you all do with Ultra!
♥️ GPT-5.6 is a major step forward for health, both at the frontier and at cost.
These models push the frontier of performance per dollar, bringing the best health intelligence to all. The smallest variant, GPT-5.6 Luna, evaluated at the lowest reasoning effort, outperforms GPT-5.5 at the highest reasoning effort–despite costing 25x less. The largest variant, GPT-5.6 Sol, sets a new high bar at cost.
Another especially cool result: physicians found fewer flaws in GPT-5.6 responses than physician-written responses.
We collected diverse tasks that remain difficult for recent OpenAI models, across patient-facing and clinician-facing use cases. We asked speciality-matched physicians to write responses to these tasks with unlimited time and web access. We then asked other physicians to compare responses side-by-side, blinded to their source. Physicians were asked to comment on areas of improvement across five axes: accuracy, communication, completeness, instruction following, and health decision helpfulness. We then reported the fraction of responses across sources rated perfectly across all axes, across 20,000 total axis ratings. GPT-5.6 Sol appeared strongest, although all GPT-5.6 models performed significantly better than physicians.
BREAKING: ANTHROPIC HAS JUST RELEASED CONTENT ON HOW TO BUILD A COMPANY WITHOUT EMPLOYEES
- CEO: 1 person.
- employees: claude agents
it lasts 30 minutes and it's free
Bookmark this so you don't miss it.
@skdh There is no such thing on earth like an empty emergency unit in a hospital which is air-conditioned, I assume you must have dreamed it or was it on a different planet?