Introducing Kimi K3: Open Frontier Intelligence
🔹 2.8 Trillion Parameters, 1 Million Context, Native Multimodal
🔹 Kimi Delta Attention enables up to 6.3x faster decoding in million-token contexts
🔹 Attention Residuals deliver ~25% higher training efficiency at <2% additional cost
🔹 Built for long-horizon agentic coding and self-evolving workflows
Kimi K3 is now live on on https://t.co/zrk6zZxZUo, Kimi Work, Kimi Code, and the Kimi API.
Open Weights by July 27, 2026.
🔗 API: https://t.co/XCrgjXAqMw
🔗 Tech blog: https://t.co/YTfiMSNM1f
If you use LLM-as-judge, this one is worth reading.
(bookmark it)
It's actually one of the most effective ways to use LLM-as-a-Judge for evals.
Holistic judge scores hide both their reasoning and their ceiling effects.
BINEVAL decomposes each evaluation criterion into atomic yes-or-no questions, answers each independently per output, then aggregates the verdicts into calibrated multi-dimensional scores.
Every question-level verdict is inspectable, so you can diagnose exactly why an output scored low, and the same verdicts feed straight back as targeted prompt-improvement signal.
Across SummEval, Topical-Chat, and QAGS, it matches or beats UniEval and G-Eval, training-free, with especially strong results on factual consistency.
Paper: https://t.co/oar6BZcasm
Learn to build effective AI agents in our academy: https://t.co/1e8RZKs4uX
Welcome to Gemini 3.5 Flash, our most powerful model to date. It pushes the frontier of intelligence, speed, and cost putting 3.5 Flash in a class of its own.
We spent the last 6 months making sure Flash is great for real world use cases. It's available everywhere now!
Introducing TurboQuant: Our new compression algorithm that reduces LLM key-value cache memory by at least 6x and delivers up to 8x speedup, all with zero accuracy loss, redefining AI efficiency. Read the blog to learn how it achieves these results: https://t.co/CDSQ8HpZoc
Demis Hassabis just defined the real test for AGI. It’s more brutal than anyone expected.
Train AI on all human knowledge. Cut it off at 1911. See if it independently discovers general relativity like Einstein did in 1915.
If it can, we have AGI. If not, we’re still building pattern matchers.
Hassabis: “My definition of AGI has never changed. A system that can exhibit all the cognitive capabilities that humans can.”
Not bar exams. Not coding competitions. All cognitive capabilities.
Hassabis: “The brain is the only existence proof we have, maybe in the universe, of a general intelligence.”
That’s why DeepMind studies neuroscience. Not for inspiration. For data. The human brain is the only confirmed evidence that general intelligence is physically possible.
If you want to build it, you study the only example that exists.
Hassabis: “True creativity, continual learning, long-term planning. They’re not good at those things.”
Current systems are impressive and broken simultaneously.
Hassabis: “They can get gold medals in international math olympiad questions, but they can still fall over on relatively simple math problems if you pose it in a certain way.”
Jagged intelligence. Brilliant in narrow domains. Incompetent when approached differently.
That inconsistency is the tell. A true general intelligence doesn’t spike in one direction and collapse in another.
The Einstein test cuts through all of it. No benchmarks. No leaderboards. No carefully curated evals.
Just a model, a knowledge cutoff, and the question of whether it can do what one human did alone in 1915.
Hassabis: “Training an AI system with a knowledge cutoff of 1911 and seeing if it could come up with general relativity like Einstein did in 1915. That’s the true test of whether we have a full AGI system.”
Current models can’t. They remix brilliantly. They don’t generate paradigm-shifting theories from first principles.
Hassabis: “I think we’re still a few years away from that.”
A few years. Not decades.
The system that can be Einstein once can be Einstein a thousand times simultaneously across every domain.
That’s not AGI anymore. That’s the beginning of something we don’t have words for yet.
When that test gets passed, we won’t need a press release to know what happened.
Claude in PowerPoint is now available on the Pro plan.
It also now supports connectors, bringing context from your daily tools directly into your slides.
Try it here: https://t.co/N52VtGwcPj
Today, we’re introducing Pomelli’s latest feature update, ‘Photoshoot’
With Photoshoot, you can start from a single image of your product and easily create high quality, customized product shots to elevate your marketing.
Available free of charge in the US, Canada, Australia & New Zealand! Get started with Pomelli today at https://t.co/SbeT00ToNx
A good internet identity is especially important for career entry, as there are not many criteria of evaluation for employers besides the degree. It is an important tool to consciously highlight strengths. #makingsocialmedia
@jpschroeder@grok wouldn’t it be way smarter of anthropic to gain significant market share with publishing their best model and being the Sota ahead of every other ai company instead of focusing on short term optimization of their margin?