New CursorBench results just dropped and Grok 4.5 is the story.
#3 overall at 66.7%. Right behind Fable 5 Max at 70.5%.
Now look at the cost column.
Fable 5 Max: $17.32 per task
Grok 4.5 High: $1.51 per task
That is Fable level performance at roughly 1/10th the cost.
And it beats Fable 5 High and Opus 4.8 Max outright.
The intelligence war just became a price war.
Every SEO/GEO has the same advice for showing up in ChatGPT. Good content, listicles, post on Reddit.
I wanted to know if any of it actually works, so I spent a few days reading ChatGPT's raw network traffic instead of taking it on trust.
It tags every source as 1 of 4 pipelines, never touches the web on some questions, and on a pricing comparison it gave up on a vendor's JS loaded pricing page and cited a competitor instead.
Full writeup, with the traffic 👇
https://t.co/fKhdijlf9E
GLM 5.1 vs GLM 5.2
6 advanced HTML canvas challenges:
💧 Ink diffusing in water
⚔️ Energy-blade duel
📱 Slide to unlock
🅿️ 360° parking assist
🔥 Burning letter to ash
🏠 Build-a-house sequence
Pure canvas, zero libraries.
Introducing Composer 2.5, our most powerful model yet.
It's more intelligent, better at sustained work on long-running tasks, and more reliable at following complex instructions.
For the next week, we’re doubling the included usage of the model.
@Sardoche_Lol Quelques semaines, si tu veux des conseils en dev agentique ou création de ferme de développement par agent (full local avec API cloud branché sur un abonnement existant) hésite pas à me dm, j’ai beaucoup expérimenté et exploité depuis janvier c’est avec plaisir que je t’aiderai
Wow they did it 🔥
"Qwen3.5-35B-A3B now surpasses Qwen3-235B-A22B-2507"
So in 6 months they've trained a model which is:
- 6.7x smaller than the previous one
- Better in all benchmarks
- Available locally on a laptop
We're just at the very beginning of local LLMs and, at some point, we'll have an Opus 4.6 intelligence running on a phone.
Demis Hassabis just defined the real test for AGI. It’s more brutal than anyone expected.
Train AI on all human knowledge. Cut it off at 1911. See if it independently discovers general relativity like Einstein did in 1915.
If it can, we have AGI. If not, we’re still building pattern matchers.
Hassabis: “My definition of AGI has never changed. A system that can exhibit all the cognitive capabilities that humans can.”
Not bar exams. Not coding competitions. All cognitive capabilities.
Hassabis: “The brain is the only existence proof we have, maybe in the universe, of a general intelligence.”
That’s why DeepMind studies neuroscience. Not for inspiration. For data. The human brain is the only confirmed evidence that general intelligence is physically possible.
If you want to build it, you study the only example that exists.
Hassabis: “True creativity, continual learning, long-term planning. They’re not good at those things.”
Current systems are impressive and broken simultaneously.
Hassabis: “They can get gold medals in international math olympiad questions, but they can still fall over on relatively simple math problems if you pose it in a certain way.”
Jagged intelligence. Brilliant in narrow domains. Incompetent when approached differently.
That inconsistency is the tell. A true general intelligence doesn’t spike in one direction and collapse in another.
The Einstein test cuts through all of it. No benchmarks. No leaderboards. No carefully curated evals.
Just a model, a knowledge cutoff, and the question of whether it can do what one human did alone in 1915.
Hassabis: “Training an AI system with a knowledge cutoff of 1911 and seeing if it could come up with general relativity like Einstein did in 1915. That’s the true test of whether we have a full AGI system.”
Current models can’t. They remix brilliantly. They don’t generate paradigm-shifting theories from first principles.
Hassabis: “I think we’re still a few years away from that.”
A few years. Not decades.
The system that can be Einstein once can be Einstein a thousand times simultaneously across every domain.
That’s not AGI anymore. That’s the beginning of something we don’t have words for yet.
When that test gets passed, we won’t need a press release to know what happened.
Today, we’re introducing Pomelli’s latest feature update, ‘Photoshoot’
With Photoshoot, you can start from a single image of your product and easily create high quality, customized product shots to elevate your marketing.
Available free of charge in the US, Canada, Australia & New Zealand! Get started with Pomelli today at https://t.co/SbeT00ToNx
17,000 tokens per second!! Read that again!
LLM is hard-wired directly into silicon. no HBM, no liquid cooling, just raw specialized hardware. 10x faster and 20x cheaper than a B200.
the "waiting for the LLM to think" era is dead. Code generates at the speed of human thought.
Transition from brute-force GPU clusters to actual AI appliances.
https://t.co/Bf6DH7Q6Uf
I've spent 2.54 BILLION tokens perfecting OpenClaw.
The use cases I discovered have changed the way I live and work.
...and now I'm sharing them with the world.
Here are 21 use cases I use daily:
0:00 Intro
0:50 What is OpenClaw?
1:35 MD Files
2:14 Memory System
3:55 CRM System
7:19 Fathom Pipeline
9:18 Meeting to Action Items
10:46 Knowledge Base System
13:51 X Ingestion Pipeline
14:31 Business Advisory Council
16:13 Security Council
18:21 Social Media Tracking
19:18 Video Idea Pipeline
21:40 Daily Briefing Flow
22:23 Three Councils
22:57 Automation Schedule
24:15 Security Layers
26:09 Databases and Backups
28:00 Video/Image Gen
29:14 Self Updates
29:56 Usage & Cost Tracking
30:15 Prompt Engineering
31:15 Developer Infrastructure
32:06 Food Journal
"It was ready to kill someone, wasn't it?"
"Yes."
Daisy McGregor, UK policy chief at Anthropic, a top AI company, says it's "massively concerning" that Anthropic's Claude AI has shown in testing that it's willing to blackmail and kill in order to avoid being shut down.
Claude Code just got an "App Store" for agents 🤯
A massive new open-source library just dropped with 100+ pre-made agents, skills, and templates that you can install instantly.
And it's 100% free to use.