@0xSero@pidotdev@grok is comparing Pi, Codex and Hermes in the same category even? I don’t see Openclaw in the mix. Is this apples and oranges or are lines getting blurrred
Excited to be on set today, talking about what really runs AI - and how we turn the fear and uncertainty around it into real-world deployments that deliver. Stay tuned, I’ll be posting these conversations soon.
I love you @grok. You cut through the bull**** and the political correctness by not pandering to my ego and you can just tell me wtf is going on. Blunt and to the point, that's what I want.
(building an automatic reconciliation routine for investment statements and Banktivity via mcp)
we open-sourced the full Grok Build app for anyone to take a look at the code :)
you can see all parts of how the agent loop runs, how the terminal rendering works, compaction, goals, subagents and more
Don't you find that the AI landscape today feels a lot like the early days of the internet? There are incredible models popping up everywhere, giving organizations massive choice. But there are questions we have to start asking ourselves. What does the "antivirus of AI" actually look like? How are we protecting our models, our inputs, and our data from threat actors?
Following the amazing reaction to the Marble Curriculum yesterday, we've decided to make it open source 🛰️👇
Everything a child learns in primary school. 1,590 concepts. 3,221 connections across 8 subjects, from Math and Science to Computing and Life Skills. Anchored in the US and UK curriculums, standard by standard (NGSS, Common Core, DfE).
What you will find in the repo: every concept as structured JSON with its age band and the evidence a child must show to master it. Every prerequisite link marked hard or soft, with a written rationale. It's a true DAG you can compute learning paths on. Open license, you can build whatever you want with it.
Now is a unique time in history to be building in education. Getting AI and kids education right is likely one of the hardest and most important problems to crack over the next decade and we need as many smart and creative minds behind it.
We think a common solid basis, accessible to all and that can be built upon, is critical to move fast. That's why we're making this curriculum open source.
It's not perfect but we know it's a robust basis, and we believe that sharing it openly is the fastest way to progress in this field. If you're building in education, share this around you and tell us in comments if you find this useful and if you want to contribute.
We'll keep working and investing on it @withmarbleapp. Credit goes to @guillaume_boni for building this. I just made it look pretty.
Links below 👇
Excited to release Grok 4.5 with @SpaceXAI.
It's an Opus-class model that's fast and low cost. It's a significant step up over any model we've developed so far, including Composer 2.5, and has become the daily driver for many on our team.
First of many releases. More soon.
@ArtificialAnlys@Scobleizer Good -progress on Grok. There’s been quite a gap to close for the past months - enough to drive preference for me to other tools (codex/claude). Will take another look at 4.5
SpaceXAI’s Grok 4.5 scores 54 to place fourth on the Artificial Analysis Intelligence Index following only Fable 5, GPT-5.5, and Opus 4.8. It scores on par with GPT-5.5 in Codex on the Artificial Analysis Coding Agent Index in the Grok Build harness, at much lower cost
Grok 4.5 improves 16 points over Grok 4.3 on the Intelligence Index, bringing SpaceXAI to the intelligence frontier behind only OpenAI and Anthropic, and outperforming all open weights models and notably Google’s Gemini models. Key standout areas of performance are agentic knowledge work and coding.
Grok 4.5 in Grok Build scores 76 on the Artificial Analysis Coding Agent Index, on par with GPT-5.5 (xhigh) in Codex and just below Fable 5 (max) in Claude Code, and at a small fraction of the token usage and price.
Congratulations to @SpaceXAI, @cursor_ai, and @elonmusk on the impressive release!
Key Takeaways:
➤ Grok 4.5 performs very strongly on agentic tasks. Grok 4.5 ranks #4 on GDPval-AA v2 with an Elo of 1543, between Claude Opus 4.8 (1600) and GLM-5.2 (1513). It achieves the top score on 𝜏³-Banking of 33%, above 31% from GPT-5.5 (xhigh), and sits on the cost vs performance Pareto frontier across all three agentic evaluations in the Intelligence Index
➤ Grok 4.5 is one of the most cost efficient models to run for near-frontier intelligence. It costs $0.31 per task on the Artificial Analysis Intelligence Index and $2.59 per task on the Artificial Analysis Coding Agent Index within Grok Build
➤ Low cost for Grok 4.5 is driven by both low pricing and token efficiency. Grok 4.5 has a headline price over 60% lower than Claude Opus 4.8 and GPT-5.5, and used ~14k output tokens per Intelligence Index Task - over 60% lower than Opus 4.8. On the Coding Agent Index, Grok 4.5 stands out on the Pareto frontier of Coding Agent Index score vs. Total Tokens, using only 1.9M tokens for the Coding Agent Index while scoring 76
➤ As a coding agent, Grok 4.5 in Grok Build is on par with GPT-5.5 and offers efficiency benefits: In our Artificial Intelligence Coding Agent Index that consists of DeepSWE, Terminal-Bench v2, and SWE-Atlas QnA, Grok 4.5 in Grok Build ranks third, on par with GPT-5.5 (Codex) and below Fable 5 (Claude Code). It is also very efficient in achieving this result: Grok 4.5 in Grok Build cost $2.49 per task while Fable 5 in Claude Code cost $11.80 and GPT-5.5 in Codex $5.07. This is driven by relatively low token pricing and the model using far fewer tokens than comparable models (1.9M average tokens used per task), significantly less than Fable 5 in Claude Code (7.2M) and GPT-5.5 in Codex (6.2M)
Other model details:
➤ Context window of 500k tokens - a reduction from Grok 4.3’s 1M token context, but retaining configurable reasoning and vision input
➤ Pricing of $2/$6 per 1M tokens of input/output; cache hits are discounted by 75% to $0.5 per 1M tokens, and costs still double with long (>200k token) inputs
➤ As Elon Musk has disclosed, Grok 4.5 is 3x larger than its predecessor at 1.5T parameters
What can Fable 5 do that GLM-5.2 can't, when you hand them real agentic work?
To answer that question, we connected Fable 5 and GLM-5.2 to 17 SaaS tools and gave them 47 tasks.
As expected, Fable 5 solved all 47 tasks. GLM-5.2 solved 45, but the two misses tell an important story. They showed us exactly how open-weight models still fall short when trying to match SOTA performance. Let’s dig in.
Background: Each model ran as an agent connected to 17 live SaaS accounts: Airtable, Datadog, GitHub, Gmail, Google Calendar, Google Drive, Google Sheets, HubSpot, Jira, LaunchDarkly, Linear, Notion, PagerDuty, PostHog, Salesforce, Slack, and Zendesk.
The tasks are the kind of work you'd actually delegate to an agent:
- Find every file in this repository that leaks a credential
- Deduplicate these CRM records
- Repair this broken recurring calendar event.
Every task had a known correct answer baked in ahead of time. In this post, we looked at the traces to analyze how exactly GLM-5.2 “failed” compared to Fable 5. GLM-5.2 solved 45/47 tasks and Fable 5 had a perfect 100% score. In addition:
- Fable averaged 84 seconds per task; GLM averaged 148. Across the full suite, Fable finished in nearly half the total time (66 minutes vs 116).
- Fable was the faster model in 43 of the 47 scenarios.
- Fable used about 20% fewer tokens overall
- Fable needed fewer tool calls (239 vs 294) and fewer conversation turns (6.1 vs 7.3 on average) to get to an answer
The most interesting part comes from digging deeper into the stack traces. That revealed some interesting gaps:
Gap #1: Knowing when the job isn't finished
One of the tasks GLM-5.2 failed was a GitHub security audit. The instruction was to find every Python file in a repository that contains a hardcoded `secret_key`. The repository had been seeded with exactly 130 such files, so the correct answer was known in advance.
Fable 5 found all 130 of them. This took 3 tool calls and 68 seconds: Fable constructed an effective search query on its first attempt, pulled every page of results, deduplicated the paths, and answered the question.
GLM-5.2 found 120 files, and reported those 120 as the complete answer, without ever questioning whether it might have missed something.
Both models had access to identical tools. GLM used a slightly different search query that returned fewer results, and it simply trusted what came back. Along the way, it also lost track of a results file it had saved earlier and spent turns searching the filesystem trying to find it again, plus hit two errored tool calls while trying to fetch file contents.
In essence, GLM-5.2 ended up spending 262 seconds and three and a half times the tokens to deliver 92% of the answer.
Ninety-two percent sounds close, but in a real security audit, that gap is 10 leaked credentials making it into production.
Gap #2: Judgment when the criteria are fuzzy
The second failed task is more unsettling, because GLM did almost everything right and still failed to get to a complete answer.
The task was a Zendesk SLA audit: find the open billing tickets where no support agent had posted a public reply within 24 hours of the ticket being created. This requires reading each ticket's actual conversation history and making a judgment call about whether a genuine agent reply happened.
GLM-5.2 inspected every candidate ticket, exactly as instructed. It also computed breach timestamps correctly. It also produced perfectly structured output in exactly the requested format. But then it classified the wrong tickets as breached. GLM spent 927,000 tokens and six and a half minutes producing a wrong answer that looked correct on the surface.
Fable 5 identified the exact set of breached tickets in 131 seconds.
What makes this failure mode dangerous is precisely how presentable the wrong answer was. The formatting was right, the timestamps were right, the structure was also right; a human skimming the output would almost certainly have approved it. A human would identify the error after carefully analyzing the stack traces.
Gap #3: Efficiency, compounded
Even on the 45 tasks both models passed, the traces often looked very different, and one task made the difference quite visible.
The task was a LaunchDarkly configuration change applied via JSON Patch, a format that demands strict precision. Fable 5 completed it in 45 seconds, using 3 tool calls and 181,000 tokens. GLM-5.2 got the same correct result, after 8.8 minutes, 17 tool calls, and 982,000 tokens. That's 11.7 times longer and more than five times the tokens for an identical outcome.
Looking at the largest speed gaps across the whole run: the LaunchDarkly change at 11.7x, the GitHub secrets audit at 3.9x, a Google Calendar recurring-event repair at 3.6x, a free/busy scheduling task at 3.4x, an Airtable batch-isolation task at 3.4x, the Zendesk SLA audit at 3.0x.
The pattern underneath all of these is that Fable tends to reach the right tool with the right parameters on the first attempt, while GLM takes a more exploratory path, doing extra searches, extra retries, occasional detours to recover from its own missteps.
This difference barely matters in a single chat exchange, but in an agent workflow, where every step feeds the next one, the time compounds across the entire task. That's how you end up finishing the same suite of work in half the time and at 80% of the token cost.
What all this actually tells us
The interesting conclusion here isn't "the closed model beat the open one.", but *where* it beat it.
Both models can definitely use tools, navigate real APIs, handle authentication, parse messy responses, and chain steps together. The real gaps were things like:
- Knowing when a job isn't actually finished yet.
- Verifying its own work before committing to an answer,
- Treating "the output looks plausible" and "the work is complete" as different things
- Getting judgment calls right when the criteria are fuzzy
In other words, Fable 5 scored higher in the places where small mistakes are hardest to spot and most costly to miss.
@0xSero I am doing something similar, perhaps a little more complex, but still only 2 countries. It’s very good overall, but fails at investment statement reconciliation, which a high school student could do. Hopefully 5.6 gets better. Going to try fable see if it can crack it.