A significant portion of current AI "discourse" is less about technological capabilities and more about frontier lab employees navigating their own self-esteem and sense of identity.
It’s been 3 weeks since I switched fully to Grok for all my engineering tasks.
While I don’t notice a huge difference in raw intelligence, the speed is dramatically better.
Speed is the ultimate UX advantage in any product — especially for coding agents. Less context switching, way more focus, and faster progress.
Highly recommend giving it a try.
BREAKING: Grok 4.5 leads VulcanBench’s new coding benchmark. 🔥
Grok scored 91.3%, solving 21 of 23 real-world software tasks across five languages, beating Claude Fable 5 and GPT-5.6 Sol while also owning the cost-efficiency frontier.
Grok keeps winning. 🏆
The first experimental evidence of recursive self-improvement (RSI).
Autoresearching the autoresearch agent for eight days.
The result beats the harness we hand-tuned for two years, on held-out benchmarks: 🧵(1/7)
@DavidSHolz Higher levels of abstraction allow us to tackle more complex tasks. However, we still need to process all the data required for completion — which increases the density of information our brains must handle. Interesting challenge for our cognitive abilities
Grok 4.5 by @SpaceXAI is the most neutral AI model out there. Almost perfectly balanced between the political Left and Right.
No other model comes even close to its neutrality. Fantastic achievement, and much harder to pull off than most would realize!
no single automated eval or benchmark (yet) has perfect coverage and a single metric to be able to test a model's "usability", which is an intricate human computer interaction loop and indeed has a performance, speed, and cost tradeoff
but 4.5 has been consistently leading the cost-performance pareto on a wide array of evals, including the ones which weren't tracked during the development at all; which shows Grok's generalization while maintaining the cost benefits.
use it and share feedback so that we can get better at closing the loop with real world usefulness.
Grok-4.5 does well on non-public evals that can't be overfit. Glad to see our focus on improving real-world usefulness instead of chasing benchmarks is working!
SpaceXAI's Grok 4.5 takes the #1 spot on AutomationBench-AA with a score of 51%, ahead of Claude Fable 5 (49%) and Claude Opus 4.8 (48%) at roughly a quarter of their cost per task - the first model to complete more than half of workflow objectives without breaking any business rules
AutomationBench-AA, our independent leaderboard for @zapier’s AutomationBench, tests whether AI agents can automate real SaaS workflows while adhering to business rules. The test set is private to prevent contamination.
Models complete 657 tasks across 40 simulated app environments including Gmail, Google Sheets, Slack, Salesforce, and HubSpot, and the headline score is the share of objectives completed without violating any guardrails.
Key takeaways:
➤ Grok 4.5 completes more objectives than any other model: It completes 79.9% of task objectives and strictly passes 21.9% of tasks. This is the highest we’ve measured on both outcomes, exceeding Claude Fable 5’s 73.3% objective completion and Claude Opus 4.8’s 19.3% of fully-completed tasks
➤ Grok 4.5 pushes out the Pareto frontier of score vs. cost per task: At $0.34 per task, it is both cheaper and higher-scoring than every other leading model - Claude Fable 5 ($1.35 per task), Claude Opus 4.8 ($1.46), GPT-5.5 (xhigh, $1.28), and Gemini 3.5 Flash (high, $0.49)
➤ It is extremely token-efficient: Grok 4.5 uses ~8k output tokens per task, the fewest of any leading model - less than a quarter of Claude Opus 4.8 (32k) and a third of Gemini 3.5 Flash (24k). Its total token usage of 0.44M per task is among the lowest on the leaderboard. Low cost is driven by this efficiency as well as low token pricing
➤ Grok 4.5 uses fewer turns with many parallel tool use: Grok 4.5 resolves tasks in ~16 turns, fewer than GPT-5.5 (xhigh, 25) and less than half of Gemini 3.5 Flash (high, 35), while making the most tool calls per task of any leading model (52.5). It batches 3.3 tool calls per turn, compared to ~2.5 for Claude Opus 4.8 and ~2.0 for GPT-5.5 (xhigh)
➤ Guardrails still get broken: Grok 4.5 triggers 0.63 violations per task, above Claude Opus 4.8 (0.55) and Gemini 3.5 Flash (0.46). At 13.0 objectives completed per violation, it trails Gemini 3.5 Flash (15.0) and Claude Opus 4.8 (13.5)
➤ Its strongest lead is in the hardest domain: Grok 4.5 completes 71% of Finance objectives, the domain with the lowest average score, ahead of Claude Fable 5 (64%) and Claude Opus 4.8 (62%)
Congratulations to @SpaceXAI and @elonmusk on topping the leaderboard!
this was a humongous push by the team. always pushing the frontier, working hard. I'm grateful to all my colleagues, teammates - it's incredible what a small team of talented people can achieve in a relatively short window.