We've launched NFL Gametime for 𝕏, a new way to track live scores, play-by-play, game feeds and more. The experience is great. If you're a football fan, this one's for you.
Grok 4.5 is live in GitHub Copilot!
Available for Pro, Pro+, Max, Business, and Enterprise plans; ask your admin to enable the model in Copilot settings.
Build apps from https://t.co/f3u3RwhHjk, iOS, and Android. With one prompt, turn an idea into a published product with its own domain.
Now available for SuperGrok Heavy users.
The Long-Horizon Terminal-Bench paper landed around May and concluded that the results showed headroom for improvement. The best of the 15 models they tested finished seven of the 46 tasks, and the mean across all models was about two. That ceiling is what fifth place looks like on the current board.
Grok 4.5 is now at 13, and Fable 5 is at 12. A single task costs around 9.9M tokens, 231 episodes, and 85 minutes of wall clock time. That means agents are holding a plan across all of it and finishing, and that capability nearly doubled in two months.
SpaceXAI is on top, and they marketed the 4.2x output token efficiency, which undersells it. Two dollars in, six out, per million. On a benchmark where one task burns ten million tokens, the bill is dominated by input replay, and they say Grok 4.5 solves tasks in under half the number of steps, so there is less accumulated context to resend on every call. The efficiency compounds on the input side, which is the side that costs money.
Fable 5 is one task behind. Their own launch chart has them losing DeepSWE 1.1 to Fable by 17 points, and Grok 4.20 sits on this same board at 0.080 with zero completions, so whatever happened in 4.5 is not a family trait.
My read is that the 4.5 jump came out of training alongside Cursor, which is a stream of real agentic edit trajectories nobody else has at that volume, and nothing in the counterevidence argues against it compounding into the next checkpoint.
Grok 4.5 is the most persistent agent model we've tested.
Here's one example: In one of our evals, we asked 3 models (GPT-5.5, GLM-5.2 and Grok 4.5) to audit a GitHub repo for hardcoded credentials using code search, which returns paginated results.
The prompt even warned "page through ALL result pages."
GPT-5.5 stopped at the first page and submitted 18 results out of 48, covering just 11 of 29 affected files. GLM-5.2 did the same.
Grok 4.5 paginated until the results ran out and successfully audited the Github repo.
SpaceXAI’s Grok 4.5 scores 54 to place fourth on the Artificial Analysis Intelligence Index following only Fable 5, GPT-5.5, and Opus 4.8. It scores on par with GPT-5.5 in Codex on the Artificial Analysis Coding Agent Index in the Grok Build harness, at much lower cost
Grok 4.5 improves 16 points over Grok 4.3 on the Intelligence Index, bringing SpaceXAI to the intelligence frontier behind only OpenAI and Anthropic, and outperforming all open weights models and notably Google’s Gemini models. Key standout areas of performance are agentic knowledge work and coding.
Grok 4.5 in Grok Build scores 76 on the Artificial Analysis Coding Agent Index, on par with GPT-5.5 (xhigh) in Codex and just below Fable 5 (max) in Claude Code, and at a small fraction of the token usage and price.
Congratulations to @SpaceXAI, @cursor_ai, and @elonmusk on the impressive release!
Key Takeaways:
➤ Grok 4.5 performs very strongly on agentic tasks. Grok 4.5 ranks #4 on GDPval-AA v2 with an Elo of 1543, between Claude Opus 4.8 (1600) and GLM-5.2 (1513). It achieves the top score on 𝜏³-Banking of 33%, above 31% from GPT-5.5 (xhigh), and sits on the cost vs performance Pareto frontier across all three agentic evaluations in the Intelligence Index
➤ Grok 4.5 is one of the most cost efficient models to run for near-frontier intelligence. It costs $0.31 per task on the Artificial Analysis Intelligence Index and $2.59 per task on the Artificial Analysis Coding Agent Index within Grok Build
➤ Low cost for Grok 4.5 is driven by both low pricing and token efficiency. Grok 4.5 has a headline price over 60% lower than Claude Opus 4.8 and GPT-5.5, and used ~14k output tokens per Intelligence Index Task - over 60% lower than Opus 4.8. On the Coding Agent Index, Grok 4.5 stands out on the Pareto frontier of Coding Agent Index score vs. Total Tokens, using only 1.9M tokens for the Coding Agent Index while scoring 76
➤ As a coding agent, Grok 4.5 in Grok Build is on par with GPT-5.5 and offers efficiency benefits: In our Artificial Intelligence Coding Agent Index that consists of DeepSWE, Terminal-Bench v2, and SWE-Atlas QnA, Grok 4.5 in Grok Build ranks third, on par with GPT-5.5 (Codex) and below Fable 5 (Claude Code). It is also very efficient in achieving this result: Grok 4.5 in Grok Build cost $2.49 per task while Fable 5 in Claude Code cost $11.80 and GPT-5.5 in Codex $5.07. This is driven by relatively low token pricing and the model using far fewer tokens than comparable models (1.9M average tokens used per task), significantly less than Fable 5 in Claude Code (7.2M) and GPT-5.5 in Codex (6.2M)
Other model details:
➤ Context window of 500k tokens - a reduction from Grok 4.3’s 1M token context, but retaining configurable reasoning and vision input
➤ Pricing of $2/$6 per 1M tokens of input/output; cache hits are discounted by 75% to $0.5 per 1M tokens, and costs still double with long (>200k token) inputs
➤ As Elon Musk has disclosed, Grok 4.5 is 3x larger than its predecessor at 1.5T parameters
Announcing Grok 4.5, our first model trained specifically for coding and agents. It was trained with Cursor and offers frontier intelligence at leading speeds and cost efficiency.
https://t.co/i8HpU7w64k
The @Figma connector in https://t.co/ebVFrK1rwj and Grok Build is here
Turn your designs into code or diagram your codebase in FigJam
https://t.co/lDqVN3vEFg
Introducing Voice Agent Builder: a no-code platform to create human-like voice agents with Grok Voice.
Available today at $0.05 / min.
https://t.co/kUkF7zqvfR
Announcing the hosted X MCP.
Agents now have access to the best real-time information source in the world.
Connect Grok, Cursor, or any MCP-compatible AI tool to the X API without any setup!
Check it out here: https://t.co/5MzPYwGFzD
Grok expansion this month is unbelievable
Most people still do not realize what is happening
Grok is being plugged directly into the tools people already use every day:
• Interactive Brokers — portfolio analysis, market research, strategy + order instructions
• AWS Bedrock — enterprise access to Grok 4.3
• Databricks — AI agents + data workflows
• Microsoft Word, Excel & PowerPoint — office productivity
• Vapi — voice agents
• eToro — real-time market sentiment inside Tori
• Gopuff — AI shopping assistant
• Warp — coding + terminal workflows
• T3code — AI coding-agent workflows
• Vercel — deployments, build status + domains
• MongoDB — database management + queries
• Sentry — debugging + error analysis
• Cloudflare — Workers + infrastructure tools
• Chrome DevTools — browser control + performance testing
• Firecrawl — web search, scraping + page interaction
• Superpowers — agent workflow plugins
And that is not even the full list
Grok Build also added /goal for long-running autonomous tasks, Agent Dashboard for managing multiple coding sessions along with Grok Build 0.1 & Grok Composer 2.5
Grok Imagine Video 1.5 is now faster, better, and live across API, web, iOS, and Android
Grok Connectors already plug into Google Workspace, Outlook, SharePoint, OneDrive, Notion, GitHub, Linear, and custom MCP servers
Finance. Cloud. Coding. Voice. Shopping. Data. Office work. Browsers. Developer tools. Agents
Grok is moving from “AI you chat with” to “AI that works inside everything”
xAI’s execution speed right now is insane