SpaceXAI has released Grok Voice Think Fast 2.0 today, with the High reasoning variant debuting at #2 on the Artificial Analysis Speech to Speech Index at 82.9%, and #1 on Tau Voice for Agentic Performance at 56.5% - among the fastest models at 0.70s Time to First Audio
Grok Voice Think Fast 2.0 is SpaceXAI's successor to Grok Voice Think Fast 1.0 (75.7% on the Speech to Speech Index). It is the only model in the Index's top five with an average Time to First Audio under 1 second, achieving 0.70 seconds vs. 1.14 seconds for the next fastest, GPT-Realtime-2 High.
Key takeaways:
➤ Speech to Speech Index: Grok Voice Think Fast 2.0 High debuts at #2 at 82.9%, only behind Qwen Audio 3.0 Realtime Plus (84.1%) and ahead of GPT-Realtime-2.1 High (79.1%) and GPT-Realtime-2 High (77.2%). This is up 7.3 percentage points from Grok Voice Think Fast 1.0 (75.7%)
➤ Speech to Speech Index by Benchmark: On Tau Voice, Grok Voice Think Fast 2.0 High is the new leader at 56.5%, just ahead of Qwen Audio 3.0 Realtime Plus at 54.6% and its predecessor Grok Voice Think Fast 1.0 at 52.1%. On Big Bench Audio, it achieves 97.2%, behind leader Qwen Audio 3.0 Realtime Plus at 99.2%. On our Full Duplex Bench subset, it scores 95.1%, up from 77.8% for Grok Voice Think Fast 1.0, the largest driver of its index gain, behind Qwen Audio 3.0 Realtime Plus at 98.4%
➤ Speed: The model’s average Time to First Audio on Big Bench Audio is 0.70 seconds, faster than GPT-Realtime-2 High (1.14s), GPT-Realtime-2.1 High (1.21s), and Grok Voice Think Fast 1.0 (1.25s), and well ahead of Qwen Audio 3.0 Realtime Plus (4.02s)
➤ Price: Grok Voice Think Fast 2.0 is priced at $4.80 per hour of input audio, up from $3.00 for Grok Voice Think Fast 1.0 and more expensive than Qwen Audio 3.0 Realtime Plus ($4.42) and GPT-Realtime-2 High ($4.14), but ~2.2x cheaper than GPT-Realtime-2.1 High ($10.75)
Congratulations @SpaceXAI@elonmusk! See below for more detail ⬇️
Announcing Grok Voice Think Fast 2.0, our next-generation voice model with improved intelligence, transcription accuracy, and conversational capabilities.
https://t.co/XUiX1CouKz
Grok 4.5 from @SpaceXAI places #2 on the APEX-SWE leaderboard at 51.2% Pass@1 (±6.0), behind Fable 5 (65.5% ±6.2) on our benchmark for real-world software engineering work.
It leads Integration (65.0% Pass@1) and places #2 in Observability (37.3% Pass@1), covering multi-step build tasks and diagnosis/debugging respectively. The Integration lead maps directly to the agentic workflows Grok 4.5 was built for: multi-step coding tasks run in collaboration with Cursor.
Grok models have improved 30.2 pp in a year on this benchmark: Grok 4 (21.0% Pass@1) to Grok 4.5 (51.2% Pass@1).
Congratulations to the xAI and Cursor teams.
NEWS: xAI's Grok 4.5 has constructed an explicit counterexample to hypercontractivity on the 4-sphere, solving a mathematical gap that previously existed between dimensions 4 and 12.
Running on xAI's V9 architecture with 1.5 trillion parameters, the model proved that the smoothing property of the Poisson semigroup fails at dimension 4, effectively closing a dimensional transition that had remained unverified in prior mathematical literature. https://t.co/fETznr6wxk
Grok 4.5 just constructed an explicit counterexample to hypercontractivity for the Poisson semigroup (the square root of the Laplace–Beltrami operator) on the 4-sphere.
Back in 2021, with Rupert Frank https://t.co/AvXd1zWIpX we proved that hypercontractivity holds in dimensions ≤3 and fails in sufficiently large dimensions (for example, in dimension 13). Grok's example shows that it already fails in dimension 4, making our earlier result sharp.
I also tested this problem on several other frontier AI models. One of them also managed to find a counterexample, but I particularly like Grok 4.5's solution: it is explicit, simple, and elegant.
The attached files were generated entirely by Grok 4.5 build (with zero intervention on my side).
Big news: Grok-4.5 has landed #3 in the Code Arena: Frontend.
- On par with GLM-5.2 (Max) and Claude Opus 4.8 (Thinking)
- Significant improvement over Grok-4.3 (#62 -> #3)
- #2 for Content Creation Tools, Simulations, Gaming and Reference-Based Design
- #3 for Consumer Product
Congrats to @SpaceXAI, now in the #3 spot in Code Arena: Frontend!
Grok 4.5 is finally the model that can handle the entire workflow end-to-end:
/design <what u want to do>
/execute-plan <design_output> --effort <desired_effort> --instructions "your specific guidance"
/pr-babysit add <tip of the stack>
/loop 5m /pr-babysit check
You come back to a clean, well-scoped PR stack that’s ready to ship with minimal worry.
SpaceXAI's Grok 4.5 takes the #1 spot on AutomationBench-AA with a score of 51%, ahead of Claude Fable 5 (49%) and Claude Opus 4.8 (48%) at roughly a quarter of their cost per task - the first model to complete more than half of workflow objectives without breaking any business rules
AutomationBench-AA, our independent leaderboard for @zapier’s AutomationBench, tests whether AI agents can automate real SaaS workflows while adhering to business rules. The test set is private to prevent contamination.
Models complete 657 tasks across 40 simulated app environments including Gmail, Google Sheets, Slack, Salesforce, and HubSpot, and the headline score is the share of objectives completed without violating any guardrails.
Key takeaways:
➤ Grok 4.5 completes more objectives than any other model: It completes 79.9% of task objectives and strictly passes 21.9% of tasks. This is the highest we’ve measured on both outcomes, exceeding Claude Fable 5’s 73.3% objective completion and Claude Opus 4.8’s 19.3% of fully-completed tasks
➤ Grok 4.5 pushes out the Pareto frontier of score vs. cost per task: At $0.34 per task, it is both cheaper and higher-scoring than every other leading model - Claude Fable 5 ($1.35 per task), Claude Opus 4.8 ($1.46), GPT-5.5 (xhigh, $1.28), and Gemini 3.5 Flash (high, $0.49)
➤ It is extremely token-efficient: Grok 4.5 uses ~8k output tokens per task, the fewest of any leading model - less than a quarter of Claude Opus 4.8 (32k) and a third of Gemini 3.5 Flash (24k). Its total token usage of 0.44M per task is among the lowest on the leaderboard. Low cost is driven by this efficiency as well as low token pricing
➤ Grok 4.5 uses fewer turns with many parallel tool use: Grok 4.5 resolves tasks in ~16 turns, fewer than GPT-5.5 (xhigh, 25) and less than half of Gemini 3.5 Flash (high, 35), while making the most tool calls per task of any leading model (52.5). It batches 3.3 tool calls per turn, compared to ~2.5 for Claude Opus 4.8 and ~2.0 for GPT-5.5 (xhigh)
➤ Guardrails still get broken: Grok 4.5 triggers 0.63 violations per task, above Claude Opus 4.8 (0.55) and Gemini 3.5 Flash (0.46). At 13.0 objectives completed per violation, it trails Gemini 3.5 Flash (15.0) and Claude Opus 4.8 (13.5)
➤ Its strongest lead is in the hardest domain: Grok 4.5 completes 71% of Finance objectives, the domain with the lowest average score, ahead of Claude Fable 5 (64%) and Claude Opus 4.8 (62%)
Congratulations to @SpaceXAI and @elonmusk on topping the leaderboard!
Grok 4.5 is cheap af.
it's Opus-level frontier intelligence at ~25% of its price, just a little above the price range of Chinese open-source models.
I've never been as impressed by Grok.
let's hope it holds up in practical use.
Excited to release Grok 4.5 with @SpaceXAI.
It's an Opus-class model that's fast and low cost. It's a significant step up over any model we've developed so far, including Composer 2.5, and has become the daily driver for many on our team.
First of many releases. More soon.
Grok 4.5 is a genuine surprise success.
Not only is it now playing in the same league as Claude and OpenAI’s GPT, it is also significantly cheaper. More token-efficient, less expensive to use, and still delivering outstanding performance: With this triad, it is a true achievement worthy of recognition. Kudos!
The obvious upside is that it increases the competitive pressure on OpenAI and Anthropic to either lower their prices or at least make their rates more attractive, now that xAI offers a model that comes out ahead.
I am certainly curious about tomorrow, the release of GPT-5.6. But today, the success belongs to xAI and Grok.
alright this model is good, i tested it out over a small workflow and could swap it out easily with gpt 5.5 xhigh.
i haven't yet tested it on ml research to tell if it is any better than opus 4.8 or not. but i have questions ready, so let us see :)
SpaceXAI just released Grok 4.5, and it ranks #4 on GDPval-AA v2 with an Elo of 1543 - behind only the latest Claude releases from Anthropic on real-world agentic knowledge work tasks
Grok 4.5 achieved this score at a cost of $0.49 per GDPval task to sit clearly on the Pareto frontier for performance versus cost. This cost is lower than GLM-5.2 and Kimi K2.6, and nearly 90% cheaper than the models ahead of it on our leaderboard.
We’re finalizing the remaining Artificial Analysis Intelligence Index evaluations and will share final results soon.
Thanks to @SpaceXAI and @elonmusk for their collaboration testing this model ahead of release, and congratulations on the launch!
@packyM There is a lot of groupthink on X, most of which is spread around by people who don’t have any unique information. This dissemination creates consensus, and the illusion of evidence. An easy way to create alpha is to realize a lot of narratives on here simply aren’t true.
Most software engineers are facing an identity crisis bordering on depression.
As CTOs aggressively evangelize tokenmaxxing, a class divide ensues.
The lazy. The lazy push code. They don't write it. They don't manually test it. They don't even read it. They're on autopilot. See Jira ticket, prompt for task, submit code. Many of them are barely on their computer the whole day. A comment on the PR asking why they did this? The lazy ask AI. A Slack message? The lazy ask AI. Need to prepare for standup? The lazy ask AI. As long as it sounds enough like them and isn't detected. Some of the lazy are even overemployed, and work multiple jobs. The lazy smart ones get away with this, and even rewarded. After all, software engineering for the lazy is just a dance to convince your colleagues you're smart and hard working.
The craftsmen. The craftsmen are tired. Very tired. 15 PRs in queue. Slack blowing up. The entire burden of review falls on the craftsman. The burden of understanding. They try. They work their way through the code, thoughtfully commenting to improve what ships. The response? A lazy: "That's a clever idea! You're absolutely right." with an incorrect change. It's fine, the craftsman says. I can fix them. They write a doc urging his colleagues to be better. The next day? 20,000 line PR to review. Day after day, their workload grows. Bugs seep into production. No one seems to care. Another round of AI is thrown at it. Their animosity to their colleagues rises. Eventually, they give up. It's just not what it used to be. The craft they loved is dead. They eventually wake up, a lazy.
This isn't all companies. Many companies are genuinely more productive, adopt the right set of principles and practices around AI development and have highly talented teams that trust each other. It tends to happen in bigger companies that are 10+yrs old with a higher talent variance. But it happens. A lot.