Anthropic’s Economics team is sharing a new model of how AI might affect economic growth, jobs, wages, and more by 2030.
Explore the scenarios, tell us what you think will happen, and see how your answers compare to more than 10,000 Americans. https://t.co/AvQlEZNxR0
Does anyone seriously care about these benchmarks anymore?
They lost their credibility with initial ones about Astra and now keep updating them until the models reach their perceived state in public view
Announcing Artificial Analysis Intelligence Index v4.3, upgrading Terminal-Bench to 4.0 and adding AutomationBench-AA, an agentic workflow automation benchmark with a private test set. This is a continuation of our rollout of Intelligence Index v5
Changelog (Index v4.2 → Index v4.3):
➤ Terminal-Bench: 2.1 → 4.0, completing our upgrade to the latest version of Terminal-Bench
➤ Replacing 𝜏³-Banking with AutomationBench-AA, our implementation of Zapier's business workflow automation benchmark
We are continuing to prioritize keeping Intelligence Index as useful as possible by bringing forward a subset of the changes we had planned for Index v5. Each change in v4.2 and v4.3 stands on its own merits and brings the Index closer to real-world problem solving, adds more private test sets to prevent gaming, and reduces saturation
Intelligence Index v4.3 raises the difficulty of agentic coding tasks and broadens the types of agentic workflows tested. Because we use a held-out test set for AutomationBench-AA, in collaboration with @zapier, the weight assigned to evaluations with private tasks or answers increases from 40% to 45%. Category weights are unchanged from v4.2: Agents 30%, Coding 20%, General 30%, Scientific Reasoning 20%
Detailed changes:
➤ Upgraded Terminal-Bench 2.1 to 4.0: 66 multi-step tasks testing agents on tasks run in agent sandboxes driven via the terminal, including tasks involving software engineering, machine learning, science, and operations. The 4.0 update recalibrates compute and time allowances, and improves task instructions and verification. We have changed from the Terminus 2 harness to mini-SWE-agent, a minimal, model-agnostic harness. We will also be updating our Coding Agent Index, where we test model and harness pairs, to include Terminal-Bench 4.0 soon
➤ Replaced 𝜏³-Banking with AutomationBench-AA: Our implementation of Zapier’s AutomationBench tests agents on 657 business workflows across simulated applications such as Gmail, Slack, Salesforce, and Jira. Agents must complete task objectives while following business rules. AutomationBench-AA uses Zapier’s private set of 657 tasks, and is built on v1.0.6
Key results:
➤ Claude Fable 5.1 and GPT-6 Astra lead the Intelligence Index: Both Claude Fable 5.1 (max with fallback) and GPT-6 Astra (max) score 53 on Intelligence Index v4.3, followed by Claude Opus 5 (max, 51), Claude Fable 5 (with fallback, 50), Muse Spark 1.3 (max, 48) and GPT-5.6 Sol (max, 47)
➤ GLM-5.3 and Kimi K3 continue to lead open weights models (both at 44): GLM-5.3-Flash (42) is the third strongest open weights model, followed by Qwen3.8 2.4T A95B (40) and DeepSeek V4 Pro 0813 (max, 36)
➤ 4 labs occupy the Intelligence vs. Cost per Task Pareto frontier: OpenAI occupies the majority of the cost-efficiency frontier, with all five reasoning efforts of the recently released GPT-6 Astra offering the lowest Cost per Task at their respective levels of intelligence. Claude Fable 5.1 (xhigh, max, 53), GLM-5.3-Flash (42) and MiMo-V2.5-Pro (26) round out the rest of the frontier
You and me are not going to switch to Gemini, Muse or some other subscription until they can match or replace the value of OpenAI/Anthropic
Spending API credits is out of question too of-course
This makes it really hard for other companies to monetize their models even if they're getting good
We all have limited personal money we choose to spend toward these AI subscriptions
Currently, both OpenAI and anthropic are worth getting for coding
Grok 4.6 is generously available if you use Cursor
This leaves models like Gemini, Muse and others in the "meh" area. Why?
You are supposed to ask these SOTA frontier models about specific problems
If you tell them to read your entire codebase then no shit you're going to run out of usage
I see a lot complaining about Astra burning usage on their $20 plan but like, what did you expect?
At least Astra is available as an option compared to what Anthropic has done:
- no Fable at all
- %50 usage allowed on Fable for $100+ plans
I don't think so. OpenAI made GPT-Sol more efficient few weeks ago and it's really good to use
Just use reasoning efforts better. low for implementation, high+ for thinking/design/review
Codex Plus users, just use Luna
Just accept that Sol and Astra are not for the Plus plan
This reality. Sol and Astra are very smart but very expensive models. $20 doesn’t go a long way
I’m not saying stay away from Astra or Sol, I love these models! They get things done fast!
I will get Codex Pro soon once I’ve rebalanced my AI subs
@quxiaoyin I have been building on 3 $20 plans and have to think before prompting to make the most of it.
1b+ burn is encouraging burning tokens without thought which is ehhh
It's been nearly 2 months but @Kimi_Moonshot still isn't accepting new subscribers...
Kimi K3 just melted their capacity. Not to mention the surge of local users due to Claude/GPT unavailability in China