AutomationBench is based on two metrics: Score and Tasks Completed.
On the first, GPT-6 Astra scores 68.5%
On the second, only 41.6% of workflows are completed without a guardrail violation. Claude Fable 5.1 only completes about 32.1% of workflows.
Opus 5 (one of my favorites) completes 89% of objectives given regardless of guardrail violations. If you take into account guardrail violations, it only completes about 28.3% of workflows.
I am still diving deep into Intelligence Analysis Index v4.3 and I learnt something quite interesting.
This upgrade is not just about ensuring the AI gets the job done. It is more about getting the job done without breaking the rules.
If an agent violates a guardrail, AutomationBench-AA gives it a zero for that particular workflow. I love it!
Ideally, you don't want an agent that does 99% of the job but then sends confidential data to the wrong person, changes or deletes data, creates a duplicate record, modifies a financial record, etc.
That will no longer be 99% of the job done.
45% of the Artificial Analysis v4.3 now uses private or held out tasks.
These are basically test questions that no AI company has prepared for in advance. It's like an exam for AI models with goal of testing if AI can solve problems it has never encountered before.
This might push AI systems to become adaptable when it comes to solving new unforeseen and unanticipated problems. It will also result in better handling of uncertainty taking away reliance on patterns from known benchmarks.
What matters more to you in an AI?
- ACTION; or
- THINKING
Answering that question is going to be very important because you will have to choose between:
- GPT-6 Astra the Executioner; or
-Fable 5.1 the Mastermind.
AI model specialization at it's best.
What's the cheapest AI model per task?
According to Artificial Analysis Intelligence Index v4.3, these is how the top models rank:
- GPT-5.6 Luna (max) - $0.182
- GLM-5.3-Flash - $0.253
- GPT-5.6 Terra (max) - $1.404
- GPT-6 Astra (max) - $3.265
Claude Fable 5.1 (max + fallback) - $7.63
If you needed proof that AI is going past answering questions to actually doing the work, then Artificial Analysis Intelligence Index v4.3 has your answer.
The new category weights are:
- Agents 30%
- Coding 20%
- General 30%
- Scientific Reasoning 20%
Less chatbot. More coworker. More execution.
Open-weight has made some significant leaps but it is still far from the frontier models.
Given the latest Artificial Analysis Intelligence Index v4.3, they rank as follows:
Closed (Proprietary)
โข Claude Fable 5.1 โ 53
โข GPT-6 Astra โ 53
โข Claude Opus 5 โ 51
โข Claude Fable 5 โ 50
โข Muse Spark 1.3 โ 48
โข GPT-5.6 Sol โ 47
Open-Weight
โข GLM-5.3 โ 44
โข Kimi K3 โ 44
โข GLM-5.3-Flash โ 42
โข Qwen3.8 2.4T โ 40
โข DeepSeek V4 Pro โ 36
The best closed models sit at 53.
The best open-weight models sit at 44.
Thatโs a massive 9-point gap.
Artificial Analysis Intelligence Index v4.3 upgraded two benchmarks:
- Terminal-Bench 4.0; and
- AutomationBench-AA
Here's what each means and why it is important:
1. Terminal-Bench 4.0
A hard test of real agentic coding and terminal work.
66 multi-step tasks in actual sandboxes โ software engineering, ML experiments, science, system operations. The AI has to plan, run commands, debug, and finish the job under tighter time and compute limits. Harder instructions and better verification than the old version.
2. AutomationBench-AA
A private test of everyday business automation.
657 real workflows across tools people actually use (Gmail, Slack, Salesforce, Jira, etc.). The AI must complete the full objective while following company rules โ no shortcuts, no breaking guardrails. Built with Zapierโs held-out test set so models canโt memorize the answers.
GPT-6 Astra and Fable 5.1 are the two frontier models tying at an Artificial Analysis Intelligence score of 53. This is how each scored on each of the upgraded benchmarks:
Terminal-Bench 4.0
โข GPT-6 Astra (max) โ 59.1%
โข Claude Fable 5.1 (max + fallback) โ 52.0%
AutomationBench-AA
โข GPT-6 Astra (max) โ 68.5%
โข Claude Fable 5.1 (max + fallback) โ ~59%
Astra leads Fable 5.1 in both practical agentic tests.
Simply put:
- Claude and OpenAI are level when it comes to pure intelligence.
- Astra is the best model for real-world agent work. If you are looking for a model that can write and fix code in a terminal, or run a multi-app business, Astra is your best bet.
So now, the best AI tool for your or your busines depends on your need:
If you need the strongest general reasoning coupled with knowledge work then go for Fable 5.1.
If you need an AI a software engineer or office automation agent then Astra is your go to choice.
Oh! You may also want to think about cost and speed. Astra is significantly cheaper per task.
It's official! Artificial Analysis Intelligence Index v4.3 is here.
This is how every model lines up on the AI leaderboard at the moment:
๐ฅ Claude Fable 5.1 โ 53
๐ฅ GPT-6 Astra โ 53
๐ฅ Claude Opus 5 โ 51
4๏ธโฃ Claude Fable 5 โ 50
5๏ธโฃ Muse Spark 1.3 โ 48
6๏ธโฃ GPT-5.6 Sol โ 47
7๏ธโฃ GLM-5.3 โ 44
7๏ธโฃ Kimi K3 โ 44
9๏ธโฃ GLM-5.3-Flash โ 42
๐ Qwen3.8 2.4T A95B โ 40
11. DeepSeek V4 Pro 0813 โ 36
What changed in this version of the benchmark?
Artificial Analysis upgraded Terminal-Bench from 2.1 to 4.0 and added AutomationBench-AA, which tests AI agents across 657 business workflows involving simulated applications like Gmail, Slack, Salesforce and Jira.
GPT-6 Astra had an impressive score:
59.1% on Terminal-Bench 4.0.
68.5% on AutomationBench-AA.
And that's not all. Astra is also cheaper than Fable 5.1.
They both score 53, but Artificial Analysis estimates Astra's cost per Intelligence Index task at $3.26 vs $7.63 for Fable 5.1.
That's a 57% lower cost per task on GPT-6 Astra.
It's not just about cost. Consumers are looking at a mix of:
Intelligence ร reliability ร autonomy ร cost.
Chinese LLMs have been unusually silent. Moonshot is not updating us on what benchmarks Kimi might beat. DeepSeek keeps getting mlre and more silent. Qwen is almost non-existing. I wonder what's happening down there.
@deel is about to release @get_akai and I am super stoked about this one.
Akai is the best agentic AI you will come across:
- it operates compliantly
- get's cheaper over time
- no code workflow creation
- easily works with existkng systems
- has human in the loop controls
And the best part? It has already worked internally.
Akai currently handles 100,000+ cases per month and saves 91,000+ hours of manual work every month within Deel.
This month, Akai is premiering for all!
Since the 1950s, the world has experienced multiple boom-and-bust AI cycles . It's only until the post-ChatGPT era that the hype cycles have become faster and extreme.
But looking at it deeper, the structure is familiar:
It starts with an Innovation Trigger. Big leaps are made in terms of what AI can do.
Then in comes a Peak of Inflated Expectations. That's when you hear things like โThis changes the game completely" or much recently "AGI is here.โ
Before you know it, we are caught in Trough of Disillusionment. Limitations start showing (think about all the vibe coding talk, the high cost of AI, etc), ROI does not add up, and organizations start making quiet walk-backs.
The wheels of AI keep moving and then the Slope of Enlightenment comes. We start experiencing real-world practical use cases.
Eventually, we land at the Plateau of Productivity. AI becomes integrated into the society and consequentially becomes useful infrastructure.
But there's something different about this new era of AI. We're no longer experiencing clean winters. The cycles keep overlapping. New models are arriving faster and faster before the previous hype fully dies.
I don't think there's anything more dystopian than witnessing an AI agent complete 90% of a complex workflow successfullyโฆ then it proceeds to fail on the remaining 10% confidently. The worst part is that as it fails, it still consumes your token limit.
Since we're past prompting and now into vibe coding, the newest skill one needs to learn is "steer coding."
This is the skill where you know when to let the model run and when to hit stop before $200 in tokens goes down the drain doing something silly.