Indeed!
Context:
Megha personally ran many of the evals you are seeing on the Gemini 4 Argon model card today!
Lots of hard work goes into making a release happen, thank you for giving it your all 😌
And thank you for organizing our Boba Tea surprise today 🧋🧋
I work at Google DeepMind and yes I have tried the model known externally as “Argon”.
I can say this:
I don’t reach for Opus anymore for any work (like I did when 3.8 Flash was our top model).
I hope that it will soon be generally available for everyone to enjoy!
I don’t know if you really want the internal version. All the bugs go through us first 😄. What surprised me recently - the harness isn’t as important these days to agents’ capabilities . I ran large internal evals on different popular harnesses (open and proprietary) and found out that the pass rate difference was mostly marginal, and when it was not - it was fixable by AGENTS.md files. It’s the human UX that is different, which an eval cannot measure well (yet).
@iamhirusha To be fair, if I need to do a well-scoped, simple task that takes a lot of time to do (like writing tests or splitting commits), I’ll pick 3.8 Flash over Opus anytime because it’s blazing fast and will be done x2 times faster than Opus. Giving it serious work though….nope
@UnsoldBanana I personally agree that we should follow a different release pattern from what we have now (and I could be wrong). However I’m not the person making decisions like that at Google 🤷♂️
That is a fair point. Worth noting that benchmarks themselves have evolved (both externally and internally). They became much harder and much better at discriminating model performance. At DeepMind now we look at pass-rate and traces of individual tasks across many runs.
btw these were the official benchmarks for Gemini 3.1 Pro when it released, practically destroying Opus 4.6 across the board.
I'm not saying Gemini 4 Argon will be bad, but it's crazy that no one has the slightest bit of skepticism given Google's track record.
@patchwanders Of course. Trivially - tool call failure rate. Some models (like earlier versions of GPT from this year) had some issues with calling a non-existent file tools. That is likely an artifact of their post-training environment where the model was overfitted to certain tools.
Google’s new Gemini 4 Argon equals GPT-6 Astra on the Artificial Analysis Intelligence Index at 60% of the Cost per Task with discounted prices. Google is now back to being one of the top three labs in intelligence achieved
Gemini 4 Argon is @GoogleDeepMind’s first proprietary model above the Flash class in over 7 months. With high reasoning (the highest available), it scores 53 on the Artificial Analysis Intelligence Index, matching GPT-6 Astra (max, 53) and 1 point ahead of GPT-6.1 Sol (max, 52), with gains driven by lower hallucinations and stronger agentic capabilities.
At its current 50% pricing discount and with cache discounts increased to 95%, Gemini 4 Argon costs $1.99 per Intelligence Index task, 60% of GPT-6 Astra (max), but 2.7x GPT-6.1 Sol (max). After the discount ends, this will rise to $3.98 (~1.2x GPT-6 Astra (max)).
Gemini 4 Argon is currently being rolled out to selected users and is not publicly available. The 50% discount is an initial promotion. Google has not yet confirmed the promotion end date
Key benchmarking results for Gemini 4 Argon with high reasoning:
➤ Google returns as one of the top three labs on intelligence: Gemini 4 Argon (high) scores 53 on the Artificial Analysis Intelligence Index, matching GPT-6 Astra (max, 53) and 1 point ahead of GPT-6.1 Sol (max, 52). This is 23 points above Google’s previous non-Flash model, Gemini 3.1 Pro Preview (30) and 12 points ahead of Gemini 3.8 Flash (high)
➤ Launch discounts of 50% make Gemini 4 Argon competitive on Cost per Task: At current discounted pricing, Gemini 4 Argon (high) costs $1.99 per Intelligence Index task, 60% of GPT-6 Astra (max, $3.26) for a comparable level of intelligence. This cost efficiency is driven by lower token prices, rather than reduced token use, with Gemini 4 Argon averaging 62k output tokens per task, compared with 27k for GPT-6 Astra (max). Google has not yet confirmed the promotion end date, but on standard pricing, Cost per Task will increase to $3.98
➤ Stronger agentic performance: Historically a weaker area for Gemini models, Gemini 4 Argon shows improvements across agentic evaluations. It ranks #1 on AutomationBench-AA at 77.5%, 6 points ahead of Claude Sonnet 5.5 (max, 71.3%). On Terminal Bench 4, Gemini 4 Argon achieves 57%, a +53 point improvement from Gemini 3.1 Pro Preview, only behind Claude Sonnet 5.5 (max, 64%), Claude Opus 5.5 (max, 60%) and GPT-6 Astra (59%). On AA-Briefcase, it reaches 1494 Elo. This is driven by a 65% rubric pass rate, the highest we have recorded, but lower Analytical Quality (1576 Elo) and Presentation Quality (1308 Elo)
➤ Lowest hallucination rate among leading models: On AA-Omniscience, Gemini 4 Argon has a 15% hallucination rate, the lowest of any model scoring 45+ on the Intelligence Index, compared with 51% for GPT-6 Astra (max) and 54% for GPT-6.1 Sol (max). This means Argon is much more likely to acknowledge when it does not know an answer rather than guess incorrectly. On accuracy, Gemini 4 Argon scores 50%, a 5 point decrease from Gemini 3.1 Pro Preview, and 13 points below GPT-6 Astra (max, 63%). With this slightly lower accuracy, its overall AA-Omniscience score of 42 remains in line with GPT-6 Astra (43) and GPT-6.1 Sol (42)
Key model details:
➤ Context Window: 1M tokens
➤ Multimodality: Text, image, video, and speech input, with text output
➤ Pricing: $4/$20 per 1M input/output tokens at standard pricing, currently discounted 50% to $2/$10. Cached input tokens receive a 95% discount ($0.10 per 1M at discounted pricing), up from 90% on Gemini 3.8 Flash
➤ Long Decode Continuation: We tested Gemini 4 Argon with Long Decode Continuation, a new Gemini API feature that pauses long responses and resumes them across follow-up calls. This lets reasoning run up to 1M output tokens without request timeouts
Today we’re introducing Gemini 4 Argon.
It delivers frontier performance in complex workflows across real-world software engineering, knowledge work, and cybersecurity defense with an industry-leading 1M token output limit.