@flowersslop AGI doesn't need to get 100% on all of it. It just needs to exhibit all human cognitive abilities like continual learning, sample efficiency, to learn through any modality, the ability to understand the physical world, the ability to learn any skill with certain amount of data
More big news from @GoogleDeepMind: Gemini 4 Argon (High) is #8 in Agent Arena with a +7.92% net improvement score, and has reshaped the Pareto frontier with a $0.62 cost per task! See its placement below.
Gemini 4 Argon (High) is a 4.96 percentage point improvement over Gemini 3.8 Flash (High), at #19 with +2.96% net improvement.
By key signals, Gemini 4 Argon (High) stands out in:
- #1 in Steerability with +15.88% (the model’s ability to course-correct when you push back)
- #2 in Confirmed Success with +14.15% (explicit user feedback that the task worked)
- #4 in Praise vs Complaint with +27.74% (implicit sentiment in user reactions)
By category Gemini 4 Argon (High) is especially strong in Chat, landing at #3 with +11.58% net improvement.
With 3k real-world agentic sessions so far, this score is preliminary. Stay tuned as more traces come in from our global community of users.
Congrats to the @GoogleDeepMind team on this release!
@jpshrodinger The usage till now was based on flash models. This should consume 2-3 times more usage than flash.
Whatever it might be, I really loved the generous usage limits in the Google AI Pro plan.
More big news from @GoogleDeepMind: Gemini 4 Argon (High) is #8 in Agent Arena with a +7.92% net improvement score, and has reshaped the Pareto frontier with a $0.62 cost per task! See its placement below.
Gemini 4 Argon (High) is a 4.96 percentage point improvement over Gemini 3.8 Flash (High), at #19 with +2.96% net improvement.
By key signals, Gemini 4 Argon (High) stands out in:
- #1 in Steerability with +15.88% (the model’s ability to course-correct when you push back)
- #2 in Confirmed Success with +14.15% (explicit user feedback that the task worked)
- #4 in Praise vs Complaint with +27.74% (implicit sentiment in user reactions)
By category Gemini 4 Argon (High) is especially strong in Chat, landing at #3 with +11.58% net improvement.
With 3k real-world agentic sessions so far, this score is preliminary. Stay tuned as more traces come in from our global community of users.
Congrats to the @GoogleDeepMind team on this release!
Big news: Gemini 4 Argon (High) by @GoogleDeepMind just landed #1 in Text Arena with 1525 pts, and #8 in Code Arena: WebDev with 1679 pts!
This release has reshaped the Text Arena Pareto frontier with a blended $8/MToken! Gemini 4 Argon (High) is now the most cost efficient model, see its placement on Pareto frontier below.
In the Text Arena, Gemini 4 Argon (High) ranks #1 in Coding, Hard Prompts, Instruction Following, Longer Query, and Creative Writing. It also leads every occupational domain evaluated, with additional #1 spots in English, Non-English, Chinese, and Russian.
This model is +20 points above the #2 ranked Claude Opus 4.6 (High), and a huge leap from Google’s previous release, Gemini 3.8 Flash (High) at #11!
In Code Arena: WebDev, Gemini 4 Argon (High) gained +96 points from Gemini 3.8 Flash (High), and went from #29 to #8.
Congrats to the @GoogleDeepMind team on this impressive frontier release!
On Artificial analysis - high settings G4A
#8 overall
It scores 1 point below Opus 5.5 (high) while being 11% more expensive. (Meaning doesn't land on Pareto Frontier)
> It makes the Pareto Frontier on Coding Agent Index at 64. (Behind Opus 5.5 Max - 66)
On Artificial analysis - high settings G4A
#8 overall
It scores 1 point below Opus 5.5 (high) while being 11% more expensive. (Meaning doesn't land on Pareto Frontier)
> It makes the Pareto Frontier on Coding Agent Index at 64. (Behind Opus 5.5 Max - 66)
@haider1 This is true across all labs. For every lab their bigger model is more token efficient than their smaller ones. At this point it's like a law of sorts. Nothing suprising here. It still doesn't beat OAI models in efficiency cuz they are in a league to their own rn.
TPS in Antigravity have dropped. Currently less than 1/3rd of the previous speeds.
Obviously this has to do something with the launch of Gemini 4 Argon. Maybe something else