I work at Google DeepMind and yes I have tried the model known externally as “Argon”.
I can say this:
I don’t reach for Opus anymore for any work (like I did when 3.8 Flash was our top model).
I hope that it will soon be generally available for everyone to enjoy!
I do not always use Gemini Argon, even though I have unlimited & free access to it.
Working at Google DeepMind, I choose Argon when I need to perform the initial design phase, as well as critical de-risking exploration/POCs. The goal is to front-load the highly cognitive work and go/no-go decisions, the critical stuff.
Later when implementing discrete sections of work (and splitting massive commits into reviewable chunks) I will use Gemini 3.8 Flash...because it is just so darn fast! At this point the project is well-defined and scoped.
Another example where I won't use Argon is agentic testing:
"Validate that the UI changes work end-to-end by recording me a video demo of the entire feature with local changes"
Why would I wait for a big and brainy model to do this easy work if Flash does the job and is much faster at that?
Same goes for my agentic review tooling that main agents call.
For example, my "minimalist" SKILL scans a given technical doc for complicated sentence structures; the model does not need to be a genius, it just needs to follow a straightforward SKILL I had written for it, which Flash can do just fine (and fast).
Of course, some agentic tooling I have needs a big model for that extra reasoning boost. Take my "abstraction-police" SKILL meant to review written code and suggest better abstractions ("borrowed" from @karpathy 😁).
I pair this skill with Argon (before I used Opus 5.5/5/4.8/etc), as the ability to navigate ambiguity and make the right calls is paramount.
In my opinion, it is important to pick the right tool for the job.
Gemini 4 argon is frontier level at adapting to new information, trial and error, and overall generalizing. It is just not as fast and learning efficient as the top 2 models.
@RudyXFantasy I do notice the model is much better at this, and you're right we've been having enough checkpoints that it's a gradual change, not a sudden difference like y'all will see with 3.1 Pro -> 4 Argon.
I find it peculiar that many books I find inside Google DeepMind’s main office are old Soviet math textbooks.
Here are works on Ordinary Differential Equations, Group Theory, and Abstract Algebra.
Either a coincidence or there’s a mysterious explanation for it 🤷♂️
I do appreciate having them around though, always a pleasant read.
AI performing in Google3 feels like the ultimate final boss test.
The monorepo is full of bespoke and highly complex software that is no model is trained on. The models have to rely on their true reasoning capabilities to solve problems inside Google.
@file_mutex With a ton of effort and swe-hours.
It’s really not as uniform as one might think. But there are some commonalities like some shared libraries and tooling.
@Greg_GL_87 google3 codebase is not used to train the model, that would leak out our internal codebase to everyone else.
We tried to have a separate “Google-internal-knowledge” model at first, but it became pretty apparent that we need to use what we ship, not a separately-trained variant.
I work at Google DeepMind and yes I have tried the model known externally as “Argon”.
I can say this:
I don’t reach for Opus anymore for any work (like I did when 3.8 Flash was our top model).
I hope that it will soon be generally available for everyone to enjoy!
Figma doesn’t want you to use MCP
Amazon doesn’t want you to shop with agents
Reddit won’t let Claude access threads
X won’t let ChatGPT read tweets
It’s starting.
This happened with APIs 8 years ago.
Indeed!
Context:
Megha personally ran many of the evals you are seeing on the Gemini 4 Argon model card today!
Lots of hard work goes into making a release happen, thank you for giving it your all 😌
And thank you for organizing our Boba Tea surprise today 🧋🧋