Gemini 3 models from @Google@GoogleDeepMind have made a significant 2X SOTA jump on ARC-AGI-2 (Semi-Private Eval)
Gemini 3 Pro:
31.11%, $0.81/task
Gemini 3 Deep Think (Preview):
45.14%, $77.16/task
We’ve been intensely cooking Gemini 3 for a while now, and we’re so excited and proud to share the results with you all. Of course it tops the leaderboards, including @arena, HLE, GPQA etc, but beyond the benchmarks it’s been by far my favourite model to use for its style and depth, and what it can do to help with everyday tasks.
Introducing Gemini 3 ✨
It��s the best model in the world for multimodal understanding, and our most powerful agentic + vibe coding model yet. Gemini 3 can bring any idea to life, quickly grasping context and intent so you can get what you need with less prompting.
Find Gemini 3 Pro rolling out today in the @Geminiapp and AI Mode in Search. For developers, build with it now in @GoogleAIStudio and Vertex AI.
Excited for you to try it!
An advanced version of Gemini with Deep Think has officially achieved gold medal-level performance at the International Mathematical Olympiad. 🥇
It solved 5️⃣ out of 6️⃣ exceptionally difficult problems, involving algebra, combinatorics, geometry and number theory. Here’s how 🧵
A cool detail from the Gemini 2.5 paper: I liked their fault-tolerant scheduling system where they train on the ~97% remaining TPU slices when one slice goes down, instead of waiting for another healthy slice to become available for scheduling.
It's a cool throwback to the early DNA of Google where they used inexpensive commodity hardware and wrote software to make it fault-tolerant
https://t.co/syooPlOqck
It’s been an amazing few months of relentless building, shipping, and optimising our models incorporating your feedback. Excited for more users and developers to try out the incredible Gemini 2.5 series!
excited to finally share on arxiv what we've known for a while now:
All Embedding Models Learn The Same Thing
embeddings from different models are SO similar that we can map between them based on structure alone. without *any* paired data
feels like magic, but it's real:🧵
Today Demis announced Deep Think which marks our progression to greater test-time compute and stronger reasoning capabilities in Gemini 💎
Highlighting USAMO which is a very challenging set of held-out math problems, we're now at 49% accuracy. This is equivalent to the top quartile of entrants. Progress is fast month-by-month, back in March we were at 24% (2.5 Pro preview), and our best Thinking model from Jan (2.0 Flash Thinking) was at 6%.
🚨Breaking: @GoogleDeepMind’s latest Gemini-2.5-Pro is now ranked #1 across all LMArena leaderboards 🏆
Highlights:
- #1 in all text arenas (Coding, Style Control, Creative Writing, etc)
- #1 on the Vision leaderboard with a ~70 pts lead!
- #1 on WebDev Arena, surpassing Claude for the first time
This is the first-ever sweep across text, vision, and WebDev by any model!🥇
Huge congrats to @GoogleDeepMind on this incredible breakthrough!
Very excited to share the best coding model we’ve ever built! Today we’re launching Gemini 2.5 Pro Preview 'I/O edition' with massively improved coding capabilities. Ranks no.1 on LMArena in Coding and no.1 on the WebDev Arena Leaderboard.
It’s especially good at building interactive web apps - this demo shows how it can be helpful for prototyping ideas. Try it in @GeminiApp, Vertex AI, and AI Studio https://t.co/7FbP3R1cmF
Enjoy the pre-I/O goodies !
Gemini 2.5 Flash just dropped. ⚡
As a hybrid reasoning model, you can control how much it ‘thinks’ depending on your 💰 - making it ideal for tasks like building chat apps, extracting data and more.
Try an early version in @Google AI Studio → https://t.co/sYsLJyIhCz
Gemini 2.5 Pro physics simulations in Three.js!
All of these started out as "one-shot prompts" but I continued to query Gemini for better results.
Clone with GitHub below 👇
#threejs#Physics
Gemini 2.5 Pro is taking off 🚀🚀🚀
The team is sprinting, TPUs are running hot, and we want to get our most intelligent model into more people’s hands asap.
Which is why we decided to roll out Gemini 2.5 Pro (experimental) to all Gemini users, beginning today.
Try it at no cost at https://t.co/lc7BAJqH5u
Introducing Gemini 2.5 Pro Experimental.
The 2.5 series marks a significant evolution: Gemini models are now fundamentally thinking models.
This means the model reasons before responding, to maximize accuracy -- and it’s our best Gemini model yet.
Blog - https://t.co/Gx35wlkomq
BREAKING: Gemini 2.5 Pro is now #1 on the Arena leaderboard - the largest score jump ever (+40 pts vs Grok-3/GPT-4.5)! 🏆
Tested under codename "nebula"🌌, Gemini 2.5 Pro ranked #1🥇 across ALL categories and UNIQUELY #1 in Math, Creative Writing, Instruction Following, Longer Query, and Multi-Turn!
Massive congrats to @GoogleDeepMind for this incredible Arena milestone! 🙌
More highlights in thread👇