It’s over. Gemini 3 Deep Think achieved a 3300 rating on LiveCodeBench Pro almost surpassing all humans (99.99%) and is leading GPT-5.2 by a massive margin of 1000 points. Gemini is insanely strong!
Link:https://t.co/jC5SHHeSEB
Proud to be part of the Deepthink team as we continue to push the frontiers of AI. For the first time, our model has reached a Codeforces ELO of 3455, breaking into the top 10 of the human rankings!
The latest Deep Think moves beyond abstract theory to drive practical applications.
It’s state-of-the-art on ARC-AGI-2, a benchmark for frontier AI reasoning.
On Humanity’s Last Exam, it sets a new standard, tackling the hardest problems across mathematics, science, and engineering — making it a genuine collaborator for heavy-duty analysis.
It achieved an Elo of 3455 on Codeforces, demonstrating the ability to solve complex, real-world coding tasks - while earning gold medal-level results on the written portion of the 2025 Physics and Chemistry Olympiads.
In our latest work, we show autoregressive pixel modeling scales predictably, but is massively compute-hungry compared to text LLMs. But with 4-5x annual compute growth, raw pixel modeling will become feasible in ~5 years. So see you in 2030!
For years, RAW pixel space pretraining has been sidelined: too compute-expensive.
Our new @GoogleDeepMind paper 📜 dives into the scaling trends of raw pixel models to answer the question “how far are we from scaling up next-pixel prediction?”
https://t.co/OU1u9TnTIR
Forecast: Raw next-pixel modeling will reach competitive ImageNet classification (>80% top1 accuracy) and generation metrics (90 Fr’echet Distance) in five years!
Threads 👇
Gemini 3 Deep Think is next level. Deep Think was the the engine behind our gold medal-level wins at IMO and ICPC, and now powers an even stronger version of Gemini 3. SOTA above SOTA. More to come soon!
Incredibly proud to be part of the team that took home Gold medal at the ICPC World Finals with Gemini! Conquering the complexity of competitive coding is a huge win, and it fuels our drive toward the ultimate challenge: Building AGI! 🚀
Very excited to see our Gemini models getting better and better at coding! An advanced version of Gemini 2.5 Deep Think at the 2025 International Collegiate Programming Contest (ICPC) World Finals achieved gold-medal level performance! 🎉
https://t.co/yCQKOagnkm
Following its IMO gold-level win, @GoogleDeepMind is sharing Gemini Deep Think with mathematicians for feedback. Excited to see what they discover! 🧠
Plus, an updated Gemini 2.5 Deep Think is now rolling out for Google AI Ultra subscribers.
Learn more: https://t.co/W6vWpFvxIu
Amazing progress in reasoning! 🚀 Gemini 2.5 Pro Deep Think hitting 49.4% on USAMO – a feat I'd have considered impossible just a couple of years ago – & Gemini 2.5 Flash achieving 1424 Elo are huge leaps. So proud our team's research ideas contributed to this moment!
https://t.co/PWJ4O5HWBk
Today we’re sharing an early look at our latest Gemini update for I/O!
Introducing the updated Gemini 2.5 Pro (I/O edition), which ranks #1 on WebDev Arena and surpasses our previous 2.5 Pro model by +147 Elo points. 🏆
https://t.co/jIPslRklGd
Check out our strongest model Gemini 2.5! Happy to have been driving its multimodal reasoning capability and now we are also SOTA on MMMU for the first time! https://t.co/CL6r9JLFod
🥁Introducing Gemini 2.5, our most intelligent model with impressive capabilities in advanced reasoning and coding.
Now integrating thinking capabilities, 2.5 Pro Experimental is our most performant Gemini model yet. It’s #1 on @lmarena_ai leaderboard. 🥇
We’ve been thrilled by the positive reception to Gemini 2.0 Flash Thinking we discussed in December.
Today we’re sharing an experimental update (gemini-2.0-flash-thinking-exp-01-21) with improved performance on math, science, and multimodal reasoning benchmarks 📈:
• AIME: 73.3%
• GPQA: 74.2%
• MMMU: 75.4%
Gemini's deeper thinking unlocks new possibilities! Proud to be part of it and have contributed to its enhanced multimodal reasoning. Give it a try! https://t.co/7akhbUKgWD
Introducing Gemini 2.0 Flash Thinking, an experimental model that explicitly shows its thoughts.
Built on 2.0 Flash’s speed and performance, this model is trained to use thoughts to strengthen its reasoning.
And we see promising results when we increase inference time computation!
1/3 Today, an anecdote shared by an invited speaker at #NeurIPS2024 left many Chinese scholars, myself included, feeling uncomfortable. As a community, I believe we should take a moment to reflect on why such remarks in public discourse can be offensive and harmful.
Congratulations to Ilya Sutskever, Oriol Vinyals and Quoc V. Le for their paper “Sequence to Sequence Learning with Neural Networks", for winning the #NeurIPS2024 Test of Time award!
Award announcement: https://t.co/eH5YOfkbD9
Read the paper: https://t.co/XX5n2hF2DL
What a way to celebrate one year of incredible Gemini progress -- #1🥇across the board on overall ranking, as well as on hard prompts, coding, math, instruction following, and more, including with style control on.
Thanks to the hard work of everyone in the Gemini team and elsewhere at Google! 🎊
Woah, huge news again from Chatbot Arena🔥
@GoogleDeepMind’s just released Gemini (Exp 1121) is back stronger (+20 points), tied #1🏅Overall with the latest GPT-4o-1120 in Arena!
Ranking gains since Gemini-Exp-1114:
- Overall #3 → #1
- Overall (StyleCtrl): #5 -> #2
- Hard Prompts (StyleCtrl): #3 → #1
- Coding: #3 → #1
- Vision: #1
- Math: #2 → #1
- Creative Writing #2 → #1
Congrats again @GoogleDeepMind! The LLM race is on fire — progress is now measured in days!
See more analysis below👇
Introducing #HaloQuest, our latest effort towards improving hallucination in multimodal foundation models! We hope that HaloQuest, as both a challenging benchmark and an open-sourced dataset, will enable the field's progress in advanced reasoning! Also glad that HaloQuest played a small part in the great news yesterday that #Gemini is now #1 on lmsys & accepted at #ECCV2024.
Project page: https://t.co/xh33sLQqlF
Paper: https://t.co/zZ6pcQtptP
Github: https://t.co/ZafyKC7dpS
Exciting News from Chatbot Arena!
@GoogleDeepMind's new Gemini 1.5 Pro (Experimental 0801) has been tested in Arena for the past week, gathering over 12K community votes.
For the first time, Google Gemini has claimed the #1 spot, surpassing GPT-4o/Claude-3.5 with an impressive score of 1300 (!), and also achieving #1 on our Vision Leaderboard.
Gemini 1.5 Pro (0801) excels in multi-lingual tasks and delivers robust performance in technical areas like Math, Hard Prompts, and Coding.
Huge congrats to @GoogleDeepMind on this remarkable milestone!
Gemini (0801) Category Rankings:
- Overall: #1
- Math: #1-3
- Instruction-Following: #1-2
- Coding: #3-5
- Hard Prompts (English): #2-5
Come try the model and let us know your feedback!
More analysis below👇
We’re presenting the first AI to solve International Mathematical Olympiad problems at a silver medalist level.🥈
It combines AlphaProof, a new breakthrough model for formal reasoning, and AlphaGeometry 2, an improved version of our previous system. 🧵 https://t.co/SYaLPSbIyj
Big news – Gemini 1.5 Flash, Pro and Advanced results are out!🔥
- Gemini 1.5 Pro/Advanced at #2, closing in on GPT-4o
- Gemini 1.5 Flash at #9, outperforming Llama-3-70b and nearly reaching GPT-4-0125 (!)
Pro is significantly stronger than its April version. Flash’s cost, capabilities, and unmatched context length make it a market game-changer!
Huge congrats to @GoogleDeepMind on the incredible
Gemini launches! Can't wait to see what new applications Gemini unlocks!
More breakdown analysis below👇
Gemini 1.5 Model Family: Technical Report updates now published
In the report we present the latest models of the Gemini family – Gemini 1.5 Pro and Gemini 1.5 Flash, two highly compute-efficient multimodal models capable of recalling and reasoning over fine-grained information from millions of tokens of context, including multiple long documents and hours of video and audio.
Our latest report details notable improvements in Gemini 1.5 Pro within the last four months.
Our May release demonstrates significant improvement in math, coding, and multimodal benchmarks compared to our initial release in February.
Furthermore, the 1.5 Pro Model is now stronger than 1.0 Ultra.
The latest Gemini 1.5 Pro is now our most capable model for text and vision understanding tasks, surpassing 1.0 Ultra on 16 of 19 text benchmarks and 18 of 21 of the vision understanding benchmarks. The table below highlights the improvement in average benchmark performance for different categories in 1.5 Pro since Feb, and also shows the strength of the model relative to the 1.0 Pro and 1.0 Ultra models. The 1.5 Flash model also compares very well against the 1.0 Pro and 1.0 Ultra models.
One clear example of this can be seen on MMLU
On MMLU we find that 1.5 Pro surpasses 1.0 Ultra in the regular 5-shot setting scoring 85.9% versus 83.7%. However with additional inference compute, via majority voting on top of multiple language model samples, we can get a performance of 91.7% versus Ultra’s 90.0%, which extends the known performance ceiling of this task.
@OriolVinyalsML and I are very proud of the whole Gemini team, and it’s fantastic to see this progress and to share these highlights from our Gemini Model Family.
Read the updated report here: https://t.co/CTzTHND4nQ