A computer scientist from Bangkok, Thailand. Research Scientist @GoogleDeepmind. Previous: @GoogleAI, @ShaLabUSC at @CSatUSC, & @BrownCSDept. He/Him. 🏳️🌈
Looking for a multilingual visual question answering dataset to test your model? Here's MaXM. English, French, Hindi, Hebrew, Romanian, Thai, and Chinese!
https://t.co/8xkz2ZDfCV
1/ Today at #GoogleIO, we’re releasing Gemini 3.5, our latest family of models combining frontier intelligence with action.
We’re starting by releasing 3.5 Flash, which is built to help you execute complex, long-horizon agentic workflows.
Gemini 3.5 Flash is our strongest model for coding and agent https://t.co/m62cBJhIjJ outscores 3.1 Pro on agentic and coding benchmarks like Terminal-Bench and MCP Atlas, while running 4x faster than other frontier models.
Used in Google Antigravity, 3.5 Flash is even further optimized to be up to 12x faster. It’s a powerful engine to deploy sub-agents that collaborate, run high-frequency iterative loops, and solve real-world problems at scale.
Some highlights we’re excited about 🔽
If you want to see our multimodal models in full flow, check out how we bring Roman aqueducts to life with custom interactive images and timelines. Honestly, it’s such a fun way to dive into history (and other topics) without a dusty textbook! 🏛️✨
One of my favorite projects I get to be a part of! 🚀 Josh Woodward just showed off our new Neural Expressive design on the Google I/O 2026 stage.
🔗https://t.co/aACHDluIGk
A conversation with some of the research folks behind nano-banana 🍌 (aka Gemini 2.5 Flash Image) on how we got here, what it took to build this model, and where we go next!
So much fun to hang with: @19kaushiks@robertriachi@m__dehghani@nbrichtova
Image generation with Gemini just got a bananas upgrade and is the new state-of-the-art image generation and editing model. 🤯
From photorealistic masterpieces to mind-bending fantasy worlds, you can now natively produce, edit and refine visuals with new levels of reasoning, control and creativity.
A quick dive into Gemini 2.5 Flash’s capabilities 🧵
Our Aeneas AI model gives historians valuable new insights into ancient inscriptions & ancient history that may have taken years to uncover otherwise. Published in @Nature today: https://t.co/U9ZcIKcfaF
Big update to our MathArena USAMO evaluation: Gemini 2.5 Pro, which was released *the same day* as our benchmark, is the first model to achieve non-trivial amount of points (24.4%). The speed of progress is really mind-blowing.
I'm delighted to have joined my good friend and colleague @NoamShazeer for a 2+hour conversation with @dwarkesh_sp about a wide range of topics (early Google, ML hardware, training trillion token LLMs in 2007, model sparsity, continual learning, and more).
Thanks for a fantastic conversation, Noam and Dwarkesh! 🙏
Gemini 2.0 Flash Experimental has the ability to produce native audio in a variety of styles and languages - all from scratch. 🗣️
Here’s how this is different to traditional text-to-speech systems ↓ https://t.co/FRWb3q3KHe
Gemini 2.0 Flash ⚡️ has arrived!
2.0 Flash > 1.5 Pro (again!) 📈
Interacts with a browser 🤖
Native image generation 🖼️
and much more!
Try it out https://t.co/IMN3bqKgJS
As a preview of what is possible, wishing you all a Drastic Holiday powered by 2.0!
A super useful blog.
"7 examples of Gemini’s multimodal capabilities in action"
1. Detailed Image Descriptions - Can analyze and describe images, adjusting style and format based on prompts
2. Long PDF Understanding - Processes 1000+ page PDFs, including tables, layouts, charts, diagrams, and handwritten text
3. Real World Document Reasoning - Extracts information from receipts, labels, signs, notes, and whiteboard sketches
4. Webpage Data Extraction - Extracts structured data from webpage screenshots, including text and visual content
5. Object Detection - Detects objects and generates bounding box coordinates in images
6. Video Summarization - Processes 90-minute videos, generating transcripts, summaries, and answering questions
7. Video Information Extraction - Extracts structured data from videos for cataloging and entity detection, though currently limited by 1FPS sampling
A nice new benchmark for long video understanding by Tobias Weyand @0xtob and others. This is likely to be one of the new frontiers of capabilities for large-scale multimodal models, and it's great to have a new benchmark to assess others in this area.
The #AlphaFold 3 model code and weights are now available for academic use. We @GoogleDeepMind are excited to see how the research community continues to use AlphaFold to address open questions in biology and new lines of research.
https://t.co/kVB9hWJZTI
🚀 Join the Gemini Multilinguality team @GoogleDeepMind 🌐 We’re looking for researchers passionate about making LLMs helpful for all. Dramatically improve model quality, coverage, and cultural relevance across hundreds of languages. #NLProc#MultilingualAI#i18n#LLMs https://t.co/RzwiRM5zWr
🚀 New Scale Product 🚀
Today, we're launching Expert Match!
Expert Match enables AI developers to connect with Experts (doctors, lawyers, PhDs) to collaborate on their AI projects:
🔎 Advanced Search
💼 View Qualifications
⭐️ Streamlined Screening + Selection
Read more 👇
All languages covey information at a similar rate when spoken (39bits/s).
Languages that are spoken faster have less information density per syllable!
One of the coolest results in linguistics.