Gemini Audio Live last night was SO much fun. Thanks @GoogleDeepMind
for an amazing event!
🍣 Used live transcription to order sushi across English/Japanese
⛳ Gemini coached my golf swing — and my second swing was actually better
🎶 Lyria made a custom song about my night and pressed it onto vinyl
Also loved talking with founders building on these models + celebrating with the Gemini Audio teams, who have been on an incredible launch streak lately.
Seeing transcription, real-time audio, music, and reasoning all come together in actual experiences was very cool. 💫
At the Gemini Audio event tonight from @GoogleDeepMind and WOW what a turn out! What an amazing set of launches from that team this week, can't wait to see what folks make with these new models (I've been having a lot of fun myself)!
Highlight of the night: Ordering sushi in English to someone who only speaks Japanese and having Gemini Live translate our conversation in real time 🤯 ...then actually getting sushi that was flown in fresh from Japan this morning 🛩️🍣
Create and deploy custom audio with our new text-to-speech models:
🔵 Gemini 3.8 Flash TTS: Design unique voices with distinct accents and characteristics.
🔵 Gemini 3.8 Flash-Lite TTS: Built for efficiency and scale, choose from your created styles or our expansive production-ready library.
Help us build the future of voice-based experiences -- come join us for Gemini Audio | At Night in SF on Sept 24!
It'll be an evening dedicated to the next frontier of voice-first AI. Meet the product and research teams behind our latest Gemini Audio models, experience hands-on demos, and network with builders and founders.
More info here:https://t.co/E6s4pLiC2S
Gemini 3.8 Live: built for scale and cost efficiency, combining conversational intelligence with fluid dialogue and visual grounding.
Gemini 3.8 Live Extended Thinking: built for high-complexity tasks, with increased intelligence and multi-step reasoning.
Google has released Gemini 3.8 Live, its new Speech to Speech model, with the Extended Thinking (High) variant debuting at #1 on the Artificial Analysis Speech to Speech Index at 82.6, and #1 on our Tau Voice benchmark implementation at 68.6%
Gemini 3.8 Live is @GoogleDeepMind's successor to Gemini 3.1 Flash Live, a Speech to Speech model that executes tools and API calls in the background while continuing the conversation. It comes in two variants: the standard Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, which supports configurable reasoning effort. We evaluated the standard model and the Extended Thinking variant at High reasoning effort through the Gemini Live API.
Key takeaways:
➤ Speech to Speech Index: Gemini 3.8 Live Extended Thinking (High) debuts at #1 at 82.6, ahead of GPT-Live-1 (Astra, medium) at 81.5, Grok Voice Think Fast 2.0 High at 81.3 and GPT-Live-1 (Sol, low) at 80.1. The standard Gemini 3.8 Live debuts at #5 at 76.0, with both variants up on Gemini 3.1 Flash Live High at 71.5 (+11.1 and +4.5 points). The Index averages Speech Reasoning (Big Bench Audio), Agentic Performance (Tau Voice), Arena Preference and Arena Task Success Rate
➤ Speech Agent Arena: Gemini 3.8 Live ranks #2 in preference at Elo 1083, behind Gemini 3.1 Flash Live (1096) and ahead of GPT-Live-1 (Sol, low) at 1053, and #2 on Task Success Rate at 93.2%, behind Grok Voice Think Fast 2.0 High at 94.6%. Gemini 3.8 Live Extended Thinking (High) trails at Elo 990 with 89.1% task success
➤ Tau Voice: Gemini 3.8 Live Extended Thinking (High) takes the top spot on our Tau Voice benchmark implementation at 68.6%, ahead of GPT-Live-1 (Astra, medium) at 67.9%, GPT-Live-1 (Sol, low) at 59.3% and Grok Voice Think Fast 2.0 High at 56.5% - up from 37.7% for Gemini 3.1 Flash Live High. The standard Gemini 3.8 Live scores 30.1%
➤ Big Bench Audio: Gemini 3.8 Live Extended Thinking (High) scores 97.7% on audio reasoning, ahead of Grok Voice Think Fast 2.0 High at 97.2% and behind Qwen Audio 3.0 Realtime Plus at 99.2%. The standard Gemini 3.8 Live scores 91.7%
➤ Speed: Average Time to First Audio on Big Bench Audio is 1.18 seconds for Gemini 3.8 Live and 1.35 seconds for Extended Thinking (High), both well ahead of Gemini 3.1 Flash Live High (2.99s) and in line with GPT-Live-1 (Sol, low) at 1.24s and GPT-Live-1 (Astra, medium) at 1.34s, though behind Grok Voice Think Fast 2.0 High at 0.70s
➤ Cost: Gemini 3.8 Live costs $0.84 per hour of input audio, the cheapest model in the Index and roughly half the $1.75 of Gemini 3.1 Flash Live High. Extended Thinking (High) costs $3.50 per hour - cheaper than GPT-Live-1 (Sol, low) at $4.47, Grok Voice Think Fast 2.0 High at $4.80 and GPT-Live-1 (Astra, medium) at $5.83, and ~3.1x cheaper than GPT-Realtime-2.1 High at $10.75
See below for more detail ⬇️
We built the 3.8 Live series with enterprise voice agents in mind.
Our latest models can combine language switching, visual input, and async function calling for a seamless conversational experience. Build with it on AI Studio today!
Google has released Gemini 3.8 Live, its new Speech to Speech model, with the Extended Thinking (High) variant debuting at #1 on the Artificial Analysis Speech to Speech Index at 82.6, and #1 on our Tau Voice benchmark implementation at 68.6%
Gemini 3.8 Live is @GoogleDeepMind's successor to Gemini 3.1 Flash Live, a Speech to Speech model that executes tools and API calls in the background while continuing the conversation. It comes in two variants: the standard Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, which supports configurable reasoning effort. We evaluated the standard model and the Extended Thinking variant at High reasoning effort through the Gemini Live API.
Key takeaways:
➤ Speech to Speech Index: Gemini 3.8 Live Extended Thinking (High) debuts at #1 at 82.6, ahead of GPT-Live-1 (Astra, medium) at 81.5, Grok Voice Think Fast 2.0 High at 81.3 and GPT-Live-1 (Sol, low) at 80.1. The standard Gemini 3.8 Live debuts at #5 at 76.0, with both variants up on Gemini 3.1 Flash Live High at 71.5 (+11.1 and +4.5 points). The Index averages Speech Reasoning (Big Bench Audio), Agentic Performance (Tau Voice), Arena Preference and Arena Task Success Rate
➤ Speech Agent Arena: Gemini 3.8 Live ranks #2 in preference at Elo 1083, behind Gemini 3.1 Flash Live (1096) and ahead of GPT-Live-1 (Sol, low) at 1053, and #2 on Task Success Rate at 93.2%, behind Grok Voice Think Fast 2.0 High at 94.6%. Gemini 3.8 Live Extended Thinking (High) trails at Elo 990 with 89.1% task success
➤ Tau Voice: Gemini 3.8 Live Extended Thinking (High) takes the top spot on our Tau Voice benchmark implementation at 68.6%, ahead of GPT-Live-1 (Astra, medium) at 67.9%, GPT-Live-1 (Sol, low) at 59.3% and Grok Voice Think Fast 2.0 High at 56.5% - up from 37.7% for Gemini 3.1 Flash Live High. The standard Gemini 3.8 Live scores 30.1%
➤ Big Bench Audio: Gemini 3.8 Live Extended Thinking (High) scores 97.7% on audio reasoning, ahead of Grok Voice Think Fast 2.0 High at 97.2% and behind Qwen Audio 3.0 Realtime Plus at 99.2%. The standard Gemini 3.8 Live scores 91.7%
➤ Speed: Average Time to First Audio on Big Bench Audio is 1.18 seconds for Gemini 3.8 Live and 1.35 seconds for Extended Thinking (High), both well ahead of Gemini 3.1 Flash Live High (2.99s) and in line with GPT-Live-1 (Sol, low) at 1.24s and GPT-Live-1 (Astra, medium) at 1.34s, though behind Grok Voice Think Fast 2.0 High at 0.70s
➤ Cost: Gemini 3.8 Live costs $0.84 per hour of input audio, the cheapest model in the Index and roughly half the $1.75 of Gemini 3.1 Flash Live High. Extended Thinking (High) costs $3.50 per hour - cheaper than GPT-Live-1 (Sol, low) at $4.47, Grok Voice Think Fast 2.0 High at $4.80 and GPT-Live-1 (Astra, medium) at $5.83, and ~3.1x cheaper than GPT-Realtime-2.1 High at $10.75
See below for more detail ⬇️
Today we’re launching Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking.
They bring asynchronous tool use and deeper multi-step reasoning to complex tasks while keeping the conversation flowing.
Very proud of the Gemini Audio team and the many teams who made this possible!
We’re introducing Gemini 3.8 Live and 3.8 Live Extended Thinking – our best conversational AI.
The models talk, think, and handle tasks in the background without breaking your flow. 🧵
We’re introducing Gemini 3.8 Live and 3.8 Live Extended Thinking – our best conversational AI.
The models talk, think, and handle tasks in the background without breaking your flow. 🧵
Say hello to Gemini 3.5 Transcribe!
- Build apps that understand user speech / intent, even w/ multiple speakers!
- Auto-detection of 85+ languages out of the box
- Custom vocab adaptation for specialized jargon... SGTM:)
API available now in @GoogleAIStudio and Gemini Enterprise, or try it in the Gemini app on macOS or Rambler on Android!
More details: https://t.co/AduutCb3M7
Introducing Gemini 3.5 Flash Live Translate, our real time speech to speech translation model which supports more than 70 languages (both in and out), and is so natural.
It is available in the Gemini API, AI Studio, & Google Translate right now + coming soon to Google Meet!!
Typing is great, but speaking is faster. Bring Gemini directly into your desktop workflow to seamlessly reformat text between apps using just your voice.
Watch it pull raw details from a plain document and transform them into a festive, emoji-filled email invite exactly where you need it.
We’re expanding Google @Antigravity’s agentic surfaces and features — and they’re all available now for you to try.
🔹 Antigravity CLI
🔹 Antigravity SDK
🔹 Native voice support with Gemini Audio models
🔹 @Antigravity 2.0 desktop application
🔹 Integrations with @GoogleAIStudio, @Android, @Firebase, and the web
#GoogleIO
The stage is set. The tech is ready. Are you? 🚀
Join us tomorrow for #GoogleIO as we unveil the breakthroughs, tools, and innovations shaping the future of AI.
Tune in live right here on @X from 10am PT: https://t.co/u4s3nkfrlT
Today at the @Android Show (I/O edition) we announced Gemini Intelligence - bringing the best of Gemini to our most advanced devices.
Automate multi-step tasks across apps and Chrome, fill out forms in a single tap, turn spoken thoughts into polished text with Rambler, build custom widgets & loads more.
Our most expressive and steerable TTS model yet! Designed to give builders granular control over AI-generated speech, Gemini 3.1 Flash TTS is really fun to play with! Available in preview today - for devs via the Gemini API & @GoogleAIStudio + for enterprises on Vertex AI
Launched Gemini 3.1 Flash Live. It’s capable of handling the nuances of live speech, like tone and interruptions, that are critical for real-world interactions. You can experience it on Gemini Live and Search Live!
Say hello to Gemini 3.1 Flash Live. 🗣️
Our latest audio model delivers more natural conversations with improved function calling – making it more useful and informed. Here’s what’s new 🧵