Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking
By Jakub Antkiewicz
•2026-09-16T13:08:10Z
Google Focuses on Fluidity and Background Tasking with New Voice Models
Google has released two new advanced dialogue models, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, aimed at making voice interactions with AI more natural and capable. The models are designed to handle complex user requests that involve real-time visual context and background tool execution without halting the conversational flow. This move signals a deliberate focus on the user experience of voice agents, shifting the industry's attention toward fluid, uninterrupted interaction rather than just turn-based responses.
The two models serve distinct purposes, with Gemini 3.8 Live engineered for scale and cost efficiency, while 3.8 Live Extended Thinking is built for higher-complexity tasks requiring multi-step reasoning. According to Google, the Extended Thinking model has achieved top performance on industry benchmarks, including the #1 spot on Artificial Analysis' Speech to Speech Quality Index (82.6) and strong results on agentic task benchmarks like τ-Voice (68.6%). Both models introduce significant technical capabilities for voice agents.
- Parallel Processing: Executes tools and API calls in the background while continuing the live conversation, using verbal cues like “Let me check that…” to maintain flow.
- Visual Grounding: Processes visual inputs in near real-time to enrich conversations with on-screen context.
- Multilingual Support: Automatically detects and transitions between 97 supported languages mid-conversation.
- Transparent Generation: All generated audio is watermarked with SynthID to ensure AI-generated content remains detectable.
The release impacts the entire AI ecosystem, from individual users to large enterprises. The models are rolling out immediately via the Gemini API and Google AI Studio for developers, with integrations planned for partners like Agora, LangChain, and Vercel. Enterprise customers like Salesforce and Genspark are also leveraging the new capabilities. For consumers, the technology will be integrated into the Gemini app, Search Live, and across Google Workspace applications, enabling users to execute complex commands using only their voice.
With this launch, Google is making a strategic bet that the future of AI assistants hinges not just on raw intelligence, but on the agent's ability to seamlessly manage complex, multi-modal tasks in the background without breaking the cadence of a natural human conversation.