Google launched two native speech-to-speech artificial intelligence models on Tuesday, introducing Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking to handle spoken conversations with simultaneous background task processing.
Key points
- Google announced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on September 15, 2026.
- The models run tools and API calls in the background while continuing natural voice conversations.
- Gemini 3.8 Live Extended Thinking scored 82.6 to lead Artificial Analysis' Speech to Speech Quality Index.
- The Live API costs $0.005 per minute for audio input and $0.018 per minute for audio output.
- Standard Gemini 3.8 Live now powers conversational voice queries in Google Search Live globally.

Announced by Gemini Audio Team engineers Tom Ouyang and Malini Jaganathan, the releases expand Google’s voice portfolio across the Gemini app, Google Search, developer APIs, and Google Workspace. The release aims to remove common delays in voice assistants by allowing models to reason and call software tools while speaking, rather than pausing conversational audio output.
Dual Architectures for Speed and Multi-Step Tasks
Google developed the two models for distinct operational needs. The standard Gemini 3.8 Live focuses on lower operating costs and fast response times for high-volume deployments. The larger Gemini 3.8 Live Extended Thinking model targets complex, multi-step tasks that require internal reasoning before providing answers.
Unlike traditional voice assistants that switch between speech recognition, text analysis, and speech generation pipelines, both models operate natively on incoming audio signals. According to Google’s announcement, this architecture enables the models to register user interruptions, paralanguage, and tone variations without resetting conversational context.
Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking are our most advanced live dialogue models yet. Major upgrades in intelligence and parallel reasoning make them more intuitive to collaborate with and use to execute complex tasks using your voice.
Parallel Tool Calling and Visual Inputs
Both models can trigger asynchronous application programming interface calls while actively speaking. During extended operations, Gemini 3.8 Live Extended Thinking uses conversational filler phrases, such as explaining that it is looking up data, while executing the required software functions in the background.
In technical demonstrations, Google showed the extended model generating React code from physical paper sketches and booking restaurant reservations through external services during a voice exchange. Both models also analyze visual inputs at up to one frame per second, allowing users to show camera feeds and receive spoken answers in real time. The system automatically switches between 97 supported spoken languages without requiring users to alter input settings.
Benchmark Scores and Safety Features
Google reported benchmark results placing the models at the top of several voice evaluation indexes. On Artificial Analysis’ Speech to Speech Quality Index, Gemini 3.8 Live Extended Thinking achieved an 82.6 score to rank first overall. The standard Gemini 3.8 Live model placed second in the Speech Agent Arena.
- Artificial Analysis Speech to Speech Quality Index: 82.6
- Big Bench Audio: 97.7%
- τ-Voice benchmark: 68.6%
- Sierra τ-Voice-banking benchmark: 35.1%
For synthetic audio safety, Google confirmed that all audio generated by both models contains imperceptible SynthID digital watermarks. These embedded signals identify model-generated audio files during algorithmic inspection without impacting human auditory quality.
Pricing, Platform Integration, and Availability
Standard Gemini 3.8 Live is now live globally in Search Live within the Google app, according to Rajan Patel, vice president of engineering for Google Search. Google Workspace business customers also received access to Extended Thinking features within Google Docs, Gmail, and Google Keep.
Developers can access the models through the Gemini Live API over WebSocket connections in Google AI Studio. Technical specifications and pricing include:
- Input Context Window: 131,072 tokens
- Output Window: 65,536 tokens
- Audio Format: 16-bit PCM input at 16kHz; 24kHz audio output
- API Pricing: $0.005 per minute for audio input; $0.018 per minute for audio output
Media streaming platforms including LiveKit, Vercel, Pipecat, and Agora have added technical support for the API, while enterprise partners Salesforce, Genspark, and Lumeris have begun deploying the systems.



