Google released Gemini 3.8 Live and 3.8 Live Extended Thinking, native speech to speech models built for production voice agents.
Both run in the Gemini Live API and replace cascaded ASR plus LLM plus TTS pipelines with a single model. The defining capability: the conversation never pauses while reasoning and tool calls finish in the background.
✅ #1 on Artificial Analysis Speech to Speech Index, 82.6
✅ Executes tools and API calls in the background mid conversation
✅ Auto switches between 97 supported languages mid conversation
✅ $0.005/min audio input, $0.018/min audio output
Extended Thinking still scores just 35.1% on Sierra's τ-Voice-banking benchmark, so hard multi-step voice workflows are far from solved. Weights are not open; access is API only, with enterprise access in private preview.
Full analysis: https://www.marktechpost.com/2026/09/15/google-releases-gemini-3-8-live-and-3-8-live-extended-thinking-for-production-grade-voice-agents/
Technical details: https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-live-gemini-3-8-live-extended-thinking/
Post #1550
759

- 👍 1
- 🔥 1