← All guides

Home Wi-Fi

Two new Google Gemini models have launched, capable of executing complex tasks via voice.

Google Gemini has launched two new models capable of near-instantaneous reasoning and engaging in spoken, AI-driven conversations. These include Gemini 3.8 Live—built for scalability and cost-efficiency—which combines conversational intelligence with fluid dialogue and visual grounding; and a version designed for highly complex tasks, featuring enhanced intelligence and multi-step reasoning capabilities. These models can be integrated into the Gemini app, Google Workspace, and Google Search, enabling smoother, more collaborative conversations with Gemini and allowing users to easily handle complex tasks using just their voice.

Despite calls from AI giants to slow down the pace of language model development, Google has updated its “Flash” model series three times in just the past few weeks. Following the early September launch of Gemini 3.8 Flash and a version specialized for cybersecurity, the company released Gemini 3.8 Live within three weeks—underscoring the pressure Google faces to avoid falling behind in the race to develop language models.

Google claims that Gemini 3.8 Live Extended Thinking delivers enterprise-grade task completion capabilities and intelligence. It ranked first on Artificial Analysis’s Speech-to-Speech Quality Index with a score of 82.6 and led the τ-Voice and Sierra τ-Voice-banking benchmarks with a 68.1% completion rate. Additionally, the model boasts powerful reasoning capabilities while remaining cost-competitive.

Gemini 3.8 Live ranks second in the voice agent sector and offers exceptional cost-effectiveness, providing developers and enterprises with a powerful, efficient, and scalable model.

Enhanced with advanced voice conversation capabilities, Gemini 3.8 Live processes visual input in near real-time and incorporates contextual information into conversations to deliver more useful responses. It supports the automatic detection of and switching between 97 languages ​​during interactions. The model can execute tools and API calls in the background while maintaining the conversation; it acknowledges requests and continues chatting as background tasks are completed.