Google Makes Gemini 3.8 Live Generally Available With Real-Time AI Avatars
Google has released Gemini 3.8 Live with Live Avatar for enterprise use, combining low-latency voice conversations with synchronized video avatars, live visual understanding, tool calling and multilingual speech. The system is aimed at customer service, claims intake, shopping assistants and other applications where an AI agent must see, speak and act in real time.
By StoryBreak
Published September 25, 2026 at 12:59 AM

Google has made Gemini 3.8 Live with Live Avatar generally available to enterprise customers, introducing a version of its conversational AI that can speak, process live visual input and appear as a synchronized digital character.
The release, announced September 24, 2026, is available through Gemini Enterprise and Google’s developer tools. Google says the model supports enterprise deployments through U.S. and European Union endpoints, with provisioned throughput, compliance features and data-governance controls.
The most visible change is the Live Avatar capability. Instead of returning only audio, Gemini 3.8 Live can generate a talking avatar with synchronized speech, facial expressions and lip movements. Google’s technical documentation describes the video output as 24 frames per second, allowing developers to build virtual concierges, tutors, customer-service representatives and other interactive agents without adding a separate avatar-rendering system.
The model is also designed to take in more than a spoken question. It can process audio alongside camera feeds, screen shares and text, enabling an agent to respond to what a user is showing it. Google’s examples include insurance claims intake, where a customer can describe damage while showing it on camera, and online vehicle shopping, where an assistant can highlight information on a screen while helping a shopper compare cars.
That combination of perception and action is more consequential for businesses than the avatar itself. Gemini 3.8 Live supports asynchronous tool calling, meaning it can contact backend systems or APIs while continuing the conversation. An agent might acknowledge that it is checking an order, policy or account record rather than falling silent until the transaction is complete. Google’s developer documentation also describes automatic cancellation and interruption handling for certain blocking calls, allowing the system to change course when a user speaks again.
The model supports automatic language detection and can switch languages during a conversation. Google says Live Avatar can maintain synchronized speech and video across 97 languages. That could make the system useful for international support operations, although the company’s announcement does not provide independent performance measurements for accuracy, latency or customer outcomes.
There are also important limits. Organizations can choose from a library of preset avatars, but custom avatar creation is restricted to enterprise customers that receive Google’s approval through an allowlisting process. Google says custom avatars can be created from a reference image while preserving a character’s appearance or brand styling. The company also says generated audio and video carry an imperceptible SynthID watermark intended to help identify AI-generated media.
The underlying Gemini 3.8 Live model is positioned as a low-latency, real-time system rather than a model optimized for extended deliberation. Google’s API documentation lists support for audio, image and video inputs, audio generation, function calling, search grounding and interleaved reasoning. It does not list image generation, structured outputs, caching or code execution as supported capabilities for this model.
That distinction matters for companies evaluating the technology. A polished avatar may make an interaction feel more human, but the practical value will depend on whether the agent can reliably identify what a user shows it, retrieve accurate business data and complete transactions without introducing new privacy, security or compliance risks. Google’s examples show the intended workflow; they do not independently establish how well the system performs at scale.
For now, the clearest change is strategic: Google is moving its real-time Gemini system from voice-only assistance toward embodied, multimodal agents that can maintain a visible presence while working with enterprise software in the background. Developers can access the model through Gemini Enterprise and the Gemini Live API, while custom-avatar deployments require further approval from Google.
Sources & Further Reading
- Google Blog — Introducing Gemini 3.8 Live with Live AvatarPrimary source
- Google Cloud Blog — Gemini 3.8 Live with Live Avatar is now generally availablePrimary source
- Google AI for Developers — Gemini 3.8 LivePrimary source
- Google Cloud Documentation — Developer's guide to Gemini 3.8 LivePrimary source
- Android Authority — Google's new Gemini Live Avatars want to make support bots feel more human
StoryBreak
Independent digital news and reporting, updated throughout the day.
This article was researched and drafted with AI assistance and reviewed as part of StoryBreak's editorial process before publication. Read our editorial standards.






