This is the English edition. 한국어판 and 日本語版 are also available.

Gemini 3.8 Live Explained: Voice AI Becomes a Real-Time Working Agent

2026-09-23 · AI · United States · Zoogom Editorial

#Gemini 3.8 Live#voice AI#AI agents#Live API#SynthID

Live audio waves, camera context and connected tools converging on a real-time AI agent

Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on September 15, 2026. This is more than an upgrade to conversational voice. The models can combine speech with visual context, invoke tools without ending the conversation and, in the Extended Thinking version, continue multi-step reasoning in the background. Voice AI is moving from answering a turn to managing a live task.

Three things to know

  1. Gemini 3.8 Live prioritizes low latency and natural back-and-forth; Extended Thinking adds background reasoning for complex workflows.
  2. The standard model accepts text, images, audio and video and supports function calls and search, but it does not natively provide every capability associated with text models.
  3. Production quality depends on interruption handling, permission boundaries, tool verification and data retention—not only how human the voice sounds.

How the two models differ

A comparison of Gemini 3.8 Live and Extended Thinking across latency, tools, conversation flow and use cases

How the two models differ: Dimension, Gemini 3.8 Live, Live Extended Thinking, Selection rule

The standard model identifier is gemini-3.8-live. Google’s developer page lists a 131,072-token input limit and a 65,536-token output limit. It can receive text, image, audio and video and return text or audio. Function calling, the Live API and search grounding are supported; caching, code execution, file search, image generation and structured outputs are listed as unsupported for this model.

Extended Thinking adds background reasoning to a live audio session. Instead of going silent during a longer plan or tool sequence, it can provide short conversational progress updates. The server exposes in-progress and idle states because one user request can produce several spoken updates before the entire interaction is complete.

Asynchronous tools change the voice-agent lifecycle

A spoken request moving through perception, reasoning, external tools, confirmation and verified completion

Traditional voice assistants often listen, answer once and stop. Gemini 3.8 Live makes nonblocking function calls the default, allowing an external operation to continue while the session remains responsive. A travel assistant, for example, could search flights and compare a calendar while acknowledging a revised date from the user.

That natural flow creates engineering obligations:

Developers migrating from the earlier Gemini 3.1 Flash Live preview also need to review configuration behavior. Google says the standard 3.8 Live model does not accept the prior thinking_level setup, and asynchronous execution is now the default unless a tool is explicitly declared as blocking.

Language switching is useful, but not a quality guarantee

Google says Gemini 3.8 Live can automatically detect and move among 97 supported languages during a conversation. That matters in U.S. customer support, travel, education and workplaces where speakers switch languages without announcing it.

Language count, however, does not prove equal performance. A deployment should test:

Availability also differs by product. A capability exposed in the Gemini API or AI Studio may not appear at the same time in Search Live, the Gemini app, Workspace or an enterprise preview. General availability of model endpoints and preview labels on individual platform features should be read separately.

Voice and vision expand the privacy boundary

Live audio and video passing through consent, local processing, cloud inference, tool permissions and retention controls

A real-time agent can receive far more ambient information than a typed chatbot. A microphone captures bystanders and background audio. A camera can expose documents, faces, location clues and workplace screens. Connected tools add calendars, messages, customer records and transaction history.

Before deployment, an organization should document:

Google says audio generated through its AI products contains an inaudible SynthID watermark. Provenance signals can help identify generated audio, but they do not guarantee factual accuracy, ownership or appropriate use. A user-facing disclosure is still useful when the voice could be mistaken for a person.

Vendor benchmarks are not deployment evidence

Google reports strong results for Extended Thinking across speech quality, agent tasks and audio reasoning evaluations. Those numbers describe selected harnesses and conditions. They cannot reproduce every noisy call center, classroom, car cabin, game voice channel or accessibility use case.

A real evaluation should measure:

The best tradeoff depends on consequence. A game character or navigation hint may favor immediate response. A financial transfer, medical intake or travel purchase should favor confirmation and verifiable completion even when it feels less conversational.

Where each model fits

Gemini 3.8 Live is a strong fit for real-time support, interactive tutoring, game characters, translation and interfaces in vehicles or smart glasses. Extended Thinking is more suitable when a conversation must coordinate a sequence of searches, comparisons and tool calls.

Not every audio feature requires a live agent. Batch transcription, document summarization, exact JSON generation or long-lived file retrieval may be simpler and more reliable in another model or a conventional pipeline. Live audio should be chosen because interruption, immediate feedback and continuous context matter—not because voice is fashionable.

Gemini 3.8 Live changes the evaluation question. The standard is no longer only whether an AI voice sounds natural. A useful agent must keep state, use tools safely, expose progress, accept interruption and prove that the requested action actually happened.

Sources and use notice

This article independently analyzes specifications and availability statements in Google’s official announcement and developer documentation. It does not reproduce official charts, product interfaces or benchmark graphics, and it does not treat vendor-reported scores as independent proof of every real-world use case.

Source: Google · Includes original screenshots or graphics