Google AI: September 2026 review

Gemini 3.8 Live: voice dialogue, Extended Thinking and transcription

A voice assistant can receive questions, retrieve information and prepare next steps. It needs more than a pleasant voice. September’s Gemini Live releases make the choice between response time, reasoning effort and permitted actions particularly important.

WERKVERSTAND / CONNECTING INTELLIGENCE

The essential answer

Google released Gemini 3.8 Live and 3.8 Live Extended Thinking on September 15, 2026. Both support real-time voice dialogue; Extended Thinking adds more extensive background reasoning. Gemini Transcribe instead converts speech to text, while TTS voices written text. Choose the task first, then the model and product surface.

Two Live models for different requirements

Gemini 3.8 Live is intended for fluid conversation and low latency. Extended Thinking targets more complex tasks that require reasoning alongside the conversation. The release also describes tool use that need not constantly interrupt dialogue. The practical business question is whether an assistant should provide a short answer or check several sources before proposing a next step. Measure those requirements separately. The first audible response can be fast while a downstream task is still running. Your assistant needs to communicate that difference clearly rather than imply that speaking an acknowledgement means the work is complete.

An API release is not a universal Workspace rollout

The API release notes describe both models as generally available. Google names the Gemini API and AI Studio for developers and several product surfaces for end users. The announcement mentions Docs Live, Gmail Live and Keep Live among those experiences. Do not infer a single rollout covering every German account. An API project, a personal Gemini app and a managed Workspace account need separate checks. Record the region, account type, language, model and feature actually visible in the pilot. German availability of the Gemini API does not establish identical availability of every Workspace voice feature.

Example: prepare a service request by voice

An employee describes a problem with a device. The assistant asks for the device identifier, observed behaviour and steps already taken. It turns the conversation into a structured service-ticket draft and reads back the key details for confirmation. An explicitly selected source might be the internal manual. The first stage ends at the draft: no independent purchasing, binding commitment or changes to machinery controls. Missing identification or conflicting details trigger a hand-off to a person. This is a proposed workflow design, not a promise that Google provides that complete integration as a built-in feature.

Use Transcribe for transcripts and Live for dialogue

Gemini 3.5 Transcribe and Transcribe Live became generally available in the API on August 26, 2026. They handle speech-to-text tasks such as interview recordings, dictation or subtitles. The documentation describes automatic language recognition, speaker diarisation and timestamps. Custom vocabulary can help recognise specialist terms, but currently cannot be combined with diarisation or word-level timestamps. Decide in advance whether the output should be a readable transcript, a reliable quotation record or time-coded text. A cleaned-up version must not be presented as a verbatim account without identifying that transformation.

Make the boundary for tool actions explicit

In a voice conversation, accidental agreement can be harder to recognise than in a completed form. Before a consequential action, repeat the recipient, content and effect and require unambiguous confirmation. The application must handle interrupted conversations, duplicate requests and failed tool calls. It should announce success only after the destination system confirms the action. Read-only access is often sufficient for research. Also evaluate microphones, background noise and interruptions. These technical and operational requirements still apply when the model becomes more conversational; better speech quality does not establish reliable execution by itself.

Test a voice pilot with difficult cases

Use short, realistic test conversations containing specialist terminology, follow-up questions and a deliberate correction. Check time to first response, correctness of the final result and a dependable hand-off when information is uncertain. Repeat a case with background noise and another without the required system permission. For transcription, compare names, numbers and speaker attribution with the original recording. The AI System Check helps choose between an existing product surface, a custom API application and a simpler dictation solution. A useful solution reduces rework and remains manageable in daily operations; an impressive demonstration alone does not establish either.

Questions about Live, Transcribe and TTS

Is Extended Thinking automatically better for every task? No; additional processing needs to match complexity and acceptable response time. Can Transcribe conduct a conversation? It converts speech to text rather than providing a complete assistant dialogue. Does reading training material aloud require a Live model? Text-to-speech is usually the more relevant category. May a model submit a ticket automatically? That depends on the deliberately configured integration and granted permissions. A model release does not grant that authority. How current is this comparison? The sources and feature distinctions were checked on September 23, 2026.

Keep it verifiable

Primary sources

The next sensible step

Assess your AI workflow

Eight steps from a general interest in AI to a clearer decision for your business.

Start AI System Check
FreeProvider-neutralNo credentials