Two Live models for different requirements
Gemini 3.8 Live is intended for fluid conversation and low latency. Extended Thinking targets more complex tasks that require reasoning alongside the conversation. The release also describes tool use that need not constantly interrupt dialogue. The practical business question is whether an assistant should provide a short answer or check several sources before proposing a next step. Measure those requirements separately. The first audible response can be fast while a downstream task is still running. Your assistant needs to communicate that difference clearly rather than imply that speaking an acknowledgement means the work is complete.
An API release is not a universal Workspace rollout
The API release notes describe both models as generally available. Google names the Gemini API and AI Studio for developers and several product surfaces for end users. The announcement mentions Docs Live, Gmail Live and Keep Live among those experiences. Do not infer a single rollout covering every German account. An API project, a personal Gemini app and a managed Workspace account need separate checks. Record the region, account type, language, model and feature actually visible in the pilot. German availability of the Gemini API does not establish identical availability of every Workspace voice feature.
Example: prepare a service request by voice
An employee describes a problem with a device. The assistant asks for the device identifier, observed behaviour and steps already taken. It turns the conversation into a structured service-ticket draft and reads back the key details for confirmation. An explicitly selected source might be the internal manual. The first stage ends at the draft: no independent purchasing, binding commitment or changes to machinery controls. Missing identification or conflicting details trigger a hand-off to a person. This is a proposed workflow design, not a promise that Google provides that complete integration as a built-in feature.
Use Transcribe for transcripts and Live for dialogue
Gemini 3.5 Transcribe and Transcribe Live became generally available in the API on August 26, 2026. They handle speech-to-text tasks such as interview recordings, dictation or subtitles. The documentation describes automatic language recognition, speaker diarisation and timestamps. Custom vocabulary can help recognise specialist terms, but currently cannot be combined with diarisation or word-level timestamps. Decide in advance whether the output should be a readable transcript, a reliable quotation record or time-coded text. A cleaned-up version must not be presented as a verbatim account without identifying that transformation.
Make the boundary for tool actions explicit
In a voice conversation, accidental agreement can be harder to recognise than in a completed form. Before a consequential action, repeat the recipient, content and effect and require unambiguous confirmation. The application must handle interrupted conversations, duplicate requests and failed tool calls. It should announce success only after the destination system confirms the action. Read-only access is often sufficient for research. Also evaluate microphones, background noise and interruptions. These technical and operational requirements still apply when the model becomes more conversational; better speech quality does not establish reliable execution by itself.
Test a voice pilot with difficult cases
Use short, realistic test conversations containing specialist terminology, follow-up questions and a deliberate correction. Check time to first response, correctness of the final result and a dependable hand-off when information is uncertain. Repeat a case with background noise and another without the required system permission. For transcription, compare names, numbers and speaker attribution with the original recording. The AI System Check helps choose between an existing product surface, a custom API application and a simpler dictation solution. A useful solution reduces rework and remains manageable in daily operations; an impressive demonstration alone does not establish either.
Questions about Live, Transcribe and TTS
Is Extended Thinking automatically better for every task? No; additional processing needs to match complexity and acceptable response time. Can Transcribe conduct a conversation? It converts speech to text rather than providing a complete assistant dialogue. Does reading training material aloud require a Live model? Text-to-speech is usually the more relevant category. May a model submit a ticket automatically? That depends on the deliberately configured integration and granted permissions. A model release does not grant that authority. How current is this comparison? The sources and feature distinctions were checked on September 23, 2026.
Keep it verifiable
Primary sources
- Google: Gemini 3.8 Live and Extended ThinkingSource checked:
- Google: Real-time voice applications with Gemini AudioSource checked:
- Google: Gemini API release notesSource checked:
- Google: Audio transcription with GeminiSource checked:
- Google: AI Studio and Gemini API available regionsSource checked:

