---
格式版本: 2
标题: "Intelligent transcription with Gemini 3.5 Transcribe"
原文链接: "https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5-transcribe/"
发布日期: "2026-08-26"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:scrape:strict_html_metadata"
发布时间证据: "article:published_time: 2026-08-26"
发布时间校准原因: "规则确认唯一严格发布时间，来源 scrape:strict_html_metadata"
发布时间校准置信度: "high"
发布时间候选数量: 10
发布时间严格候选数量: 3
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-08-27T10:56:07+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-08-27T10:55:52+08:00"
入库时间: "2026-08-27T02:56:38.279Z"
来源平台: "固定入口"
搜索渠道: "fixed_url"
搜索词: "https://blog.google/"
匹配关键词:
  - "AI"
  - "performance"
  - "latency"
相关厂家:
  - "Google"
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "否"
AI打分: 43
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "正文主线是Google官方发布语音转写模型Gemini 3.5 Transcribe，新增流式与非流式API、公测状态、WER、延迟改进及语言支持等可核验事实；固定知识库未见该模型既有记录，但未命中不能证明首次出现。当前页面是一手完整来源，但内容面向语音应用与模型能力，没有AI机架、互连、供电、液冷、RAS或规模部署信息，命中“应用与模型效率”硬否决项，不具备超节点业务准入价值。"
AI质检模型: "gpt-5.6-sol"
AI质检时间: "2026-08-27T10:58:27+08:00"
AI主题相关性: 0
AI来源权威性: 15
AI新颖性: 14
AI技术细节: 1
AI商业部署信号: 3
AI完整性: 10
AI评分提示词版本: "v17-精简生产版"
AI评分提示词SHA256: "48fb9777f386026761b4873eaff30807694fb11e9b352d7c69bf2dfde750cc7d"
AI评分知识库版本: "knowledge_base_v1-20260819+runtime.6"
AI评分知识库SHA256: "da6710dec9b12b44eac7c5f65e2f6054e1d79564ad723a6342e799b618657548"
AI评分知识库检索词: "[\"Google\",\"https://blog.google/\",\"NPO\",\"RAS\",\"Intel\",\"API\",\"gemini-3.5-transcribe-live\",\"gemini-3.5-transcribe-live-preview\",\"APIs\",\"gemini-3.5-transcribe\",\"WER\",\"IDs\"]"
AI评分知识库命中: "[{\"id\":\"historical-jan-apr-02\",\"title\":\"二、Google Cloud Next '26：AI Hypercomputer 与第八代 TPU 发布\",\"sourceType\":\"curated_item\",\"time\":\"2026-01_to_2026-04\",\"matchedTerms\":[\"Google\",\"Intel\"],\"rank\":-11.217724105998673},{\"id\":\"july-correct-0033\",\"title\":\"Microsoft, Alphabet, Meta Pivot from Buy to Build in AI\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"Google\",\"RAS\",\"API\",\"APIs\",\"WER\"],\"rank\":-7.32769244979213},{\"id\":\"july-correct-0104\",\"title\":\"NVIDIA Vera Rubin 提升每瓦性能，为全球合作伙伴实现最低 Token 成本\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"Google\",\"RAS\",\"Intel\"],\"rank\":-5.985180378657868},{\"id\":\"historical-may-024\",\"title\":\"OpenAI、Microsoft等围绕MRC协议构建更大规模AI以太网训练网络\",\"sourceType\":\"curated_item\",\"time\":\"2026-05\",\"matchedTerms\":[\"Intel\"],\"rank\":-5.779665460055207},{\"id\":\"july-correct-0081\",\"title\":\"AAI 2026: 6th Gen AMD EPYC Server CPUs Power the Agentic Data Center\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"RAS\",\"Intel\",\"API\",\"WER\"],\"rank\":-5.705224088320232}]"
AI摘要: "Google推出语音转文字模型Gemini 3.5 Transcribe，开发者可在Gemini API和Gemini Enterprise Agent Platform中使用，支持实时流式及预录音频处理、自动清理口头语。"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-08-27T08:12:30.534Z"
采集批次: "2026年8月27日10点55分46秒"
采集批次ID: "20260827-105546-808"
去重键: "https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5-transcribe"
---

Today, we’re introducing Gemini 3.5 Transcribe, our most precise speech-to-text model yet, designed for intelligent voice interactions. Unlike conventional speech recognition models that struggle with background noise, complex jargon, and disfluency cleanup, Gemini 3.5 Transcribe converts raw audio directly into accurate, polished, formatted text.

Across our products like the Gemini app and on Android, we’ve seen consumers already benefiting from this transcription model with new voice capabilities like [Rambler on Android](https://blog.google/products-and-platforms/platforms/android/gemini-intelligence/) and in the Gemini app on macOS. Now, developers can build similar capabilities with Gemini 3.5 Transcribe in the [Gemini API in Google AI Studio](https://aistudio.google.com/live?model=gemini-3.5-transcribe-live) and [Gemini Enterprise Agent Platform](https://console.cloud.google.com/agent-platform/studio/multimodal-live?model=gemini-3.5-transcribe-live-preview).

We've built 3.5 Transcribe to plug seamlessly into your developer workflows, whether you’re building voice agents, real-time captioning tools, or post-call analytics pipelines. The model is available across two separate APIs:

- **Real-time streaming:** Delivers continuous, bidirectional streaming with sub-second latency for interactive voice apps via the [Live API](https://ai.google.dev/gemini-api/docs/live-api/live-transcribe) using **`gemini-3.5-transcribe-live`****.**
- **Pre-recorded audio processing:** Transcribes recorded audio, meetings, call logs, and more with speaker attribution and word-level timestamps via the [Interactions API](https://ai.google.dev/gemini-api/docs/transcribe) using **`gemini-3.5-transcribe`****.**

## Get more precise and intelligent transcription

Gemini 3.5 Transcribe is designed to capture your natural speaking style to better understand your intent and recognize custom vocabulary, so you can execute tasks with your voice.

- **Smart transcription:** Seamlessly handles self-corrections (like *"let’s meet Tuesday—no, Wednesday"*), removes filler words (“ums” and ‘“ahs"), auto-formats your text.
- **Function calling:** The model can delegate complex tasks (such as image generation and file analysis) to other Gemini models via function calls. Currently available in the [Gemini macOS app](https://blog.google/innovation-and-ai/products/gemini-app/speak-naturally-gemini-app-mac-os/).
- **More precise transcription:** As measured by Artificial Analysis, achieves an average Word Error Rate (WER) of 4.0% for streaming and 2.6% for non-streaming use-cases. It shows strong performance across noisy, real-world environments, accurately capturing alphanumeric entities like postal codes and order IDs.
- **Custom vocabulary:** Recognizes specialized jargon and unique spellings by seamlessly adapting transcriptions to your provided custom vocabulary.
- **Global language support:** Automatically detects and transcribes over 85 languages, seamlessly handling regional accents and diverse dialects.
- **Multi-speaker identification:** Accurately attributes speech in pre-recorded audio with timestamps for up to three speakers (support for 3+ speakers is experimental).

Gemini 3.5 Transcribe handles live language switches and seamless streaming transcription

Watch Gemini 3.5 Transcribe clean up speech disfluencies with smart transcription capabilities.

3.5 Transcribe delivers transcription with multi-speaker attribution and word-level timestamps.

Gemini 3.5 Transcribe’s performance represents a major advancement from our previous transcription model, Chirp 3, offering new capabilities, improved word error rates, and significantly better latency. As measured by Artificial Analysis, time to final transcription, for example, improves by 70%. On the FLEURS benchmark across a set of top languages and locales, the model delivers precise multilingual performance, improving over Chirp 3, and achieving a 5.50% WER in streaming mode and 5.04% WER in non-streaming use-cases.

## Experience smart transcription and advanced dictation

In addition to the Gemini API in the Google AI Studio and Gemini Enterprise Agent Platform, 3.5 Transcribe goes further than standard speech-to-text to make working across Google feel more natural and intuitive. By bringing context-aware understanding directly into everyday surfaces like Gboard, Antigravity, the Gemini app, and Chrome, it captures nuances, intent, and inline edits with ease.

- On **Gboard on Android,** through the new Rambler feature, 3.5 Transcribe transforms spoken thoughts into well-formatted text, filtering out filler words. You can also use your voice to make edits, correct misspellings, and change the writing style.
- On **Google Antigravity,** 3.5 Transcribe pairs screen context and chat history, with your permission, to ensure pinpoint transcription accuracy across file names, agent thoughts, and active documents.
- In **Google AI Studio**, you can access 3.5 Transcribe in Build mode to vibe code apps with your voice on the fly.
- In the **Gemini app on macOS**, 3.5 Transcribe not only transcribes your free natural speech into clean formatted text, but also enables voice commands that can pair seamlessly with screen context to power complex workflows. By calling on other Gemini models in the background to handle the heavy lifting, the model makes it effortless to summarize local files, repurpose text across apps, or generate images right at your cursor—using just your voice.
- Coming soon to **Chrome**, you’ll be able to talk to type in any web field — making it effortless to dictate replies, draft posts, or prompt Gemini in Chrome more naturally and easily with your voice.

Gemini 3.5 Transcribe lets you analyze files, generate images, and search in the Gemini app on macOS using just your voice.

See how Gemini 3.5 Transcribe uses Rambler on Android to automatically remove filler words and clean up speech.

Gemini 3.5 Transcribe leverages screen context on Google Antigravity to ensure accurate transcription accuracy.

## Read the early reviews

By leveraging the Gemini Live API, developer platforms such as [Agora](https://docs.agora.io/en/ai/models/asr/gemini), [Fishjam](https://docs.fishjam.io/tutorials/gemini-live-integration), [LangChain](https://docs.langchain.com/langsmith/trace-gemini-live), [LiveKit](https://docs.livekit.io/agents/models/stt/gemini/), [Pipecat](https://docs.pipecat.ai/api-reference/server/services/stt/google), [Vercel](https://vercel.com/docs/ai-gateway/modalities/speech-to-text), and [Vision Agents](https://visionagents.ai/integrations/stt/gemini) enable developers to build and deploy high-performance voice-driven interfaces with ease. These platforms manage complex real-time media streaming infrastructure behind the scenes, allowing developers to focus entirely on crafting the user experience.

Companies like Vivo, Intellitek Health, and Lingopal have also shared positive feedback on 3.5 Transcribe, highlighting its impressive latency, accuracy, and expansive language support.

### Start using 3.5 Transcribe today

- **For developers**: In public preview in the [Gemini API via Google AI Studio](https://aistudio.google.com/prompts/new_chat?model=gemini-3.5-transcribe) and [Google Antigravity](http://google.com/url?sa=j&url=https%3A%2F%2Fantigravity.google%2Fproduct%2Fantigravity-2&uct=1773758132&usg=eTZneyhE2yCXDuRpHPGTPF55jHQ.&opi=73833047&source=chat).
- **For enterprises**: In public preview via [Gemini Enterprise Agent Platform](https://console.cloud.google.com/agent-platform/studio/multimodal-live?model=gemini-3.5-transcribe-live-preview) and coming soon to [Gemini Enterprise for Customer Experience](https://cloud.google.com/gemini-enterprise-cx).
- **For everyone**: In Gemini app on macOS in English, [Rambler on Android](https://blog.google/products-and-platforms/platforms/android/gemini-intelligence/) in select [countries and languages](https://support.google.com/gboard/answer/17468539?hl=en#:~:text=Saturday%20to%20Sunday.%22-,Language%20support,-Rambler%20has%20been), and coming soon to Chrome.
