---
格式版本: 2
标题: "Introducing agentic video understanding with Gemini"
原文链接: "https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-agentic-video-in-gemini/"
发布日期: "2026-09-01"
发布时间校准状态: "found"
发布时间需复核: "否"
发布时间来源: "rule:scrape:strict_html_metadata"
发布时间证据: "article:published_time: 2026-09-01"
发布时间校准原因: "规则确认唯一严格发布时间，来源 scrape:strict_html_metadata"
发布时间校准置信度: "high"
发布时间候选数量: 10
发布时间严格候选数量: 3
发布时间原页读取状态: "source template page reused from URL open"
发布时间未找到原因: ""
发布时间校准时间: "2026-09-02T14:47:30+08:00"
发布时间仲裁状态: "skipped"
发布时间仲裁尝试次数: 0
发布时间仲裁耗时毫秒: 0
发现时间: "2026-09-02T14:46:59+08:00"
入库时间: "2026-09-02T06:48:05.158Z"
来源平台: "固定入口"
搜索渠道: "fixed_url"
搜索词: "https://blog.google/"
匹配关键词:
  - "AI"
  - "performance"
相关厂家:
  - "Google"
相关专家:
  []
内容类型: "网页"
抓取工具: "Free Fetch + Defuddle"
清洗工具: "Defuddle Markdown + Defuddle/Readability 正文提取"
原始附件:
  []
AI优质: "否"
AI打分: 51
AI分档: "非优质"
AI质检状态: "不通过"
AI打分理由: "正文主线是Gemini视频理解功能发布及模型推理成本优化，不是超节点、AI机柜或机架级基础设施。Google官方原文权威且明确新增动态帧率检索机制、最高88% Token降幅、66%成本降幅、7%准确率提升，以及API当日可用和后续产品推广计划；固定知识库未见同一发布，但未命中不能证明首次出现。当前页面可作为该模型功能的一手发布源，然而技术与指标均面向视频理解应用，未提供可复用的机架级通信、网络、内存、RAS、供电或散热机制，命中“应用与模型效率”硬否决，最高54分。"
AI质检模型: "gpt-5.6-sol"
AI质检时间: "2026-09-02T14:49:07+08:00"
AI主题相关性: 1
AI来源权威性: 15
AI新颖性: 16
AI技术细节: 7
AI商业部署信号: 4
AI完整性: 8
AI评分提示词版本: "v17-精简生产版"
AI评分提示词SHA256: "48fb9777f386026761b4873eaff30807694fb11e9b352d7c69bf2dfde750cc7d"
AI评分知识库版本: "knowledge_base_v1-20260819+runtime.81"
AI评分知识库SHA256: "3b93d12e47749b3512f545f51c44c011bdc0931677c2cfe61e4df59a2b1a5a48"
AI评分知识库检索词: "[\"Google\",\"https://blog.google/\",\"NPO\",\"NPU\",\"API\",\"FPS\",\"gemini-3.7-flash\",\"youtu.be/7Z5Vy9JBANs\",\"support.google.com/youtube/answer/14110396\",\"GENIE.Platform\"]"
AI评分知识库命中: "[{\"id\":\"runtime-3dcabc270e938e0c3b6d5245\",\"title\":\"Nvidia wants to bypass the CPU with an open-source AI storage overhaul\",\"sourceType\":\"ai_excellent_article\",\"time\":\"2026-08-06\",\"matchedTerms\":[\"Google\",\"API\"],\"rank\":-7.800961725461139},{\"id\":\"july-correct-0088\",\"title\":\"StrataCL: Fabric-Native Communication Library for Production Supernodes - 智源社区论文\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"NPU\"],\"rank\":-6.523814156975848},{\"id\":\"runtime-bf4039a0342d37545e9459a2\",\"title\":\"Most Neoclouds Suck At Security\",\"sourceType\":\"ai_excellent_article\",\"time\":\"2026-08-30\",\"matchedTerms\":[\"Google\",\"https://blog.google/\",\"API\"],\"rank\":-6.148987331686944},{\"id\":\"july-correct-0111\",\"title\":\"StrataCL: Fabric-Native Communication Library for Production Supernodes\",\"sourceType\":\"labeled_article\",\"time\":\"2026-07\",\"matchedTerms\":[\"NPU\"],\"rank\":-6.037951843305323},{\"id\":\"historical-jun-025\",\"title\":\"英伟达、谷歌与国产超节点的三种网络选择\",\"sourceType\":\"curated_item\",\"time\":\"2026-06\",\"matchedTerms\":[\"NPU\"],\"rank\":-6.023003238134376}]"
AI摘要: "Google 在 Gemini 3.7 Flash、3.6 Flash 和 3.5 Flash-Lite 中推出 agentic video understanding，可通过 Gemini API 使用，动态检索视频中的视觉帧。"
AI摘要模型: "ali-deepseek-v4-flash"
AI摘要时间: "2026-09-02T08:11:01.811Z"
采集批次: "2026年9月2日14点42分57秒"
采集批次ID: "20260902-144257-884"
去重键: "https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-agentic-video-in-gemini"
---

Today, we’re launching [agentic video understanding](https://ai.google.dev/gemini-api/docs/video-understanding#agentic-video-understanding) across our latest models: Gemini 3.7 Flash, 3.6 Flash and 3.5 Flash-Lite. This new capability improves accuracy while dramatically reducing token usage and costs for video analysis. Similar to [agentic vision](https://blog.google/innovation-and-ai/technology/developers-tools/agentic-vision-gemini-3-flash/), which combines code execution with Gemini models’ native image understanding, agentic video understanding uses Gemini’s native video tools to improve performance and unlock new capabilities for video processing like sub-second moment retrieval, more accurate anomaly detection, precise counting and more.

The feature is available today for video uploads and YouTube videos via the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform.

## Benchmarks

Unlike current ‘static’ processing, where the model ingests the video at a fixed frames-per-second rate (default 1 FPS, adjustable via API), agentic video understanding pairs the model’s core reasoning with native video tools to dynamically search, scan, and inspect target video segments across visual frames, audio, and transcripts. Across standard video analysis benchmarks, Gemini models with agentic video understanding **reduce analysis costs by up to 66% and token consumption by up to 88%, while improving accuracy by up to 7%.**

These efficiency gains are especially pronounced on long-form video (from 10-minute how-to guides to 90-minute lectures and multi-hour recordings), where static processing forces developers to choose between high token costs or techniques that drop critical details.

Activating agentic video understanding drops token consumption by up to 88% and boosts accuracy by up to 7% with Gemini 3.7 Flash.

While these gains span all three supported models, Gemini 3.7 Flash with agentic understanding offers the best possible quality overall and the best combination of quality and cost efficiency, putting it at the accuracy-to-cost pareto frontier among tested models for video understanding.

Using agentic video understanding places Gemini 3.7 Flash at the accuracy-to-cost pareto frontier for video analysis.

## How it works

Instead of static processing where the model ingests media streams at a fixed frame rate, agentic video understanding enables Gemini to take an active, goal-directed role in determining *what* to watch, at *what* speed, and through *which* modality (frames, audio, or transcript), fetching only the moments and signals needed. While developers could previously do this manually, with agentic video understanding, Gemini can accomplish it through an agentic loop, invoking an internal tool to load the relevant part of the video file, significantly reducing development overheads.

## Capabilities and use cases

Agentic video understanding transforms how developers can process long-form video content across a variety of demanding applications.

- **Sub-second moment retrieval**: Pinpoint split-second state changes and tight cut boundaries that are easily missed at 1 FPS, making precise automated video editing possible.
- **Long-form needle-in-a-haystack search**: Answer complex queries across multi-hour videos without consuming millions of tokens.
- **Anomaly detection**: Resample interesting time windows at higher FPS to inspect rapid motion and subtle visual artifacts.
- **Counting action & object**: Accurately track repeated physical movements and distinct objects over time.

***Token-efficient long-form video analysis***

*See how Gemini 3.7 Flash performs with and without agentic video understanding on LongVideoBench, a long-form video understanding benchmark. Notice the large token reductions and accuracy improvements.*

***Accurate fast action analysis with dynamic FPS***

*With agentic video understanding, 3.7 Flash is able to accurately count a fast-paced movement by scanning and rewatching the video at different frames per second, as needed.*

**Token-efficient needle-in-a-haystack search**

*Using agentic video understanding, Gemini 3.7 is able to accurately answer complex questions based on the content of the video while consuming a significantly lower number of tokens compared to static analysis.*

## Real-world results

Many of our early access partners saw strong performance while testing with agentic video understanding. Here’s what they have to say:

## Getting started

Agentic video understanding is available via the Gemini API in [Google AI Studio](https://ai.google.dev/gemini-api/docs/video-understanding#agentic-video-understanding) and [Gemini Enterprise Agent Platform](https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/capabilities/video-understanding), launching across Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite. It uses standard Gemini API token pricing with no additional feature fee.

To enable it, simply set processing to "agentic" in the API configuration. Read our [developer guide](http://ai.dev/learn/agentic-video-understanding-with-gemini) to get more insights into the feature and how to get started.

```py
from google import genai

client = genai.Client()

interaction = client.interactions.create(
    model="gemini-3.7-flash",
    input=[
        {
            "type": "video",
            "uri": "https://youtu.be/7Z5Vy9JBANs",
            "processing": "agentic"
        },
        {
            "type": "text",
            "text": "What are the 3 most important announcements in this keynote?",
        },
    ],
)

print(interaction.output_text)
```

We are also bringing the efficiency and quality improvements of agentic video understanding to billions of users across Google products. The feature will roll out to all users in the Gemini app across Flash and Flash-Lite models soon. And in the coming months, agentic video understanding will also power YouTube's ‘ [Ask YouTube’](https://support.google.com/youtube/answer/14110396?hl=en&co=GENIE.Platform%3DAndroid) feature on the video watch page, leveraging Gemini to deliver higher-quality answers grounded in the visuals.

***Acknowledgement for their contribution to this work:*** *Sergi Caelles, Filip Pavetić, Ahmet Iscen, Suhas Yogin, and the Agentic Vision team.*
