Google DeepMind: Agentic Video Understanding for Gemini
Google DeepMind launches a new feature for Gemini that dynamically analyzes video content to improve accuracy and reduce costs.

The update
Google DeepMind has introduced agentic video understanding for its Gemini models. The new capability allows the models to dynamically scan video segments, improving accuracy while cutting token usage by up to 88% and costs by up to 66%. The feature is available today for video uploads and YouTube videos via the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform.
Why it matters
Unlike previous ‘static’ processing methods that ingest video at a fixed frames-per-second rate, this new approach pairs the model’s core reasoning with native video tools. This dynamic scanning across visual frames, audio, and transcripts unlocks new capabilities for video processing, such as sub-second moment retrieval, more accurate anomaly detection, and precise counting.
What to watch
Developers can start using this feature by setting their API configuration to ‘agentic’ in Google AI Studio or the Gemini Enterprise Agent Platform. The improvement in accuracy and cost-efficiency could shift how developers approach video analysis tasks.
Sources
- Google DeepMind Blog — Primary source for the announcement, feature details, and performance metrics.
- Google Blog — Secondary confirmation of the announcement and feature availability.
