Gemini 3 Pro
Google Gemini
· gemini-3-pro
Multimodal Understanding
Production
Recommended, Best for Enterprise
Long-context multimodal model that reads video, audio, images and documents in a single request.
Best use cases: Audience insight, asset analysis, research synthesis
9.2
Quality
7
Speed
6.5
Cost Eff.
9
Brand Safety
7.5
Control
9
Adherence
Input cost
$2.00 / 1M tokens
Output cost
$12.00 / 1M tokens
Generation cost
—
Max context / duration
2M tokens
Aspect ratios
—
Resolutions
—
Input types
Text, images, video, audio, files
Output types
Text
API endpoint
generativelanguage.googleapis.com
Commercial use
Allowed
Data retention
Configurable via Vertex AI
Regions available
Global
API connected
Connected
Last updated
Release date
Documentation