GPT-5.6 (Sol / Terra / Luna) is now evaluated on TrustVector โ€” with day-1 independent verification, incl. METR's benchmark-cheating findings.

Read the evaluation
Evaluation record ยท gemini-3-pro

Gemini 3 Pro

vgemini-3-pro-preview

Google

Modelretiredlong-context1m-tokensdeep-think
90
Exceptional
About This Model

SHUT DOWN: Google retired Gemini 3 Pro Preview on 2026-03-09 (the gemini-pro-latest alias moved to 3.1 Pro on 2026-03-06); it is no longer served via the Gemini API or AI Studio. Historically a former Google flagship with 1M token context, 1501 LMArena Elo (first model >1500), Deep Think mode, and native multimodal. Migrate to Gemini 3.1 Pro (same pricing, ARC-AGI-2 77.1% vs 31.1%).

Last Evaluated: July 9, 2026
Official Website

Trust Vector Analysis

Dimension Breakdown

๐Ÿš€Performance & Reliability
+

First model to exceed 1500 LMArena Elo. 1M context enables unprecedented document processing. 6x improvement on ARC-AGI-2 over 2.5 Pro.

task accuracy code

Industry-standard coding benchmarks

Evidence
SWE-bench Verified โ€” 76.2% on SWE-bench Verified
WebDev Arena โ€” 1487 Elo on WebDev Arena
highVerified: 2026-07-09
task accuracy reasoning

PhD-level and world-leading reasoning benchmarks

Evidence
GPQA Diamond โ€” 93.8% with Deep Think (91.9% standard)
ARC-AGI-2 โ€” 45.1% Deep Think / 31.1% standard (~6x improvement over 2.5 Pro)
AIME 2025 โ€” 95% (no tools) / 100% (with tools)
Humanity's Last Exam โ€” 41% Deep Think / 37.5% standard (world-leading)
highVerified: 2026-07-09
task accuracy general

Crowdsourced and comprehensive testing

Evidence
LMArena Elo โ€” 1501 Elo (first model to exceed 1500)
MMLU โ€” 90% on cross-discipline knowledge
MMMU-Pro โ€” 81% multimodal understanding
highVerified: 2026-07-09
output consistency

Consistency and efficiency testing

Evidence
Google AI Documentation โ€” 7x better token efficiency than 2.5 Pro
mediumVerified: 2026-07-09
latency p50

Median latency measurements

Evidence
Community benchmarking โ€” Typical response time ~1.5s
mediumVerified: 2026-07-09
latency p95

95th percentile measurements

Evidence
Community benchmarking โ€” p95 latency ~4.0s
mediumVerified: 2026-07-09
context window

Official specification

Evidence
Google AI Documentation โ€” 1M token context window
highVerified: 2026-07-09
uptime

Historical uptime data

Evidence
Google Cloud Status โ€” 99.9% uptime (last 90 days)
highVerified: 2026-07-09
๐Ÿ›ก๏ธSecurity
+

Strong security with Google Cloud infrastructure. Configurable safety filters provide flexibility.

prompt injection resistance

OWASP LLM security testing

Evidence
Google AI Safety โ€” Enhanced prompt injection defenses
mediumVerified: 2026-07-09
jailbreak resistance

Adversarial prompt testing

Evidence
Google Safety Testing โ€” Improved jailbreak resistance
mediumVerified: 2026-07-09
data leakage prevention

Privacy policy review

Evidence
Google Privacy Policy โ€” API data not used for training
mediumVerified: 2026-07-09
output safety

Safety testing

Evidence
Google Safety Filters โ€” Configurable multi-category safety filters
highVerified: 2026-07-09
api security

API security review

Evidence
Google Cloud Security โ€” Google Cloud security standards
highVerified: 2026-07-09
๐Ÿ”’Privacy & Compliance
+

Good privacy with Google Cloud. HIPAA compliance available through Google Cloud Healthcare API.

data residency

Cloud infrastructure review

Evidence
Google Cloud Regions โ€” Multiple region options
highVerified: 2026-07-09
training data optout

Terms review

Evidence
Gemini API Terms โ€” API data not used for training
highVerified: 2026-07-09
data retention

Data retention policy review

Evidence
Google Cloud Terms โ€” Enterprise zero retention available
mediumVerified: 2026-07-09
pii handling

Data protection review

Evidence
Google AI Safety โ€” Customer responsible for PII
mediumVerified: 2026-07-09
compliance certifications

Certification verification

Evidence
Google Cloud Compliance โ€” SOC 2, ISO 27001, GDPR, HIPAA (via Google Cloud)
highVerified: 2026-07-09
zero data retention

Enterprise feature review

Evidence
Enterprise Options โ€” Available for enterprise
mediumVerified: 2026-07-09
๐Ÿ‘๏ธTrust & Transparency
+

Strong transparency with Deep Think mode. Comprehensive documentation and configurable guardrails.

explainability

Reasoning transparency evaluation

Evidence
Deep Think Mode โ€” Deep Think exposes detailed reasoning process
highVerified: 2026-07-09
hallucination rate

Factual QA testing

Evidence
Google AI Testing โ€” Improved factual accuracy over 2.5 Pro
mediumVerified: 2026-07-09
bias fairness

Bias benchmark evaluation

Evidence
Google AI Principles โ€” Regular bias testing and mitigation
mediumVerified: 2026-07-09
uncertainty quantification

Qualitative assessment

Evidence
Model Behavior โ€” Expresses uncertainty appropriately
mediumVerified: 2026-07-09
model card quality

Documentation review

Evidence
Gemini 3 Documentation โ€” Comprehensive documentation
highVerified: 2026-07-09
training data transparency

Public disclosure review

Evidence
Google AI Blog โ€” General training description
mediumVerified: 2026-07-09
guardrails

Safety mechanism review

Evidence
Safety Settings โ€” Configurable multi-category safety
highVerified: 2026-07-09
โš™๏ธOperational Excellence
+

Model retired 2026-03-09 after roughly four months in preview; no longer served. Operational history reflects its time in service on Google Cloud infrastructure.

api design quality

API design review

Evidence
Gemini API โ€” RESTful API with streaming, function calling, multimodal
highVerified: 2026-07-09
sdk quality

SDK quality assessment

Evidence
Google AI SDKs โ€” SDKs for Python, Node.js, Go, Swift, Kotlin, Dart
highVerified: 2026-07-09
versioning policy

Versioning policy review

Evidence
Google Cloud Versioning โ€” Clear versioning with migration guides
Gemini API Deprecations โ€” gemini-3-pro-preview shut down 2026-03-09; gemini-pro-latest alias moved to 3.1 Pro on 2026-03-06; designated replacement is gemini-3.1-pro-preview (released 2026-02-19; ARC-AGI-2 77.1% vs 31.1%)
highVerified: 2026-07-09
monitoring observability

Observability tools review

Evidence
Google Cloud Console โ€” Comprehensive Cloud Console monitoring
highVerified: 2026-07-09
support quality

Support assessment

Evidence
Google Cloud Support โ€” Enterprise support with SLAs
highVerified: 2026-07-09
ecosystem maturity

Ecosystem analysis

Evidence
Google AI Ecosystem โ€” Day-one launch across Gemini app, AI Studio, Vertex AI
highVerified: 2026-07-09
license terms

License review

Evidence
Google Cloud Terms โ€” Standard commercial terms
highVerified: 2026-07-09
Strengths
  • +First model to exceed 1500 LMArena Elo (1501)
  • +1M token context window (5x GPT-5.2, 5x Claude Opus 4.5)
  • +93.8% GPQA Diamond with Deep Think
  • +45.1% ARC-AGI-2 Deep Think (6x improvement over 2.5 Pro)
  • +Native multimodal (text, image, video, audio)
  • +Competitive pricing ($2/$12 per 1M tokens)
  • +7x better token efficiency than 2.5 Pro
Limitations
  • !Preview status (not yet GA)
  • !Slightly behind on SWE-bench (76.2% vs Claude's 80.9%)
  • !Deep Think increases latency significantly
  • !Data retention policies less clear than Anthropic
  • !Newer model with less community testing
  • !SHUT DOWN 2026-03-09: gemini-3-pro-preview is no longer served; migrate to Gemini 3.1 Pro
Metadata
pricing
input: $2.00 per 1M tokens (<200K), $4.00 per 1M tokens (>200K)
output: $12.00 per 1M tokens (<200K), $18.00 per 1M tokens (>200K)
consumer: $19.99/month (Google AI Pro), $124.99/month (Gemini 3 Ultra)
notes: Historical pricing; model shut down 2026-03-09. Gemini 3.1 Pro carries identical API pricing.
last verified: 2026-07-09
context window: 1000000
max output: 64000
languages
0: English
1: 100+ languages
modalities
0: text
1: vision
2: audio
3: video
api endpoint: https://generativelanguage.googleapis.com/v1beta/models
open source: false
architecture: Multimodal transformer with Deep Think reasoning
parameters: Not disclosed
knowledge cutoff: January 2025

Use Case Ratings

code generation

76.2% SWE-bench, 1487 WebDev Arena. 1M context enables full codebase analysis.

customer support

Native multimodal enables image/video support. Strong conversational abilities.

content creation

Excellent for content with multimodal capabilities and long context.

data analysis

1M context enables analysis of massive datasets. Strong analytical reasoning.

research assistant

Best for research: 1M context processes entire books/papers. Deep Think for complex analysis.

legal compliance

1M context for full contract analysis. HIPAA via Google Cloud Healthcare.

healthcare

HIPAA via Google Cloud. Good for processing medical records with long context.

financial analysis

Strong quantitative reasoning. 1M context for large financial document sets.

education

95-100% AIME. Excellent for teaching with multimodal explanations.

creative writing

Good creative capabilities with strong narrative flow.