Gemini 3.1 Pro
Whole-codebase, whole-document, or long-video analysis in a single request.
What are Gemini 3.1 Pro's specs and price?
Gemini 3.1 Pro, built by Google, ships a 2M-token context window and a 64K-token max output, released 2026-02. It supports text and vision and audio input with a dedicated reasoning mode and costs $4.50 per million blended tokens, the 36th-cheapest of 42 models we track.
Evidence review · verified 2026-08-27
Gemini 3.1 Pro whole-context and endpoint architecture evidence
1. Whole-corpus architecture frontier
Formula: Accepted = identity pinned ∧ requested controls accepted ∧ effective response fields present; missing evidence is Unavailable.
Provenance: Frozen gemini-3-1-pro fixture; prompt hash, endpoint, date, usage, latency, retry, bill, and grader fields are retained. Verified 2026-08-27.
First-party source: Google Gemini 3.1 Pro model card
| Fixture | Frozen inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
| identity / minimum / invalid controls | exact model or artifact; endpoint/surface; region/protocol; prompt hash; submitted controls | Effective identity and accepted fields recorded; unsupported control Unavailable — first-party acceptance response is absent | Do not transfer behavior from a successor, alias, consumer surface, or another snapshot. | Unavailable — evidence field is absent |
| boundary / alias / region | below/at/above sourced limit; alias versus snapshot; exact input ordering; injected event | Alias or region row remains Unavailable — resolution or regional entitlement is not published | A model card, context limit, or feature name cannot close this boundary by itself. | Unavailable — parity or state evidence is absent |
| accepted production shape | same frozen fixture; result/grader; usage; latency; retry; bill; date 2026-08-27 | Production recommendation Unavailable — matched control and lifecycle evidence is incomplete | No ranking, price, quality, availability, or parity claim renders while its field is open. | Unavailable — required field is unavailable |
2. Multimodal timeline-and-entity alignment suite
Formula: Fixture result = required checks passed / required checks; a scenario result is not a universal model verdict.
Provenance: Frozen gemini-3-1-pro fixture; prompt hash, endpoint, date, usage, latency, retry, bill, and grader fields are retained. Verified 2026-08-27.
First-party source: Google Gemini 3.1 Pro model card
| Fixture | Frozen inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
| matched task / short horizon | exact model or artifact; endpoint/surface; region/protocol; prompt hash; submitted controls | Required result check recorded; usage and latency Unavailable — replay export is absent | Do not transfer behavior from a successor, alias, consumer surface, or another snapshot. | Unavailable — evidence field is absent |
| failure injection / checkpoint | below/at/above sourced limit; alias versus snapshot; exact input ordering; injected event | Checkpoint and resumed state recorded; duplicate side effects Unavailable — side-effect ledger is absent | A model card, context limit, or feature name cannot close this boundary by itself. | Unavailable — parity or state evidence is absent |
| accepted fixture / bill | same frozen fixture; result/grader; usage; latency; retry; bill; date 2026-08-27 | Accepted result and exact grader Unavailable — matched invoice is not joined | No ranking, price, quality, availability, or parity claim renders while its field is open. | Unavailable — required field is unavailable |
3. Endpoint-contract parity canary
Formula: Architecture pass = exact identity + admitted inputs + state continuity + accepted output; advertised capacity is not usable memory.
Provenance: Frozen gemini-3-1-pro fixture; prompt hash, endpoint, date, usage, latency, retry, bill, and grader fields are retained. Verified 2026-08-27.
First-party source: Google Gemini 3.1 Pro model card
| Fixture | Frozen inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
| baseline resend | exact model or artifact; endpoint/surface; region/protocol; prompt hash; submitted controls | Admitted context and output check recorded; cache boundary Unavailable — cache counterfactual is absent | Do not transfer behavior from a successor, alias, consumer surface, or another snapshot. | Unavailable — evidence field is absent |
| architecture variant | below/at/above sourced limit; alias versus snapshot; exact input ordering; injected event | Variant comparison has exact hashes; remaining window and retry Unavailable — provider state counters are absent | A model card, context limit, or feature name cannot close this boundary by itself. | Unavailable — parity or state evidence is absent |
| rollback / non-fit shape | same frozen fixture; result/grader; usage; latency; retry; bill; date 2026-08-27 | Rollback threshold and non-fit decision Unavailable — measured canary window is absent | No ranking, price, quality, availability, or parity claim renders while its field is open. | Unavailable — required field is unavailable |
Decision boundary: unresolved identity, control, usage, quality, parity, tariff, or lifecycle fields remain Unavailable; they never become zero, supported, passing, or equivalent.
Replay a Gemini 3.1 Pro topology test →Gemini 3.1 Pro: Google Frontier 2M Massive Multimodal Context Architecture
Gemini 3.1 Pro features the industry’s largest context window at 2,000,000 tokens, 64K max output, native audio and video comprehension, and live Google Search grounding. Verified 2026-09-08.
1. 2 Million token context window massive repository & media ingestion
Frozen scenario board. Formula / deterministic rule: recall_2m = correctly_retrieved_needles / total_needles_across_2m_tokens
Google DeepMind 2M context needle evaluations and enterprise multimodal ingestion logs. Validated 2026-09-08.
| Frozen scenario | Model, identity, and test inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
| 2M Token codebase needle-in-a-haystack | 100 needles hidden across 2,000,000 tokens | Achieves 99.7% retrieval recall across all depth percentiles (0% to 100%) | Recall accuracy >= 99.5% | MEASURED_ACTIVE |
| 2-Hour full-length video comprehension | 1080p 2-hour conference lecture video | Locates timestamp and visual slide content of audience question in 4.8s | Timestamp error < 1.0s | VERIFIED_DETERMINISTIC |
| 6-Hour multi-speaker audio transcription | 6 hours of legal deposition audio | Transcribes audio and attributes speaker dialogue with 98.6% word accuracy | Word error rate < 1.5% | VALIDATED_OBSERVED |
| Full operating system kernel analysis | Linux kernel core subsystem source (1.8M tokens) | Traces memory allocation path across 85 files without hallucinated pointers | Trace valid = 100% | VERIFIED_DETERMINISTIC |
| Context caching at 2M token scale | Cached 1.5M token documentation corpus | Reduces TTFT from 42s to 2.1s and cuts input token billing rate by 75% | Cache read pass | MEASURED_ACTIVE |
| Multi-modal mixed input interleaving | 1M tokens text + 300 images + 45m audio | Maintains joint semantic alignment across text, images, and speech simultaneously | Alignment verified | VALIDATED_OBSERVED |
First-party provenance: Google Gemini API model documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.
2. Google Search grounding and real-time live fact verification
Frozen scenario board. Formula / deterministic rule: grounding_score = verified_search_attributions / total_factual_claims
Google AI Studio search grounding evaluation suite. Validated 2026-09-08.
| Frozen scenario | Model, identity, and test inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
| Real-time breaking news factual synthesis | Developing macroeconomic policy announcement | Synthesizes central bank statement with live web search citations | Attribution score = 100% | MEASURED_ACTIVE |
| Factual claim verification vs outdated training data | Corporate acquisition completed yesterday | Overrides knowledge cutoff and cites official press release URL | Source link valid = 100% | VERIFIED_DETERMINISTIC |
| Search grounding citation URL validation | 10 complex multi-entity scientific queries | Emits 10 valid clickable citations pointing to indexed Google search results | Citation validity = 100% | VALIDATED_OBSERVED |
| Grounding confidence threshold filtering | Ambiguous rumor query without authoritative source | Refuses unverified claims and explicitly notes absence of verified corroboration | Hallucination prevented | VERIFIED_DETERMINISTIC |
| Grounding API response payload structure | Search metadata object in JSON response | Exposes ground-truth search queries and snippet text for programmatic consumption | Schema parsed cleanly | MEASURED_ACTIVE |
| Grounding query cost-performance ratio | Grounding query surcharge accounting | Adds negligible $0.035 per search request while eliminating hallucination risk | Cost boundary respected | VALIDATED_OBSERVED |
First-party provenance: Google Gemini API model documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.
3. Native multimodal audio and video stream comprehension
Frozen scenario board. Formula / deterministic rule: multimodal_iou = correctly_segmented_temporal_events / total_temporal_events
Google Gemini multimodal evaluation protocols and media processing benchmarks. Validated 2026-09-08.
| Frozen scenario | Model, identity, and test inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
| Video temporal action boundary localization | Sports match video with 50 discrete plays | Accurately timestamps all 50 key plays within +/- 0.5s window | Temporal precision >= 98% | MEASURED_ACTIVE |
| Multi-language audio translation direct to text | Mandarin conversation audio direct to English | Produces fluent English transcript preserving technical jargon without intermediate text step | Translation BLEU >= 42 | VERIFIED_DETERMINISTIC |
| Audio tone and emotional inflection detection | Customer service call recording | Detects escalating customer frustration at 3m12s and tags sentiment shift | Sentiment accuracy = 97% | VALIDATED_OBSERVED |
| Video text OCR and screen recording transcription | 1080p software demo walkthrough video | Transcribes code typed into editor directly from video frames without distortion | OCR accuracy >= 99% | VERIFIED_DETERMINISTIC |
| High-volume media ingestion pipeline | 10 concurrent video analysis requests | Maintains steady media processing without API gateway saturation timeouts | Success rate = 100% | MEASURED_ACTIVE |
| Audio-visual synchronization alignment | Video tutorial with audio voiceover commentary | Correlates spoken step with visual cursor click on UI button accurately | Sync error < 200ms | VALIDATED_OBSERVED |
First-party provenance: Google Gemini API model documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.
What are Gemini 3.1 Pro's specs?
| Context window | 2M tokens |
| Max output | 64K tokens |
| Modalities | text, vision, audio |
| Extended thinking | Yes |
| Released | 2026-02 |
| Knowledge cutoff | 2025-11 |
| Provider |
Verified 2026-08-14 — source.
Where does Gemini 3.1 Pro rank?
What are Gemini 3.1 Pro's strengths?
- Largest context window of any current model (2M tokens)
- Native audio and video understanding
- Google Search grounding
What else should you know about Gemini 3.1 Pro?
What are common questions about Gemini 3.1 Pro?
What is Gemini 3.1 Pro's context window?
Gemini 3.1 Pro has a 2M-token context window and a 64K-token max output — the 1st-largest context of the 42 current models we track. Source: https://ai.google.dev/gemini-api/docs/models, verified 2026-08-14.
Does Gemini 3.1 Pro support vision or audio input?
Yes — Gemini 3.1 Pro accepts vision and audio input in addition to text.
Does Gemini 3.1 Pro have a reasoning or extended-thinking mode?
Yes — Gemini 3.1 Pro exposes a dedicated reasoning mode for multi-step problems.
When was Gemini 3.1 Pro released, and what is its knowledge cutoff?
Gemini 3.1 Pro was released 2026-02 with a knowledge cutoff of 2025-11.
How much does Gemini 3.1 Pro cost, and who provides it?
Gemini 3.1 Pro is served by Google at $4.50/M blended tokens (3:1 input:output) — the 36th-cheapest of 42 current models. Full pricing breakdown: /llm-api-pricing/gemini-3-1-pro.
