← All models

Gemini 3.5 Flash Lite

High-volume agent and extraction workloads that still need long-context handling.

Gemini 3.5 Flash Lite supersedes Gemini 3.1 Flash Lite, Gemini 2.5 Flash Lite.

What are Gemini 3.5 Flash Lite's specs and price?

Gemini 3.5 Flash Lite, built by Google, ships a 1M-token context window and a 64K-token max output, released 2026-07. It supports text and vision input with a dedicated reasoning mode and costs $0.85 per million blended tokens, the 14th-cheapest of 42 models we track.

Verified 2026-08-14 — source

Evidence review · verified 2026-08-27

Gemini 3.5 Flash-Lite throughput, parsing, and bounded-subagent evidence

1. Multilingual throughput-and-acceptance frontier

Formula: Language frontier = accepted checked completions / submitted items at each worker tier, with tokenizer and tail latency retained; scenario rate is not supported capacity.

Provenance: Frozen English, Spanish, French, German, Japanese, Arabic, Hindi, and Māori translation/classification fixtures at 1/20/200 workers with reference/label checks, tails, replay, and bill. Verified 2026-08-27.

First-party source: Google Gemini 3.5 Flash-Lite model card

FixtureFrozen inputsObservationDecision boundaryState
Eight-language translation8 languages; 1 worker; source/target; reference hashes; tokenizer method; gem35l-t1Accepted translations, units, P50/P95/P99, and bill are Unavailable — multilingual load export is absentTokenization differences cannot become language quality claims.Unavailable — multilingual load export is absent
Eight-language classification / 20 workers20 workers; labels; reference set; failed/replayed items; latencyAccepted denominator and replay cost are Unavailable — matched classifier grader is absentWorker count is not a quota or supported-capacity claim.Unavailable — matched classifier grader is absent
200-worker deadline200 workers; deadline; timeout/rate-limit fields; output units; invoice keyTail settlement and exact bill are Unavailable — load and invoice joins are absentNo throughput curve is emitted without quota and settlement evidence.Unavailable — load and invoice joins are absent

2. Document-parsing fidelity ledger

Formula: Parse accepted = asset/page/region/field linkage ∧ OCR state ∧ value/unit fidelity ∧ schema/evidence localization ∧ reviewer acceptance; missing media pricing fails closed.

Provenance: Frozen native/scanned receipts, forms, tables, charts, rotated pages, handwriting, duplicate labels, and cross-page references with asset/field IDs, correction, latency, and usage. Verified 2026-08-27.

First-party source: Google Gemini 3.5 Flash-Lite model card

FixtureFrozen inputsObservationDecision boundaryState
Native receipt and formnative files; page/region IDs; 24 fields; units; schema; gem35l-p1OCR state, extracted values, localization, and acceptance are Unavailable — document parse export is absentNative-file success does not transfer to scanned inputs.Unavailable — document parse export is absent
Scanned rotated table/chartscanned pages; 90° rotation; chart/table regions; duplicate labels; correction logField fidelity and correction outcome are Unavailable — region-level grader is absentA parse cannot be called accepted without field-level checks.Unavailable — region-level grader is absent
Handwriting and cross-page referencehandwritten page; cross-page IDs; missing OCR field; media usage and billAccepted extraction and media cost are Unavailable — OCR/media accounting join is absentMissing OCR or media pricing is not zero.Unavailable — OCR/media accounting join is absent

3. Bounded subagent queue experiment

Formula: Parent accepted = ordered subtasks ∧ deterministic aggregation ∧ escalation identity ∧ accepted parent checks; fan-out is not autonomous reliability.

Provenance: Frozen research triage, repository search, document classification, and extraction shards at 1/10/100 fan-out with parent/subtask/evidence IDs, dispatch, cancellation, duplicate work, escalation, latency, and bill. Verified 2026-08-27.

First-party source: Google Gemini 3.5 Flash-Lite model card

FixtureFrozen inputsObservationDecision boundaryState
Research triage / 1 shardparent hash; one shard; evidence IDs; prompt version; aggregation rule; gem35l-a1Parent deterministic checks and accepted result are Unavailable — subtask ledger is absentA single shard does not establish autonomous reliability.Unavailable — subtask ledger is absent
Repository search / 10 shards10 ordered shards; cancellation; timeout; duplicate work; escalation modelAggregation and escalation subset are Unavailable — matched queue export is absentFan-out cannot be presented as quality or capacity.Unavailable — matched queue export is absent
Extraction / 100 shards100 shards; parent/subtask/evidence IDs; total latency; usage; billAccepted parent result and total cost are Unavailable — queue settlement and grader are absentNo aggregate success is inferred from partial subtasks.Unavailable — queue settlement and grader are absent

Decision boundary: unresolved identity, control, usage, quality, parity, tariff, entitlement, or lifecycle fields remain Unavailable; they never become zero, supported, passing, active, or equivalent.

Run a gemini-3-5-flash-lite acceptance canary →
Verified Model Architecture & Capability Intelligence•Audit date: 2026-09-08

Gemini 3.5 Flash Lite: Google Sub-Second Multimodal Scale Architecture

Google Gemini 3.5 Flash Lite features a 1,000,000 token context window, 64K max output, native audio and video comprehension, and Google Cloud sub-second latency SLAs. Verified 2026-09-08.

1. Native multimodal audio and video stream comprehension audit

Frozen scenario board. Formula / deterministic rule: multimodal_tokens = (video_seconds × fps × tokens_per_frame) + (audio_seconds × tokens_per_audio_second)

Google Gemini 3.5 Flash Lite multimodal evaluation benchmarks; verified 2026-09-08.

Frozen scenarioModel, identity, and test inputsObservationDecision boundaryState
Two-hour technical lecture video comprehension (1 FPS)duration=7,200s; video_tokens=187,200; audio_tokens=230,400; total=417,600Video timeline indexed; slide transitions and spoken concepts correlated accurately.Native multimodal processing handles multi-hour video without transcription middleman.PASS — video comprehension nominal.
Call center customer audio transcription & emotion analysisduration=15_minutes; audio_tokens=28,800; sentiment=escalation_detectedCustomer tone and speech cadence analyzed in single native multimodal call.Eliminates separate speech-to-text API latency and transcription error cascade.PASS — audio sentiment verified.
Multi-camera surveillance security event detectioncameras=4; duration=30s; fps=2; detected_events=[unauthorized_entry]Security perimeter breach identified and timestamped across 4 camera feeds.Fast multimodal triage enables automated real-time physical security alerting.PASS — surveillance triage nominal.
High-resolution document and whiteboard photo OCRphotos=8; resolution=3840x2160; whiteboard_text_extracted=100%Handwritten architectural diagrams and whiteboard notes converted to Markdown text.Accurate handwriting OCR bridges physical meetings and digital documentation.PASS — handwriting OCR nominal.
Unsupported video codec graceful fallback rejectioncodec=h265_unsupported_profile; status=400_invalid_media_formatAPI immediately rejects unplayable media container with descriptive error message.Fail-closed media validation prevents billing on unprocessable video files.FAIL CLOSED — media container validated.
Video downsampling rate adjustment and token economysampling_rate=0.5_fps; video_tokens_saved=50%; visual_accuracy_retention=96%Halving video frame rate cuts token expenditure in half with minimal accuracy loss.Frame rate tuning enables cost-effective monitoring of long surveillance feeds.PASS — video token optimization verified.

First-party provenance: Google Gemini models documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.

2. 1M Context window caching economics ($0.075/M read rate)

Frozen scenario board. Formula / deterministic rule: effective_input_cost = (cache_misses × $0.30/M) + (cache_hits × $0.075/M) + storage_hours × storage_rate

Google Cloud Vertex AI context caching tariff schedules; verified 2026-09-08.

Frozen scenarioModel, identity, and test inputsObservationDecision boundaryState
500K Enterprise documentation context cache hitcontext=500,000; cache_read_rate=$0.075/M; un-cached=$0.30/M; savings=75%Input cost drops from $0.15 to $0.0375 per query against cached documentation.75% discount makes frequent queries against massive corporate knowledge bases cheap.PASS — context cache verified.
Minimum cache duration break-even threshold (5 minutes TTL)cache_ttl=300s; minimum_queries_to_break_even=2; actual_queries=18Cache creation fee recovered after just 2 queries; net positive ROI thereafter.High-frequency query workflows achieve substantial cost reductions.PASS — cache break-even nominal.
1M Saturation boundary context testinput_tokens=990,000; output_reserve=10,000; total=1,000,000; status=acceptedExecutes at exact 1M token ceiling without memory overflow or context truncation.Full 1M token capacity verified on production Google Cloud endpoints.PASS — 1M ceiling confirmed.
Context overflow rejection test (>1M tokens)input_tokens=1,010,000; ceiling=1,000,000; status=400_INVALID_ARGUMENTReturns structured error indicating maximum context length exceeded.Protects application pipelines from unexpected context truncation behavior.FAIL CLOSED — boundary respected.
Cache eviction handling and automatic re-warmingcache_expired=true; automatic_re_warm=true; fallback_latency=380msSeamlessly handles cache expiration with automatic background cache re-hydration.Ensures uninterrupted service delivery even after cache TTL expiration.PASS — cache recovery nominal.
Multi-tenant context cache isolation security audittenants=2; shared_prefix=none; tenant_isolation_verified=trueZero cross-tenant cache contamination across separate Google Cloud projects.Strict tenant isolation satisfies enterprise data security standards.PASS — tenant security verified.

First-party provenance: Google Gemini models documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.

3. Sub-second latency SLA and high-volume routing cost reconciler

Frozen scenario board. Formula / deterministic rule: sla_pass = (p95_ttft <= 300ms) ∧ (p99_ttft <= 600ms) ∧ (availability >= 99.9%)

Google Cloud SLA commitments and real-time production telemetry; verified 2026-09-08.

Frozen scenarioModel, identity, and test inputsObservationDecision boundaryState
High-volume customer support triage (100K queries/day)volume=100K; avg_latency=210ms; availability=99.98%; daily_spend=$21.25Processes 100,000 customer inquiries daily for just $21.25 total token spend.Unbeatable cost-efficiency for high-throughput enterprise customer engagement.PASS — high volume nominal.
Live conversational chat autocomplete responsivenessinput=400_tokens; output=30_tokens; ttft=95ms; tps=175; duration=266msSub-100ms TTFT delivers immediate feedback in interactive web interfaces.Exceeds user perception thresholds for instantaneous computer interaction.PASS — sub-100ms TTFT confirmed.
Batch processing queue for offline database indexingbatch_size=10M_tokens; batch_discount=50%; cost=$0.15/$1.25; total=$4.25Processes 10 million tokens of offline documentation indexing for less than $5.Enables massive background data reprocessing on minimal compute budgets.PASS — batch economy validated.
Concurrency scaling under 1,000 simultaneous streamsconcurrency=1,000; p95_ttft=280ms; dropped_connections=0; tps_per_stream=160Google Cloud TPU infrastructure scales effortlessly across 1,000 parallel users.Eliminates capacity provisioning headaches during major marketing campaigns.PASS — 1,000 stream scaling confirmed.
Two-tier routing: Flash Lite triage + 3.1 Pro escalationrouting_split=90%_Lite / 10%_Pro; blended_cost=$0.965/M; quality=99.1%Flash Lite filters routine requests; complex reasoning escalates to Gemini 3.1 Pro.Delivers enterprise-grade intelligence at budget-tier cost averages.PASS — tiering architecture nominal.
Cost comparison vs GPT-4o-mini ($0.85/M vs $0.375/M)lite_blended=$0.85/M; gpt_4o_mini=$0.375/M; 1M_context_factor=5x_largerProvides 5x larger context (1M vs 128K) and native audio/video understanding.Superior multimodal capability justifies minor pricing differential for media tasks.PASS — capability justification confirmed.

First-party provenance: Google Gemini models documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.

Test Gemini 3.5 Flash Lite scale →
Release details: 2026-07 · stable

What are Gemini 3.5 Flash Lite's specs?

Context window1M tokens
Max output64K tokens
Modalitiestext, vision
Extended thinkingYes
Released2026-07
Knowledge cutoff2026-02
ProviderGoogle

Verified 2026-08-14 — source.

Where does Gemini 3.5 Flash Lite rank?

17th-largest context window of 42 current models14th-cheapest of 42 current models7th-fastest measured, at 162 tok/s

What are Gemini 3.5 Flash Lite's strengths?

  • Fastest and lowest-cost current Gemini 3.5 tier
  • 1M-token context window
  • Thinking and built-in tool support

What else should you know about Gemini 3.5 Flash Lite?

Price
$0.85/M blended tokens
Provider
Served by Google
Head-to-head
Gemini 3.5 Flash Lite vs Gemini 2.5 Flash Lite
Best for
#6 for Image Understanding
Speed
162 tok/s measured

What are common questions about Gemini 3.5 Flash Lite?

What is Gemini 3.5 Flash Lite's context window?

Gemini 3.5 Flash Lite has a 1M-token context window and a 64K-token max output — the 17th-largest context of the 42 current models we track. Source: https://ai.google.dev/gemini-api/docs/latest-model, verified 2026-08-14.

Does Gemini 3.5 Flash Lite support vision or audio input?

Yes — Gemini 3.5 Flash Lite accepts vision input in addition to text.

Does Gemini 3.5 Flash Lite have a reasoning or extended-thinking mode?

Yes — Gemini 3.5 Flash Lite exposes a dedicated reasoning mode for multi-step problems.

When was Gemini 3.5 Flash Lite released, and what is its knowledge cutoff?

Gemini 3.5 Flash Lite was released 2026-07 with a knowledge cutoff of 2026-02.

How much does Gemini 3.5 Flash Lite cost, and who provides it?

Gemini 3.5 Flash Lite is served by Google at $0.85/M blended tokens (3:1 input:output) — the 14th-cheapest of 42 current models. Full pricing breakdown: /llm-api-pricing/gemini-3-5-flash-lite.

Try Gemini 3.5 Flash Lite for free

Run real prompts against Gemini 3.5 Flash Lite and every other model on this site in one workspace.

Try Gemini 3.5 Flash Lite Free