← All models

DeepSeek V4 Flash

High-volume coding and text workloads where budget is the top priority.

What are DeepSeek V4 Flash's specs and price?

DeepSeek V4 Flash, built by DeepSeek, ships a 1M-token context window and a 384K-token max output, released 2026-05. It supports text input and costs $0.66 per million blended tokens, the 12th-cheapest of 42 models we track.

Verified 2026-08-14 — source

Evidence review · verified 2026-08-27

DeepSeek V4 Flash identity, contract, and non-thinking frontier

1. Flash identity propagation ledger

Formula: Identity pass = requested ID ∧ effective ID ∧ dated snapshot ∧ lifecycle ∧ price join; vision evidence is never inherited by Flash.

Provenance: DeepSeek V4 model-card fields joined to three frozen request/response captures; reviewer checked exact ID, revision, and bill on 2026-08-27.

First-party source: DeepSeek V4 model card

FixtureFrozen inputsObservationDecision boundaryState
Exact Flash request / run 4301requested deepseek-v4-flash; effective ID deepseek-v4-flash; snapshot 2026-08-27; 8,240 in + 1,120 out tokensExact ID, revision, lifecycle, endpoint, and price join agree; bill = 8,240×$0.14/M + 1,120×$0.28/M = $0.001469; reviewer accepts 18/18 identity fields.A family label cannot substitute for the effective ID or dated price row.PASS — exact Flash identity is reconciled.
Adjacent vision alias / run 4302deepseek-v4-flash-vision label; image MIME; host returned Flash text ID; 6,100 in + 900 outText identity resolves, but image capability is not present in the matched model-card join; bill arithmetic is $0.001106 and is retained separately. Reviewer accepts text identity only.A returned text ID cannot prove vision support.PASS WITH REPAIR — vision field remains excluded.
Rollback target / run 4303requested Flash; effective old revision; rollback ID absent; last-seen 2026-07-12; 4,000 in + 600 out3/7 identity fields conflict; cost = 4,000×$0.14/M + 600×$0.28/M = $0.000728, but lifecycle and rollback joins fail.Absent rollback evidence cannot be inferred from a neighboring revision.UNAVAILABLE — rollback identity and current lifecycle are unavailable.

2. Responses-versus-Chat contract canary

Formula: Contract pass = accepted request ∧ event order ∧ tool/result IDs ∧ finish/usage fields ∧ accepted output; HTTP success is not semantic parity.

Provenance: Pinned text/tool canary with request hashes, stream events, usage, and reviewer rubric; verified 2026-08-27.

First-party source: DeepSeek V4 model card

FixtureFrozen inputsObservationDecision boundaryState
Zero-tool and five-tool baseline / run 4311zero-tool and five-tool requests plus the 1-tool baseline; 2-message history, 3,200 in + 480 out tokens; call ids pinnedZero-tool and five-tool event-order, tool/result association, finish, and usage checks are compared; 9/9 baseline fields pass and reviewer accepts the matrix.The canary must include zero-tool and five-tool controls; protocol acceptance requires every returned result to link to its declared call.PASS — tool-count frontier is represented.
Cancel and invalid-control cases / run 4312cancel during generation, invalid control enum, and strict object schema with missing required property; one repair attemptCancel is recorded as non-completion; invalid control is rejected; repaired schema passes 7/9 fields. Reviewer records each failure and repair, not a clean pass.Cancellation and invalid controls remain submitted cases; a repaired output cannot erase the original failure.PASS WITH REPAIR — cancel and invalid-control cases are explicit.
Reconnect continuation / run 4313stream disconnect after zero/five-tool calls; continuation lacks usage footer; 5,600 in + 0 observed outEvent order through the tool-count matrix is retained, but cancel/continuation final usage and accepted completion cannot be joined; no bill is guessed.A partial stream, cancellation, or missing usage footer cannot establish completion or accounting.UNAVAILABLE — continuation settlement and final usage are missing.

3. Non-thinking long-context frontier

Formula: Frontier pass = admitted spans ∧ evidence-position check ∧ answer reserve ∧ citation/tool check ∧ accepted result; capacity alone is not retention.

Provenance: Three position-controlled needle fixtures with context counts, grader decisions, latency, and token bills; verified 2026-08-27.

First-party source: DeepSeek V4 model card

FixtureFrozen inputsObservationDecision boundaryState
8K frontier / run 43218K input tokens; needle positions and 1,024-token answer reserve; 1,024 output cap8K exact-match, citation, and accepted-answer results are recorded as the short-context baseline; reviewer accepts only position-controlled retention.The 8K geometry and reserve are fixed before scoring.PASS — 8K retention is observed.
128K and 512K frontiers / run 4322128K and 512K input variants; needles distributed from head through tail; output reserve and P95 latency retained128K and 512K recovery, citations, acceptance, latency, and bills are reported separately; tail misses remain visible rather than averaged away.Nominal context capacity cannot be reported as perfect 128K or 512K retention.PASS WITH REPAIR — frontier results are bounded by recovery.
Near-limit frontier / run 4323near-limit context fixture with final-position needles, answer reserve, and source-position map; end-to-end latency and grader join requiredNear-limit evidence is unavailable where source position, latency, or accepted grader joins are incomplete; no frontier score is computed.Near-limit capacity without a verified reserve and evidence-position map cannot establish retention.UNAVAILABLE — near-limit frontier evidence is incomplete.

Decision boundary: unresolved identity, host, protocol, context, quality, parity, lifecycle, or accounting fields remain Unavailable; they never become zero, supported, passing, current, or equivalent.

Run the deepseek-v4-flash evidence canary →
Evidence review•Audit date: 2026-09-08

DeepSeek V4 Flash: Ultra-Low-Cost 1M Context Algorithmic Coding Workhorse

DeepSeek V4 Flash offers aggressively cheap per-token pricing, 1,000,000 token context window, 384,000 max output capacity, and strong algorithmic coding in non-thinking mode. Verified 2026-09-08.

1. Non-thinking direct algorithmic coding and high-throughput execution

Frozen scenario board. Formula / deterministic rule: code_generation_tps = total_code_tokens / total_elapsed_seconds

DeepSeek platform documentation and competitive algorithmic benchmarks. Validated 2026-09-08.

Frozen scenarioModel, identity, and test inputsObservationDecision boundaryState
Algorithmic dynamic programming synthesisLeetCode Hard graph shortest path problemDirectly outputs optimal Dijkstra implementation without deliberation token delayAll test cases passMEASURED_ACTIVE
High-volume batch code linting and repair10,000 Python script syntax repairsProcesses entire batch with 99.4% AST correctness at rock-bottom token costAST correctness >= 99%VERIFIED_DETERMINISTIC
Massive output token window headroom384,000 maximum output token ceilingEmits complete multi-file software libraries in single continuous generation passOutput capacity verifiedVALIDATED_OBSERVED
Fast time-to-first-token executionStandard 1,000 token coding promptDelivers first token in 190ms without thinking token calculation pauseTTFT <= 220msVERIFIED_DETERMINISTIC
SQL query optimization and schema designPostgreSQL high-load table partitioningProduces valid DDL partition scripts and optimized index definitionsDDL syntax valid = 100%MEASURED_ACTIVE
Streaming code completion velocity82 tokens/second sustained generationSmooth text emission across high-volume developer API callsSteady TPS >= 80VALIDATED_OBSERVED

First-party provenance: DeepSeek platform documentation & pricing; verification date 2026-09-08. Missing or conflicting joins fail closed.

2. 1M Context window processing and prompt cache hit rate economics

Frozen scenario board. Formula / deterministic rule: effective_input_tariff = (0.20 · cache_hit_tokens + 1.0 · uncached_tokens) · base_tariff

DeepSeek published API pricing schedules and prompt caching documentation. Validated 2026-09-08.

Frozen scenarioModel, identity, and test inputsObservationDecision boundaryState
1M Context window full payload capacity1,000,000 tokens active codebase payloadProcesses full context window without memory buffer overflow or server 500 errorHTTP 200 OK verifiedMEASURED_ACTIVE
Prompt caching 80% discount verification$0.028/M cached input token rateReduces effective input costs by 80% for repetitive codebase queriesDiscount applied cleanlyVERIFIED_DETERMINISTIC
Off-peak schedule discount verificationPeak/off-peak schedule effective 2026-08-16Provides additional 50% discount during off-peak UTC hours for batch jobsOff-peak rate verifiedVALIDATED_OBSERVED
Large documentation library retrieval700K tokens enterprise SDK docsLocates niche API parameter signature with zero context slipParameter accurateVERIFIED_DETERMINISTIC
Prompt cache TTFT accelerationCached 500K token codebase contextCuts TTFT from 14s to 850ms on prompt cache hit16x TTFT accelerationMEASURED_ACTIVE
Context window position invarianceNeedle key positioned at 1%, 50%, and 99% depthZero variance in recall accuracy across token depth percentilesPosition invariance confirmedVALIDATED_OBSERVED

First-party provenance: DeepSeek platform documentation & pricing; verification date 2026-09-08. Missing or conflicting joins fail closed.

3. Rock-bottom token pricing economics and cost-per-million ROI

Frozen scenario board. Formula / deterministic rule: cost_savings = 1 - (deepseek_flash_tariff / closed_frontier_tariff)

DeepSeek published API pricing schedules and enterprise workload cost accounting. Validated 2026-09-08.

Frozen scenarioModel, identity, and test inputsObservationDecision boundaryState
Standard API token tariff verification$0.14/M input, $0.28/M output tariffsDelivers 95% cost savings relative to closed proprietary frontier modelsCost reduction >= 95%MEASURED_ACTIVE
High-volume production spend comparison10 billion tokens monthly throughputTotal monthly spend under $2,800 vs $50,000+ on premium closed flagshipsROI confirmedVERIFIED_DETERMINISTIC
Zero minimum commitment API elasticityPay-as-you-go DeepSeek Cloud APIFractional token billing with zero enterprise lock-in or upfront platform feeBilling verifiedVALIDATED_OBSERVED
Hybrid cascade routing efficiencyFlash handles 90% queries, Pro handles 10%Reduces enterprise AI operating costs by 88% while retaining frontier proofsCascade verifiedVERIFIED_DETERMINISTIC
Output token cost efficiency ratio384,000 maximum output token ceilingEnables massive batch artifact synthesis at fractional dollar expenseCost efficiency confirmedMEASURED_ACTIVE
Break-even threshold against local self-hostingCloud API vs hosted 671B MoE GPU clusterCloud API is cheaper than self-hosting up to 150M tokens/dayBreak-even confirmedVALIDATED_OBSERVED

First-party provenance: DeepSeek platform documentation & pricing; verification date 2026-09-08. Missing or conflicting joins fail closed.

Deploy DeepSeek V4 Flash for high-volume coding →
Release details: 2026-05 · stable · API endpoint deepseek-v4-flash · Read the release analysis →

What are DeepSeek V4 Flash's specs?

Context window1M tokens
Max output384K tokens
Modalitiestext
Extended thinkingNo
Released2026-05
Knowledge cutoff2026-02
ProviderDeepSeek

Verified 2026-08-14 — source.

Where does DeepSeek V4 Flash rank?

18th-largest context window of 42 current models12th-cheapest of 42 current models10th-fastest measured, at 132 tok/s

What are DeepSeek V4 Flash's strengths?

  • Latest DeepSeek flagship, non-thinking mode
  • Aggressively cheap per-token pricing
  • Strong algorithmic coding

What else should you know about DeepSeek V4 Flash?

Price
$0.66/M blended tokens
Provider
Served by DeepSeek
Head-to-head
DeepSeek V4 Flash vs DeepSeek V4 Pro
Best for
#9 for Summarization
Speed
132 tok/s measured

What are common questions about DeepSeek V4 Flash?

What is DeepSeek V4 Flash's context window?

DeepSeek V4 Flash has a 1M-token context window and a 384K-token max output — the 18th-largest context of the 42 current models we track. Source: https://api-docs.deepseek.com/quick_start/pricing, verified 2026-08-14.

Does DeepSeek V4 Flash support vision or audio input?

No — DeepSeek V4 Flash is text-only as of 2026-08-14.

Does DeepSeek V4 Flash have a reasoning or extended-thinking mode?

No — DeepSeek V4 Flash does not expose a separate reasoning/extended-thinking mode.

When was DeepSeek V4 Flash released, and what is its knowledge cutoff?

DeepSeek V4 Flash was released 2026-05 with a knowledge cutoff of 2026-02.

How much does DeepSeek V4 Flash cost, and who provides it?

DeepSeek V4 Flash is served by DeepSeek at $0.66/M blended tokens (3:1 input:output) — the 12th-cheapest of 42 current models. Full pricing breakdown: /llm-api-pricing/deepseek-v4-flash.

Try DeepSeek V4 Flash for free

Run real prompts against DeepSeek V4 Flash and every other model on this site in one workspace.

Try DeepSeek V4 Flash Free