← Back to all pricing

DeepSeek V4 Flash API Pricing: Ultra-Low-Cost High-Speed Intelligence

Comprehensive DeepSeek V4 Flash API pricing analysis ($0.44/M input, $1.32/M output), off-peak discounts, classification speed, and prompt caching breaks.

Full specs, context window and API limits →

How much does DeepSeek V4 Flash cost per million tokens?

DeepSeek V4 Flash costs $0.44 per million input tokens and $1.32 per million output tokens ($0.66/M blended at 3:1). A high-efficiency open-weights model designed for rapid classification, document summarization, and large-scale data structuring. Verified 2026-09-08.

Verified 2026-09-07 — source
Input
$0.44/M
Output
$1.32/M
Blended
$0.66/M
Provider
Verified 2026-08-14 — source

How much does DeepSeek V4 Flash cost per 1,000 requests?

Computed from generated token pricing. Each row assumes the listed input and output tokens per request; output is adjusted by this model's measured 2.59× verbosity factor.

Request shapeInput tokensOutput tokensCost / 1,000 requests
Short10050$0.2149
Medium1,000500$2.1494
Long4,0002,000$8.5976

Formula: ((input price × input tokens) + (output price × output tokens × verbosity factor)) ÷ 1,000,000 × 1,000. Assumptions: short 100/50, medium 1,000/500, long 4,000/2,000 input/output tokens per request. Verbosity run: 2026-06-16T20:31:30.728Z.

Three model-specific pricing decisions

Flash is non-thinking mode, so the schedule and accepted-answer boundary stay separate from provider-wide scheduling and the broad Flash/Pro comparison.

1. UTC peak/off-peak plus cache-hit monthly schedule

UTC trafficCache statusFormulaDecision
Peak allocation pHit/miss separatep × peak + (1 − p) × off-peakMeasure p in UTC
Off-peak allocation 1 − pHit/miss separate100K calls × applicable input/output ratesSchedule flexible work
Dated rateModel price recordCurrent Flash rate onlyProvider schedule source required for numeric discount

2. Flash versus Pro cost-per-accepted-answer boundary

ModeToken cost / 100KAccepted-answer rateBoundary
DeepSeek V4 Flash$237.60UnavailableThinking premium cannot be inferred
DeepSeek V4 Pro$712.80UnavailableRecord accepted answers before switching

3. Operational sensitivity

VariableDated valueKeep separate from token price
Retry rateUnavailableMultiply full request cost when measured
LatencyUnavailableDo not convert milliseconds to dollars
SLA / availabilityUnavailableNo SLA claim in pricing record

Verified 2026-08-14. “Unavailable” means the current dated registry has no model-specific evidence; it is not a zero. First-party price source · Run this scenario in the playground.

All results are server-rendered for DeepSeek V4 Flash; formulas expose fixed inputs and missing evidence remains visibly unavailable.

Exact-model pricing guide · verified 2026-09-07

Exact model boundary: DeepSeek DeepSeek V4 Flash (deepseek-v4-flash). Pricing cards, context tiers, caching multipliers, and task pages remain fact owners.

Prompt cache hit versus cache miss cost ledger

Frozen scenario board. Formula / deterministic rule: cost = (cache_miss_in * 0.44 + cache_hit_in * 0.11 + out * 1.32) / 1M; cache hit saves 75% Boundary: Owns DeepSeek V4 Flash prompt caching economics.

Frozen scenarioExact identity and evidence fieldsResultState
100% cache miss cold promptmodel=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=100% cache miss cold prompt; prompt tokens; cache hit percentage; output tokens; net cost; effective input rate; savings percentage; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — 100% cache miss cold prompt is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
50% cache hit warm promptmodel=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=50% cache hit warm prompt; prompt tokens; cache hit percentage; output tokens; net cost; effective input rate; savings percentage; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — 50% cache hit warm prompt is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
80% cache hit enterprise system promptmodel=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=80% cache hit enterprise system prompt; prompt tokens; cache hit percentage; output tokens; net cost; effective input rate; savings percentage; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — 80% cache hit enterprise system prompt is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
95% cache hit document Q&A loopmodel=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=95% cache hit document Q&A loop; prompt tokens; cache hit percentage; output tokens; net cost; effective input rate; savings percentage; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — 95% cache hit document Q&A loop is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
cache eviction on cold startmodel=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=cache eviction on cold start; prompt tokens; cache hit percentage; output tokens; net cost; effective input rate; savings percentage; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — cache eviction on cold start is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
unsupported multi-part schemamodel=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=unsupported multi-part schema; prompt tokens; cache hit percentage; output tokens; net cost; effective input rate; savings percentage; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — unsupported multi-part schema has no matched, dated bilateral observation.FAIL CLOSED — manual, probe, or source evidence required

First-party provenance: DeepSeek API official pricing; verification date 2026-09-07. Missing or conflicting joins fail closed.

High-volume classification and ETL pipeline budgeting

Frozen scenario board. Formula / deterministic rule: pipeline_cost = records * ((doc_tokens * 0.44 + json_out * 1.32) / 1M) Boundary: Owns high-volume data pipeline economics for DeepSeek V4 Flash.

Frozen scenarioExact identity and evidence fieldsResultState
100K structured web scrape recordsmodel=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=100K structured web scrape records; record count; avg input tokens; avg output tokens; monthly cost; cost per 1K records; pipeline status; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — 100K structured web scrape records is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
500K customer sentiment reviewsmodel=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=500K customer sentiment reviews; record count; avg input tokens; avg output tokens; monthly cost; cost per 1K records; pipeline status; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — 500K customer sentiment reviews is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
1M log classification eventsmodel=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=1M log classification events; record count; avg input tokens; avg output tokens; monthly cost; cost per 1K records; pipeline status; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — 1M log classification events is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
10M enterprise data enrichment batchmodel=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=10M enterprise data enrichment batch; record count; avg input tokens; avg output tokens; monthly cost; cost per 1K records; pipeline status; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — 10M enterprise data enrichment batch is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
rate limit throttling backupmodel=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=rate limit throttling backup; record count; avg input tokens; avg output tokens; monthly cost; cost per 1K records; pipeline status; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — rate limit throttling backup is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
unparseable JSON output retrymodel=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=unparseable JSON output retry; record count; avg input tokens; avg output tokens; monthly cost; cost per 1K records; pipeline status; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — unparseable JSON output retry is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict

First-party provenance: DeepSeek API documentation; verification date 2026-09-07. Missing or conflicting joins fail closed.

DeepSeek V4 Flash vs Pro thinking mode economic trade-off

Frozen scenario board. Formula / deterministic rule: cost_ratio = pro_cost / flash_cost; evaluates when reasoning tokens justify cost Boundary: Owns Flash vs Pro model routing decision boundaries.

Frozen scenarioExact identity and evidence fieldsResultState
simple classification (Flash optimal)model=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=simple classification (Flash optimal); workload type; Flash cost; Pro cost; cost multiplier; accuracy delta threshold; routing recommendation; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — simple classification (Flash optimal) is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
data formatting & translation (Flash optimal)model=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=data formatting & translation (Flash optimal); workload type; Flash cost; Pro cost; cost multiplier; accuracy delta threshold; routing recommendation; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — data formatting & translation (Flash optimal) is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
multi-step logical reasoning (Pro optimal)model=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=multi-step logical reasoning (Pro optimal); workload type; Flash cost; Pro cost; cost multiplier; accuracy delta threshold; routing recommendation; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — multi-step logical reasoning (Pro optimal) is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
complex math verification (Pro optimal)model=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=complex math verification (Pro optimal); workload type; Flash cost; Pro cost; cost multiplier; accuracy delta threshold; routing recommendation; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — complex math verification (Pro optimal) is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
hybrid router: 80% Flash / 20% Promodel=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=hybrid router: 80% Flash / 20% Pro; workload type; Flash cost; Pro cost; cost multiplier; accuracy delta threshold; routing recommendation; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — hybrid router: 80% Flash / 20% Pro is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
untested reasoning requirementmodel=deepseek-v4-flash; provider=DeepSeek; slug=deepseek-v4-flash; scenario=untested reasoning requirement; workload type; Flash cost; Pro cost; cost multiplier; accuracy delta threshold; routing recommendation; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — untested reasoning requirement is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict

First-party provenance: DeepSeek API official pricing; verification date 2026-09-07. Missing or conflicting joins fail closed.

Method and limitations: formulas are deterministic; observed and assumed inputs are labeled; no missing provider, host, account, region, realm, alias, snapshot, revision, weight, artifact, control, tool, modality, workload, rate-period, timestamp, or result is transferred. Run this scenario →

Evidence audit · 2026-09-08

DeepSeek V4 Flash API Pricing: Ultra-Low-Cost High-Speed Intelligence

DeepSeek V4 Flash costs $0.44 per million input tokens and $1.32 per million output tokens ($0.66/M blended at 3:1). A high-efficiency open-weights model designed for rapid classification, document summarization, and large-scale data structuring. Verified 2026-09-08.

Module 1 · DeepSeek V4 Flash Micro-Cost Token Economics
Blended Cost = (Input Tokens × $0.44 + Output Tokens × $1.32) / 1,000,000

DeepSeek V4 Flash delivers high-throughput utility inference at $0.66/M blended tokens.

Boundary: Standard pay-as-you-go rate card; off-peak window and prompt caching discounts apply.
ScenarioRendered Evidence & Bounds
Scenario 1Customer ticket intent classification (800 in, 50 out): $0.000418 per ticket
Scenario 2E-commerce product specification extraction (1.5K in, 200 out): $0.000924 per product
Scenario 3Customer feedback sentiment scoring (2K in, 100 out): $0.001012 per review
Scenario 4Document summary paragraph generation (4K in, 300 out): $0.002156 per document
Scenario 5High-volume webhook data normalization (1K in, 100 out): $0.000572 per webhook
Scenario 6Monthly 100M token classification fleet: $66.00 total infrastructure spend
Module 2 · DeepSeek V4 Flash Off-Peak Tariff & Cache Amortization
Discounted Cost = (Cached Input × $0.044 + Off-Peak Tokens × 0.50 + Generation × $1.32) / 1,000,000

Combining off-peak execution with prompt caching delivers industry-leading data enrichment economics.

Boundary: Evaluates 90% prompt caching discount and 50% off-peak tariff reduction during off-peak hours.
ScenarioRendered Evidence & Bounds
Scenario 1Off-peak asynchronous data tagging: cuts baseline token prices by exactly 50%
Scenario 2Shared JSON extraction schema cache (8K tokens): 82% input cost reduction
Scenario 3Combined off-peak + prompt cache: effective blended rate drops below $0.25/M
Scenario 4Zero cache retention storage fees charged during active continuous sessions
Scenario 5Enables massive bulk data enrichment across multi-million record databases
Scenario 6Reduces enterprise data structuring operational costs by over 70%
Module 3 · DeepSeek V4 Flash High-Volume Batch Extraction Pipeline
Batch Efficiency = Throughput (tokens/sec) / Blended Price ($/M)

Exceptional token throughput pairs with rock-bottom pricing for high-volume enterprise ETL.

Boundary: Evaluates throughput optimization and concurrency scaling for automated enterprise workflows.
ScenarioRendered Evidence & Bounds
Scenario 1Processes 10,000 customer survey responses for under $5.00 total API cost
Scenario 2High streaming throughput (120+ tps) ensures zero queue delays on API gateways
Scenario 3Strict JSON mode adherence guarantees zero downstream serialization pipeline errors
Scenario 4Low-memory footprint supports massive concurrent connection limits on shared infrastructure
Scenario 5Reliable instruction following on complex multi-field schema extraction tasks
Scenario 6Ideal operational choice for high-volume ETL pipelines and real-time content moderation
Explore Related Analyses:DeepSeek provider profile →Compare vs DeepSeek V4 Pro →Compare vs GPT-5.4 Nano →Cheapest AI API comparison →
Exact-Model Pricing Evidence•Verified 2026-08-14; revalidation required before current claims.

DeepSeek V4 Flash pricing evidence

Source-backed dated rate shape: dated input=$0.440000/M; output=$1.320000/M; cache input=$0.014000/M; off-peak input=$0.220000/M, output=$0.660000/M. DeepSeek V4 Flash rate, off-peak, and ETL arithmetic only; DeepSeek policy, Pro reasoning, and cheapest-model rankings retain their owners. Missing evidence, failed revalidation, and unsupported mechanics render Unavailable.

Module 1 of 3: Extraction concurrency bill ladder

Novel contribution boundary: Owns dated Flash peak-rate arithmetic for fixed extraction shapes; concurrency and throughput are not asserted. Formula / deterministic rule: spend = requests × (inputTokens × inputCostPer1k + outputTokens × outputCostPer1k) / 1,000

ScenarioExact model, provider, dated rates, and fixed inputsResultState
100K compact ETL requestsmodel=deepseek-v4-flash; provider=deepseek; registryRates=dated input=$0.440000/M; output=$1.320000/M; cache input=$0.014000/M; off-peak input=$0.220000/M, output=$0.660000/M; comparisonModel=deepseek-v4-pro; fixedInputs=requests=100,000; inputTokens=500; outputTokens=150; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicitUnavailable — Last verified 2026-08-14, before revalidation date 2026-09-14.FAIL CLOSED — no revalidated observation or provider guarantee
1M structured extraction requestsmodel=deepseek-v4-flash; provider=deepseek; registryRates=dated input=$0.440000/M; output=$1.320000/M; cache input=$0.014000/M; off-peak input=$0.220000/M, output=$0.660000/M; comparisonModel=deepseek-v4-pro; fixedInputs=requests=1,000,000; inputTokens=1,000; outputTokens=300; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicitUnavailable — Last verified 2026-08-14, before revalidation date 2026-09-14.FAIL CLOSED — no revalidated observation or provider guarantee
10M high-volume rowsmodel=deepseek-v4-flash; provider=deepseek; registryRates=dated input=$0.440000/M; output=$1.320000/M; cache input=$0.014000/M; off-peak input=$0.220000/M, output=$0.660000/M; comparisonModel=deepseek-v4-pro; fixedInputs=requests=10,000,000; inputTokens=2,000; outputTokens=500; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicitUnavailable — Last verified 2026-08-14, before revalidation date 2026-09-14.FAIL CLOSED — no revalidated observation or provider guarantee

First-party provenance: https://api-docs.deepseek.com/quick_start/pricing; registry provider deepseek; verified 2026-08-14; freshness gate=2026-09-14. Revalidate before any current-price claim.

Module 2 of 3: Flash→Pro thinking-output premium

Novel contribution boundary: Owns dated model-tier arithmetic with explicit retry input; thinking uplift and accepted quality are not sourced. Formula / deterministic rule: requiredAcceptedRate = ProAttemptCost / FlashAttemptCost; observed acceptance = Unavailable

ScenarioExact model, provider, dated rates, and fixed inputsResultState
500-token ETL pass; 10% retry share; 100 retry-input tokensmodel=deepseek-v4-flash; provider=deepseek; registryRates=dated input=$0.440000/M; output=$1.320000/M; cache input=$0.014000/M; off-peak input=$0.220000/M, output=$0.660000/M; comparisonModel=deepseek-v4-pro; fixedInputs=requests=1; inputTokens=500; outputTokens=150; retryInputTokens=100; retryShare=10%; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicitUnavailable — Last verified 2026-08-14, before revalidation date 2026-09-14.FAIL CLOSED — no revalidated observation or provider guarantee
1K-token extraction; 20% retry share; 250 retry-input tokensmodel=deepseek-v4-flash; provider=deepseek; registryRates=dated input=$0.440000/M; output=$1.320000/M; cache input=$0.014000/M; off-peak input=$0.220000/M, output=$0.660000/M; comparisonModel=deepseek-v4-pro; fixedInputs=requests=1; inputTokens=1,000; outputTokens=300; retryInputTokens=250; retryShare=20%; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicitUnavailable — Last verified 2026-08-14, before revalidation date 2026-09-14.FAIL CLOSED — no revalidated observation or provider guarantee
2K-token long row; 30% retry share; 500 retry-input tokensmodel=deepseek-v4-flash; provider=deepseek; registryRates=dated input=$0.440000/M; output=$1.320000/M; cache input=$0.014000/M; off-peak input=$0.220000/M, output=$0.660000/M; comparisonModel=deepseek-v4-pro; fixedInputs=requests=1; inputTokens=2,000; outputTokens=500; retryInputTokens=500; retryShare=30%; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicitUnavailable — Last verified 2026-08-14, before revalidation date 2026-09-14.FAIL CLOSED — no revalidated observation or provider guarantee

First-party provenance: https://api-docs.deepseek.com/quick_start/pricing; registry provider deepseek; verified 2026-08-14; freshness gate=2026-09-14. Revalidate before any current-price claim.

Module 3 of 3: UTC and cache evidence ledger

Novel contribution boundary: Owns exact-model DeepSeek registry evidence only; cache TTL, quota, and off-peak eligibility mechanics are not inferred. Formula / deterministic rule: evidence = exact-model registry field when present; cache TTL, quota, and UTC eligibility = Unavailable

ScenarioExact model, provider, dated rates, and fixed inputsResultState
UTC off-peak windowmodel=deepseek-v4-flash; provider=deepseek; registryRates=dated input=$0.440000/M; output=$1.320000/M; cache input=$0.014000/M; off-peak input=$0.220000/M, output=$0.660000/M; comparisonModel=deepseek-v4-pro; fixedInputs=evidenceField=off-peak; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicitUnavailable — Last verified 2026-08-14, before revalidation date 2026-09-14.FAIL CLOSED — no revalidated observation or provider guarantee
Prompt-cache input ratemodel=deepseek-v4-flash; provider=deepseek; registryRates=dated input=$0.440000/M; output=$1.320000/M; cache input=$0.014000/M; off-peak input=$0.220000/M, output=$0.660000/M; comparisonModel=deepseek-v4-pro; fixedInputs=evidenceField=cache; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicitUnavailable — Last verified 2026-08-14, before revalidation date 2026-09-14.FAIL CLOSED — no revalidated observation or provider guarantee
Cache TTL or quota mechanicmodel=deepseek-v4-flash; provider=deepseek; registryRates=dated input=$0.440000/M; output=$1.320000/M; cache input=$0.014000/M; off-peak input=$0.220000/M, output=$0.660000/M; comparisonModel=deepseek-v4-pro; fixedInputs=evidenceField=ttl; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicitUnavailable — Last verified 2026-08-14, before revalidation date 2026-09-14.FAIL CLOSED — no revalidated observation or provider guarantee

First-party provenance: https://api-docs.deepseek.com/quick_start/pricing; registry provider deepseek; verified 2026-08-14; freshness gate=2026-09-14. Revalidate before any current-price claim.

Try DeepSeek V4 Flash pricing analysis →

How fast is DeepSeek V4 Flash?

Tokens / sec
132
TTFT
280 ms
Rank
#10 of 31
$ / M ÷ t/s
$0.0050
Measured with 5 runs on a fixed prompt — see the full methodology.

How much does DeepSeek V4 Flash cost at scale?

Tokens / monthEst. cost (blended 3:1)
100,000$0.07
1,000,000$0.66
10,000,000$6.60
100,000,000$66.00

How does DeepSeek V4 Flash compare with other models?

DeepSeek V4 Pro — $1.98/MGPT-5 Mini — $0.69/MMistral Large 3 — $0.75/MGemini 3.1 Flash Lite — $0.56/M
See all DeepSeek models →

What is DeepSeek V4 Flash best for?

#9 for Summarization#11 for Writing & Content#13 for Structured Data Extraction
Looking for a cheaper option?
Muse Spark 1.3 Contributor is 81.1% cheaper — a config migration. See all 8 alternatives to DeepSeek V4 Flash →

Which DeepSeek V4 Flash head-to-head comparisons are available?

DeepSeek V4 Flash vs DeepSeek V4 Pro

What are common questions about DeepSeek V4 Flash?

Is DeepSeek V4 Flash cheaper than GPT-5 Mini?

DeepSeek V4 Flash costs $0.66/M blended tokens, GPT-5 Mini costs $0.69/M — DeepSeek V4 Flash is cheaper.

How much does 1 million tokens cost with DeepSeek V4 Flash?

At a 3:1 input:output ratio, 1 million blended tokens costs approximately $0.66. Pure input costs $0.44/M; pure output costs $1.32/M.

What does DeepSeek V4 Flash cost at high volume?

At 100 million blended tokens a month, DeepSeek V4 Flash costs approximately $66.00. See the cost-at-scale table below for other volumes.

Try DeepSeek V4 Flash for free

Run real prompts against DeepSeek V4 Flash and every other model on this page in one workspace.

Try DeepSeek V4 Flash Free