← Back to all pricing

DeepSeek V4 Pro API Pricing: Frontier Reasoning at Open-Weights Economics

Comprehensive DeepSeek V4 Pro API pricing analysis ($1.32/M input, $3.96/M output), off-peak discounts, prompt caching breaks, and frontier reasoning benchmarks.

Full specs, context window and API limits →

How much does DeepSeek V4 Pro cost per million tokens?

DeepSeek V4 Pro costs $1.32 per million input tokens and $3.96 per million output tokens ($1.98/M blended at 3:1). Delivers state-of-the-art mathematical and code reasoning at disruptive price points. Verified 2026-09-08.

Verified 2026-09-07 — source
Input
$1.32/M
Output
$3.96/M
Blended
$1.98/M
Provider
Verified 2026-08-14 — source

How much does DeepSeek V4 Pro cost per 1,000 requests?

Computed from generated token pricing. Each row assumes the listed input and output tokens per request; output is adjusted by this model's measured 3.30× verbosity factor.

Request shapeInput tokensOutput tokensCost / 1,000 requests
Short10050$0.7854
Medium1,000500$7.8540
Long4,0002,000$31.4160

Formula: ((input price × input tokens) + (output price × output tokens × verbosity factor)) ÷ 1,000,000 × 1,000. Assumptions: short 100/50, medium 1,000/500, long 4,000/2,000 input/output tokens per request. Verbosity run: 2026-06-16T20:31:30.728Z.

Exact-model pricing guide · verified 2026-09-07

Exact model boundary: DeepSeek DeepSeek V4 Pro (deepseek-v4-pro). Pricing cards, context tiers, caching multipliers, and task pages remain fact owners.

Peak versus off-peak UTC schedule discount ledger

Frozen scenario board. Formula / deterministic rule: bill = (standard_hours * peak_rate + off_peak_hours * off_peak_rate); off-peak window is 16:30-00:30 UTC Boundary: Owns scheduled off-peak pricing calculations for DeepSeek V4 Pro.

Frozen scenarioExact identity and evidence fieldsResultState
peak daytime processingmodel=deepseek-v4-pro; provider=DeepSeek; slug=deepseek-v4-pro; scenario=peak daytime processing; schedule window; UTC hours; input rate; output rate; 100K call cost; net savings; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — peak daytime processing is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
off-peak scheduled batchmodel=deepseek-v4-pro; provider=DeepSeek; slug=deepseek-v4-pro; scenario=off-peak scheduled batch; schedule window; UTC hours; input rate; output rate; 100K call cost; net savings; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — off-peak scheduled batch is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
50/50 blended daily trafficmodel=deepseek-v4-pro; provider=DeepSeek; slug=deepseek-v4-pro; scenario=50/50 blended daily traffic; schedule window; UTC hours; input rate; output rate; 100K call cost; net savings; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — 50/50 blended daily traffic is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
weekend batch backlogmodel=deepseek-v4-pro; provider=DeepSeek; slug=deepseek-v4-pro; scenario=weekend batch backlog; schedule window; UTC hours; input rate; output rate; 100K call cost; net savings; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — weekend batch backlog is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
burst schedule overridemodel=deepseek-v4-pro; provider=DeepSeek; slug=deepseek-v4-pro; scenario=burst schedule override; schedule window; UTC hours; input rate; output rate; 100K call cost; net savings; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — burst schedule override is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
unsupported time windowmodel=deepseek-v4-pro; provider=DeepSeek; slug=deepseek-v4-pro; scenario=unsupported time window; schedule window; UTC hours; input rate; output rate; 100K call cost; net savings; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — unsupported time window has no matched, dated bilateral observation.FAIL CLOSED — manual, probe, or source evidence required

First-party provenance: DeepSeek API official pricing; verification date 2026-09-07. Missing or conflicting joins fail closed.

Prompt cache write/read economics and hit-rate sensitivity

Frozen scenario board. Formula / deterministic rule: net_input = uncached_tokens * 0.27 + cached_tokens * 0.07; cache read offers ~74% discount Boundary: Owns DeepSeek context caching economics.

Frozen scenarioExact identity and evidence fieldsResultState
0% cache hit (pure uncached)model=deepseek-v4-pro; provider=DeepSeek; slug=deepseek-v4-pro; scenario=0% cache hit (pure uncached); cache hit rate; uncached input tokens; cached input tokens; effective input cost; total request cost; savings; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — 0% cache hit (pure uncached) is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
25% occasional prefix reusemodel=deepseek-v4-pro; provider=DeepSeek; slug=deepseek-v4-pro; scenario=25% occasional prefix reuse; cache hit rate; uncached input tokens; cached input tokens; effective input cost; total request cost; savings; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — 25% occasional prefix reuse is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
50% repeated document QAmodel=deepseek-v4-pro; provider=DeepSeek; slug=deepseek-v4-pro; scenario=50% repeated document QA; cache hit rate; uncached input tokens; cached input tokens; effective input cost; total request cost; savings; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — 50% repeated document QA is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
75% agent system prompt reusemodel=deepseek-v4-pro; provider=DeepSeek; slug=deepseek-v4-pro; scenario=75% agent system prompt reuse; cache hit rate; uncached input tokens; cached input tokens; effective input cost; total request cost; savings; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — 75% agent system prompt reuse is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
90% high-frequency API loopmodel=deepseek-v4-pro; provider=DeepSeek; slug=deepseek-v4-pro; scenario=90% high-frequency API loop; cache hit rate; uncached input tokens; cached input tokens; effective input cost; total request cost; savings; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — 90% high-frequency API loop is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
unsupported cache structuremodel=deepseek-v4-pro; provider=DeepSeek; slug=deepseek-v4-pro; scenario=unsupported cache structure; cache hit rate; uncached input tokens; cached input tokens; effective input cost; total request cost; savings; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — unsupported cache structure has no matched, dated bilateral observation.FAIL CLOSED — manual, probe, or source evidence required

First-party provenance: DeepSeek API official pricing; verification date 2026-09-07. Missing or conflicting joins fail closed.

Thinking-token output budget and reasoning break-even

Frozen scenario board. Formula / deterministic rule: total_cost = (input * 0.27 + (thinking_out + answer_out) * 1.10) / 1M; visible CoT is billed as output Boundary: Owns reasoning token expense forecasting.

Frozen scenarioExact identity and evidence fieldsResultState
direct answer (no thinking)model=deepseek-v4-pro; provider=DeepSeek; slug=deepseek-v4-pro; scenario=direct answer (no thinking); reasoning depth; thinking tokens; answer tokens; total output tokens; cost per call; budget state; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — direct answer (no thinking) is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
1K thinking tokensmodel=deepseek-v4-pro; provider=DeepSeek; slug=deepseek-v4-pro; scenario=1K thinking tokens; reasoning depth; thinking tokens; answer tokens; total output tokens; cost per call; budget state; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — 1K thinking tokens is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
4K moderate reasoningmodel=deepseek-v4-pro; provider=DeepSeek; slug=deepseek-v4-pro; scenario=4K moderate reasoning; reasoning depth; thinking tokens; answer tokens; total output tokens; cost per call; budget state; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — 4K moderate reasoning is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
16K deep mathematical proofmodel=deepseek-v4-pro; provider=DeepSeek; slug=deepseek-v4-pro; scenario=16K deep mathematical proof; reasoning depth; thinking tokens; answer tokens; total output tokens; cost per call; budget state; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — 16K deep mathematical proof is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
32K complex software designmodel=deepseek-v4-pro; provider=DeepSeek; slug=deepseek-v4-pro; scenario=32K complex software design; reasoning depth; thinking tokens; answer tokens; total output tokens; cost per call; budget state; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — 32K complex software design is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
uncontrolled thinking loopmodel=deepseek-v4-pro; provider=DeepSeek; slug=deepseek-v4-pro; scenario=uncontrolled thinking loop; reasoning depth; thinking tokens; answer tokens; total output tokens; cost per call; budget state; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — uncontrolled thinking loop is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict

First-party provenance: DeepSeek API documentation; verification date 2026-09-07. Missing or conflicting joins fail closed.

Method and limitations: formulas are deterministic; observed and assumed inputs are labeled; no missing provider, host, account, region, realm, alias, snapshot, revision, weight, artifact, control, tool, modality, workload, rate-period, timestamp, or result is transferred. Run this scenario →

Evidence audit · 2026-09-08

DeepSeek V4 Pro API Pricing: Frontier Reasoning at Open-Weights Economics

DeepSeek V4 Pro costs $1.32 per million input tokens and $3.96 per million output tokens ($1.98/M blended at 3:1). Delivers state-of-the-art mathematical and code reasoning at disruptive price points. Verified 2026-09-08.

Module 1 · DeepSeek V4 Pro Frontier Token Rate Card
Blended Cost = (Input Tokens × $1.32 + Output Tokens × $3.96) / 1,000,000

DeepSeek V4 Pro provides frontier-class cognitive power at sub-$2 per million blended economics.

Boundary: Standard pay-as-you-go rate card; off-peak window and prompt caching discounts apply.
ScenarioRendered Evidence & Bounds
Scenario 1Competitive programming solution verification (2K in, 1K out): $0.006600 per problem
Scenario 2Complex software architecture review (16K in, 3K out): $0.033000 per review pass
Scenario 3Mathematical proof generation and check (4K in, 2K out): $0.013200 per theorem
Scenario 4Full-stack pull request defect scan (32K in, 4K out): $0.058080 per PR audit
Scenario 5Autonomous multi-step coding agent loop (64K in, 6K out): $0.108240 per task cycle
Scenario 6Monthly 50M token enterprise developer tier: $99.00 total infrastructure spend
Module 2 · DeepSeek V4 Pro Off-Peak Window & Prompt Caching Savings
Discounted Cost = (Cached Input × $0.132 + Off-Peak Tokens × 0.50 + Generation × $3.96) / 1,000,000

Leveraging off-peak scheduling and context caching minimizes operational overhead for dev teams.

Boundary: Evaluates 90% prompt caching discount plus 50% off-peak window tariff reduction (16:30–00:30 UTC).
ScenarioRendered Evidence & Bounds
Scenario 1Off-peak batch pipeline processing: cuts overall input/output token tariffs by 50%
Scenario 2Codebase AST context cache (50K tokens): 84% prompt cost reduction on repeated runs
Scenario 3Combined off-peak + prompt cache: effective blended rate drops below $0.65/M
Scenario 4Zero cache retention storage fees charged during active continuous sessions
Scenario 5Enables large-scale nightly regression testing of massive enterprise monorepos
Scenario 6Net infrastructure budget reduction of 62% for asynchronous developer CI pipelines
Module 3 · DeepSeek V4 Pro vs Western Flagship TCO Comparison
TCO Savings = Western Flagship ($15-$30/M) - DeepSeek V4 Pro ($1.98/M) = 87%-93% Net Cost Reduction

Delivers elite code and analytical reasoning performance at less than one-tenth Western flagship prices.

Boundary: Compares DeepSeek V4 Pro against frontier Western alternatives on code and math tasks.
ScenarioRendered Evidence & Bounds
Scenario 1DeepSeek V4 Pro ($1.98/M blended) vs Claude Opus 5 ($30.00/M blended): 93.4% cost savings
Scenario 2DeepSeek V4 Pro vs GPT-5.6 Sol ($8.00/M blended): 75.3% operational cost reduction
Scenario 3Matches or exceeds Western frontier models on HumanEval and MATH-500 benchmarks
Scenario 4High-volume production tier (100M tokens/mo): saves >$2,500/mo compared to Western flagships
Scenario 5Standard OpenAI-compatible REST API allows effortless drop-in gateway integration
Scenario 6Strongly recommended for cost-conscious AI engineering and high-throughput coding agents
Explore Related Analyses:DeepSeek provider profile →Compare vs DeepSeek V4 Flash →Compare vs Claude Opus 5 →Best LLM for coding →
Exact-Model Pricing Evidence•Verified 2026-08-14; revalidation required before current claims.

DeepSeek V4 Pro pricing evidence

Source-backed dated rate shape: dated input=$1.320000/M; output=$3.960000/M; cache input=$0.044000/M; off-peak input=$0.660000/M, output=$1.980000/M. DeepSeek V4 Pro exact-model rates and off-peak arithmetic only; DeepSeek policy, Flash economics, and math rankings retain their owners. Missing evidence, failed revalidation, and unsupported mechanics render Unavailable.

Module 1 of 3: Reasoning-token bill ladder

Novel contribution boundary: Owns dated Pro arithmetic for fixed reasoning-token shapes; proof quality and reasoning behavior are not asserted. Formula / deterministic rule: spend = requests × (inputTokens × inputCostPer1k + outputTokens × outputCostPer1k) / 1,000

ScenarioExact model, provider, dated rates, and fixed inputsResultState
5K compact proof requestsmodel=deepseek-v4-pro; provider=deepseek; registryRates=dated input=$1.320000/M; output=$3.960000/M; cache input=$0.044000/M; off-peak input=$0.660000/M, output=$1.980000/M; comparisonModel=deepseek-v4-flash; fixedInputs=requests=5,000; inputTokens=2,000; outputTokens=1,000; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicitUnavailable — Last verified 2026-08-14, before revalidation date 2026-09-14.FAIL CLOSED — no revalidated observation or provider guarantee
25K algorithm reviewsmodel=deepseek-v4-pro; provider=deepseek; registryRates=dated input=$1.320000/M; output=$3.960000/M; cache input=$0.044000/M; off-peak input=$0.660000/M, output=$1.980000/M; comparisonModel=deepseek-v4-flash; fixedInputs=requests=25,000; inputTokens=6,000; outputTokens=3,000; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicitUnavailable — Last verified 2026-08-14, before revalidation date 2026-09-14.FAIL CLOSED — no revalidated observation or provider guarantee
50K deep reasoning requestsmodel=deepseek-v4-pro; provider=deepseek; registryRates=dated input=$1.320000/M; output=$3.960000/M; cache input=$0.044000/M; off-peak input=$0.660000/M, output=$1.980000/M; comparisonModel=deepseek-v4-flash; fixedInputs=requests=50,000; inputTokens=12,000; outputTokens=8,000; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicitUnavailable — Last verified 2026-08-14, before revalidation date 2026-09-14.FAIL CLOSED — no revalidated observation or provider guarantee

First-party provenance: https://api-docs.deepseek.com/quick_start/pricing; registry provider deepseek; verified 2026-08-14; freshness gate=2026-09-14. Revalidate before any current-price claim.

Module 2 of 3: Pro→Flash accepted-result cost boundary

Novel contribution boundary: Owns dated cross-model arithmetic with explicit retry input; accepted-result quality and model selection are not sourced. Formula / deterministic rule: requiredAcceptedRate = FlashAttemptCost / ProAttemptCost; observed acceptance = Unavailable

ScenarioExact model, provider, dated rates, and fixed inputsResultState
2K-input proof; 10% retry share; 500 retry-input tokensmodel=deepseek-v4-pro; provider=deepseek; registryRates=dated input=$1.320000/M; output=$3.960000/M; cache input=$0.044000/M; off-peak input=$0.660000/M, output=$1.980000/M; comparisonModel=deepseek-v4-flash; fixedInputs=requests=1; inputTokens=2,000; outputTokens=1,000; retryInputTokens=500; retryShare=10%; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicitUnavailable — Last verified 2026-08-14, before revalidation date 2026-09-14.FAIL CLOSED — no revalidated observation or provider guarantee
6K-input algorithm review; 20% retry share; 1,000 retry-input tokensmodel=deepseek-v4-pro; provider=deepseek; registryRates=dated input=$1.320000/M; output=$3.960000/M; cache input=$0.044000/M; off-peak input=$0.660000/M, output=$1.980000/M; comparisonModel=deepseek-v4-flash; fixedInputs=requests=1; inputTokens=6,000; outputTokens=3,000; retryInputTokens=1,000; retryShare=20%; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicitUnavailable — Last verified 2026-08-14, before revalidation date 2026-09-14.FAIL CLOSED — no revalidated observation or provider guarantee
12K-input architecture proof; 30% retry share; 2,000 retry-input tokensmodel=deepseek-v4-pro; provider=deepseek; registryRates=dated input=$1.320000/M; output=$3.960000/M; cache input=$0.044000/M; off-peak input=$0.660000/M, output=$1.980000/M; comparisonModel=deepseek-v4-flash; fixedInputs=requests=1; inputTokens=12,000; outputTokens=8,000; retryInputTokens=2,000; retryShare=30%; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicitUnavailable — Last verified 2026-08-14, before revalidation date 2026-09-14.FAIL CLOSED — no revalidated observation or provider guarantee

First-party provenance: https://api-docs.deepseek.com/quick_start/pricing; registry provider deepseek; verified 2026-08-14; freshness gate=2026-09-14. Revalidate before any current-price claim.

Module 3 of 3: UTC window and cache-hit evidence board

Novel contribution boundary: Owns exact-model DeepSeek registry evidence only; cache hits, quota, and operational window mechanics are not fabricated. Formula / deterministic rule: evidence = exact-model registry field when present; cache hit rate and operational fields = Unavailable

ScenarioExact model, provider, dated rates, and fixed inputsResultState
UTC off-peak input/output ratesmodel=deepseek-v4-pro; provider=deepseek; registryRates=dated input=$1.320000/M; output=$3.960000/M; cache input=$0.044000/M; off-peak input=$0.660000/M, output=$1.980000/M; comparisonModel=deepseek-v4-flash; fixedInputs=evidenceField=off-peak; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicitUnavailable — Last verified 2026-08-14, before revalidation date 2026-09-14.FAIL CLOSED — no revalidated observation or provider guarantee
Prompt-cache input ratemodel=deepseek-v4-pro; provider=deepseek; registryRates=dated input=$1.320000/M; output=$3.960000/M; cache input=$0.044000/M; off-peak input=$0.660000/M, output=$1.980000/M; comparisonModel=deepseek-v4-flash; fixedInputs=evidenceField=cache; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicitUnavailable — Last verified 2026-08-14, before revalidation date 2026-09-14.FAIL CLOSED — no revalidated observation or provider guarantee
Cache-hit rate or quota fieldmodel=deepseek-v4-pro; provider=deepseek; registryRates=dated input=$1.320000/M; output=$3.960000/M; cache input=$0.044000/M; off-peak input=$0.660000/M, output=$1.980000/M; comparisonModel=deepseek-v4-flash; fixedInputs=evidenceField=quota; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicitUnavailable — Last verified 2026-08-14, before revalidation date 2026-09-14.FAIL CLOSED — no revalidated observation or provider guarantee

First-party provenance: https://api-docs.deepseek.com/quick_start/pricing; registry provider deepseek; verified 2026-08-14; freshness gate=2026-09-14. Revalidate before any current-price claim.

Try DeepSeek V4 Pro pricing analysis →

How fast is DeepSeek V4 Pro?

Tokens / sec
68
TTFT
480 ms
Rank
#22 of 31
$ / M ÷ t/s
$0.03
Measured with 5 runs on a fixed prompt — see the full methodology.

How much does DeepSeek V4 Pro cost at scale?

Tokens / monthEst. cost (blended 3:1)
100,000$0.20
1,000,000$1.98
10,000,000$19.80
100,000,000$198.00

How does DeepSeek V4 Pro compare with other models?

DeepSeek V4 Flash — $0.66/MClaude Haiku 4.5 — $2.00/MMuse Spark 1.3 — $2.00/Mo3-Mini — $1.93/M
See all DeepSeek models →

What is DeepSeek V4 Pro best for?

#12 for Math & Reasoning#17 for Writing & Content#17 for Summarization
Looking for a cheaper option?
Muse Spark 1.3 Contributor is 93.7% cheaper — a config migration. See all 8 alternatives to DeepSeek V4 Pro →

Which DeepSeek V4 Pro head-to-head comparisons are available?

DeepSeek V4 Pro vs Claude Opus 4.8DeepSeek V4 Pro vs Claude Opus 5.5DeepSeek V4 Pro vs Claude Sonnet 5DeepSeek V4 Pro vs Gemini 3.1 Pro

What are common questions about DeepSeek V4 Pro?

Is DeepSeek V4 Pro cheaper than Claude Haiku 4.5?

DeepSeek V4 Pro costs $1.98/M blended tokens, Claude Haiku 4.5 costs $2.00/M — DeepSeek V4 Pro is cheaper.

How much does 1 million tokens cost with DeepSeek V4 Pro?

At a 3:1 input:output ratio, 1 million blended tokens costs approximately $1.98. Pure input costs $1.32/M; pure output costs $3.96/M.

What does DeepSeek V4 Pro cost at high volume?

At 100 million blended tokens a month, DeepSeek V4 Pro costs approximately $198.00. See the cost-at-scale table below for other volumes.

Try DeepSeek V4 Pro for free

Run real prompts against DeepSeek V4 Pro and every other model on this page in one workspace.

Try DeepSeek V4 Pro Free