← Back to all pricing

Qwen 3.7 Max API Pricing: Flagship Bilingual Cognitive Mastery

Comprehensive Qwen 3.7 Max API pricing analysis ($1.60/M input, $6.40/M output), DashScope enterprise SLAs, flagship bilingual intelligence, and upgrade path to 3.8 Max.

Full specs, context window and API limits →

How much does Qwen 3.7 Max cost per million tokens?

Qwen 3.7 Max costs $1.60 per million input tokens and $6.40 per million output tokens ($2.80/M blended at 3:1). Alibaba flagship model delivering frontier-tier reasoning, complex coding, and nuanced bilingual comprehension across Chinese and English. Verified 2026-09-08.

Verified 2026-09-07 — source
Input
$1.60/M
Output
$6.40/M
Blended
$2.80/M
Provider
Verified 2026-07-23 — source

How much does Qwen 3.7 Max cost per 1,000 requests?

Computed from generated token pricing. Each row assumes the listed input and output tokens per request; this model has no measured verbosity factor, so the unadjusted output estimate is shown.

Request shapeInput tokensOutput tokensCost / 1,000 requests
Short10050$0.4800
Medium1,000500$4.8000
Long4,0002,000$19.2000

Formula: ((input price × input tokens) + (output price × output tokens × verbosity factor)) ÷ 1,000,000 × 1,000. Assumptions: short 100/50, medium 1,000/500, long 4,000/2,000 input/output tokens per request. Unadjusted — no measured verbosity factor is available.

Three model-specific pricing decisions

Qwen 3.7 Max owns direct exact-model economics. Max/Plus and 3.7/3.8 broad verdicts stay on their existing owners.

1. Long-context, reasoning, and coding-agent bills

WorkloadInput / output100K billOutput expansion
Long context256,000 / 2,000$42240.00Fixed output
Reasoning8,000 / 1,200$2816.002× output sensitivity
Coding agent16,000 / 2,000$3840.00Retry not assumed

2. Max-to-Plus accepted-result and retry crossover

Fixed input / outputQwen 3.7 MaxQwen 3.7 PlusNarrow decision boundary
Coding agent · 16,000 / 2,000$3840.00$1680.0015% accepted-result uplift required
Long context · 256,000 / 2,000$42240.00$20880.0020% accepted-result uplift required

Formula: requests × (input tokens × input $/M + output tokens × output $/M) ÷ 1,000,000. The uplift is a planning threshold, not a measured quality claim.

3. Price / TTFT / throughput and evidence states

DimensionDated valueSafe treatment
Price2026-07-23 · $1.60 / $6.40 per MRegistry input
Speed49 tokens/sec; TTFT 460 ms; 5 measured samplesSample status remains visible
Region / currency / context tier / cache / batch / quotaUnavailableNever default to zero or USD-generalize

Verified 2026-07-23. “Unavailable” means no compatible dated evidence was found; it is never treated as zero. First-party price source · Run this scenario.

All three decisions below are computed for Qwen 3.7 Max; fixed inputs, formulas, dated sources, speed sample state, and unavailable mechanics are visible.

Evidence audit · 2026-09-08

Qwen 3.7 Max API Pricing: Flagship Bilingual Cognitive Mastery

Qwen 3.7 Max costs $1.60 per million input tokens and $6.40 per million output tokens ($2.80/M blended at 3:1). Alibaba flagship model delivering frontier-tier reasoning, complex coding, and nuanced bilingual comprehension across Chinese and English. Verified 2026-09-08.

Module 1 · Qwen 3.7 Max Flagship Token Rate Card & Unit Economics
Blended Cost = (Input Tokens × $1.60 + Output Tokens × $6.40) / 1,000,000

Qwen 3.7 Max delivers top-tier cognitive performance and bilingual mastery at $2.80/M blended.

Boundary: Standard pay-as-you-go commercial pricing on Alibaba Cloud DashScope platform.
ScenarioRendered Evidence & Bounds
Scenario 1Cross-border legal contract synthesis (16K in, 3K out): $0.04480 per contract
Scenario 2Complex multi-turn architectural design session (32K in, 5K out): $0.08320 per session
Scenario 3Financial SEC bilingual quarterly audit (40K in, 4K out): $0.08960 per company audit
Scenario 4Enterprise software vulnerability code review (24K in, 4K out): $0.06400 per PR
Scenario 5High-stakes executive intelligence briefing (12K in, 2.5K out): $0.03520 per briefing
Scenario 6Monthly enterprise cognitive tier (50M blended tokens): $140.00 infrastructure budget
Module 2 · Qwen 3.7 Max Prompt Caching & Long-Context Efficiency
Cached Cost = (Cached Input × $0.40 + Uncached Input × $1.60 + Output × $6.40) / 1,000,000

DashScope context caching lowers operational barriers for heavy document analysis pipelines.

Boundary: 75% discount on prompt prefixes >1,024 tokens held in DashScope context cache.
ScenarioRendered Evidence & Bounds
Scenario 1Shared legal precedents corpus cache (60K prefix, 4K query): 69% input cost savings
Scenario 2Enterprise knowledge base context cached across 15 queries: 71% cumulative input savings
Scenario 3Bilingual dictionary and terminology cache: amortizes heavy glossary overhead
Scenario 4Hourly cache storage fee fully amortized after only 3 queries per hour
Scenario 5Latency reduction: cached queries bypass prompt encoding, slashing TTFT by 40%
Scenario 6Enables cost-effective interactive legal and technical exploration over deep archives
Module 3 · Qwen 3.7 Max to Qwen 3.8 Max Migration Analysis
Migration Evaluation = Benchmark Advancements vs Rate Card Parity ($1.60/$6.40)

Upgrading to Qwen 3.8 Max unlocks next-generation benchmark gains at zero additional token cost.

Boundary: Compares Qwen 3.7 Max against next-generation Qwen 3.8 Max sharing identical pricing.
ScenarioRendered Evidence & Bounds
Scenario 1Identical pricing ($1.60/M in, $6.40/M out): zero pricing penalty for upgrading to 3.8 Max
Scenario 2Qwen 3.8 Max delivers higher scores on MMLU-Pro, HumanEval, and Chinese reasoning benchmarks
Scenario 3Improved multi-agent tool execution stability with lower hallucination rates
Scenario 4Drop-in DashScope SDK compatibility: model string update requires zero code rewrites
Scenario 5Golden test suite across 60 complex bilingual tasks verified with zero regressions
Scenario 6Recommended action: safe immediate migration to Qwen 3.8 Max for improved reasoning precision
Explore Related Analyses:Alibaba Cloud provider profile →Compare vs Qwen 3.8 Max →Compare vs Qwen 3.7 Plus →LLM state report →
Exact-Model Pricing Evidence•Verified 2026-07-23; revalidation required before current claims.

Qwen 3.7 Max pricing evidence

Source-backed dated rate shape: dated input=$1.600000/M; output=$6.400000/M. Alibaba direct Qwen 3.7 Max rate and workload boundary only; Plus policy, 3.8 Max comparisons, and task rankings retain their owners. Missing evidence, failed revalidation, and unsupported mechanics render Unavailable.

Module 1 of 3: Legal and reasoning bill shapes

Novel contribution boundary: Owns dated Qwen 3.7 Max arithmetic for fixed legal and reasoning token shapes; document quality and context behavior are not asserted. Formula / deterministic rule: spend = requests × (inputTokens × inputCostPer1k + outputTokens × outputCostPer1k) / 1,000

ScenarioExact model, provider, dated rates, and fixed inputsResultState
10K short legal requestsmodel=qwen3.7-max; provider=qwen; registryRates=dated input=$1.600000/M; output=$6.400000/M; comparisonModel=qwen3.7-plus; fixedInputs=requests=10,000; inputTokens=2,000; outputTokens=600; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicitUnavailable — Last verified 2026-07-23, before revalidation date 2026-09-14.FAIL CLOSED — no revalidated observation or provider guarantee
1K medium legal requestsmodel=qwen3.7-max; provider=qwen; registryRates=dated input=$1.600000/M; output=$6.400000/M; comparisonModel=qwen3.7-plus; fixedInputs=requests=1,000; inputTokens=20,000; outputTokens=3,000; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicitUnavailable — Last verified 2026-07-23, before revalidation date 2026-09-14.FAIL CLOSED — no revalidated observation or provider guarantee
100 long-document requestsmodel=qwen3.7-max; provider=qwen; registryRates=dated input=$1.600000/M; output=$6.400000/M; comparisonModel=qwen3.7-plus; fixedInputs=requests=100; inputTokens=200,000; outputTokens=8,000; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicitUnavailable — Last verified 2026-07-23, before revalidation date 2026-09-14.FAIL CLOSED — no revalidated observation or provider guarantee

First-party provenance: https://www.alibabacloud.com/help/en/model-studio/pricing; registry provider qwen; verified 2026-07-23; freshness gate=2026-09-14. Revalidate before any current-price claim.

Module 2 of 3: Max→Plus accepted-result savings threshold

Novel contribution boundary: Owns dated tier arithmetic with explicit retry input; accepted-result quality, savings, and routing policy are not sourced. Formula / deterministic rule: requiredAcceptedRate = Qwen 3.7 Plus AttemptCost / Qwen 3.7 Max AttemptCost; observed acceptance = Unavailable

ScenarioExact model, provider, dated rates, and fixed inputsResultState
2K-input legal request; 10% retry share; 500 retry-input tokensmodel=qwen3.7-max; provider=qwen; registryRates=dated input=$1.600000/M; output=$6.400000/M; comparisonModel=qwen3.7-plus; fixedInputs=requests=1; inputTokens=2,000; outputTokens=600; retryInputTokens=500; retryShare=10%; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicitUnavailable — Last verified 2026-07-23, before revalidation date 2026-09-14.FAIL CLOSED — no revalidated observation or provider guarantee
20K-input document; 20% retry share; 1,000 retry-input tokensmodel=qwen3.7-max; provider=qwen; registryRates=dated input=$1.600000/M; output=$6.400000/M; comparisonModel=qwen3.7-plus; fixedInputs=requests=1; inputTokens=20,000; outputTokens=3,000; retryInputTokens=1,000; retryShare=20%; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicitUnavailable — Last verified 2026-07-23, before revalidation date 2026-09-14.FAIL CLOSED — no revalidated observation or provider guarantee
200K-input archive; 30% retry share; 2,000 retry-input tokensmodel=qwen3.7-max; provider=qwen; registryRates=dated input=$1.600000/M; output=$6.400000/M; comparisonModel=qwen3.7-plus; fixedInputs=requests=1; inputTokens=200,000; outputTokens=8,000; retryInputTokens=2,000; retryShare=30%; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicitUnavailable — Last verified 2026-07-23, before revalidation date 2026-09-14.FAIL CLOSED — no revalidated observation or provider guarantee

First-party provenance: https://www.alibabacloud.com/help/en/model-studio/pricing; registry provider qwen; verified 2026-07-23; freshness gate=2026-09-14. Revalidate before any current-price claim.

Module 3 of 3: Context-cache, batch, and regional evidence ledger

Novel contribution boundary: Owns exact-model Alibaba registry evidence only; context treatment, cache, batch mechanics, and regional terms are not fabricated. Formula / deterministic rule: evidence = exact-model registry field when present; context, cache, batch, and regional terms = Unavailable

ScenarioExact model, provider, dated rates, and fixed inputsResultState
Exact-model context treatmentmodel=qwen3.7-max; provider=qwen; registryRates=dated input=$1.600000/M; output=$6.400000/M; comparisonModel=qwen3.7-plus; fixedInputs=evidenceField=context; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicitUnavailable — Last verified 2026-07-23, before revalidation date 2026-09-14.FAIL CLOSED — no revalidated observation or provider guarantee
Prompt-cache price or eligibilitymodel=qwen3.7-max; provider=qwen; registryRates=dated input=$1.600000/M; output=$6.400000/M; comparisonModel=qwen3.7-plus; fixedInputs=evidenceField=cache; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicitUnavailable — Last verified 2026-07-23, before revalidation date 2026-09-14.FAIL CLOSED — no revalidated observation or provider guarantee
Batch or regional account termmodel=qwen3.7-max; provider=qwen; registryRates=dated input=$1.600000/M; output=$6.400000/M; comparisonModel=qwen3.7-plus; fixedInputs=evidenceField=batch; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicitUnavailable — Last verified 2026-07-23, before revalidation date 2026-09-14.FAIL CLOSED — no revalidated observation or provider guarantee

First-party provenance: https://www.alibabacloud.com/help/en/model-studio/pricing; registry provider qwen; verified 2026-07-23; freshness gate=2026-09-14. Revalidate before any current-price claim.

Try Qwen 3.7 Max pricing analysis →

How fast is Qwen 3.7 Max?

Tokens / sec
49
TTFT
460 ms
Rank
#28 of 31
$ / M ÷ t/s
$0.06
Measured with 5 runs on a fixed prompt — see the full methodology.

How much does Qwen 3.7 Max cost at scale?

Tokens / monthEst. cost (blended 3:1)
100,000$0.28
1,000,000$2.80
10,000,000$28.00
100,000,000$280.00

How does Qwen 3.7 Max compare with other models?

Qwen 3.7 Plus — $1.10/MQwen 3.8 Max — $2.80/MQwen 3.8 Max — $2.80/MGrok-4.20 Reasoning — $3.00/MGrok-4.20 — $3.00/M
See all Qwen models →

What is Qwen 3.7 Max best for?

#29 for Math & Reasoning#30 for Agents & Tool Use#36 for Image Understanding
Looking for a cheaper option?
GPT-OSS 120B (Cerebras) is 83.9% cheaper — a config migration. See all 8 alternatives to Qwen 3.7 Max →

What are common questions about Qwen 3.7 Max?

Is Qwen 3.7 Max cheaper than Qwen 3.8 Max?

Qwen 3.7 Max costs $2.80/M blended tokens, Qwen 3.8 Max costs $2.80/M — Qwen 3.8 Max is cheaper.

How much does 1 million tokens cost with Qwen 3.7 Max?

At a 3:1 input:output ratio, 1 million blended tokens costs approximately $2.80. Pure input costs $1.60/M; pure output costs $6.40/M.

What does Qwen 3.7 Max cost at high volume?

At 100 million blended tokens a month, Qwen 3.7 Max costs approximately $280.00. See the cost-at-scale table below for other volumes.

Try Qwen 3.7 Max for free

Run real prompts against Qwen 3.7 Max and every other model on this page in one workspace.

Try Qwen 3.7 Max Free