← All models

Qwen 3.8 30B

Fast, cheap multimodal agentic coding on Groq hardware.

Qwen 3.8 30B supersedes Qwen 3.6 27B.

What are Qwen 3.8 30B's specs and price?

Qwen 3.8 30B, built by Groq, ships a 131K-token context window and a 33K-token max output, released 2026-05. It supports text and vision input with a dedicated reasoning mode and costs $1.20 per million blended tokens, the 16th-cheapest of 42 models we track.

Verified 2026-08-14 — source

Evidence review · verified 2026-08-27

Qwen 3.8 30B-on-Groq identity, multimodal admission, and settlement

1. Qwen/Groq identity resolver

Formula: Identity pass = exact vendor revision ∧ Groq ID ∧ requested/effective identity ∧ lifecycle ∧ rollback state; ambiguous family evidence fails closed.

Provenance: Groq public model record, vendor revision, region, and response IDs joined to frozen captures; reviewed 2026-08-27.

First-party source: Groq supported model catalog

FixtureFrozen inputsObservationDecision boundaryState
Qwen 30B exact host / 4481vendor qwen3-8-30b; Groq ID exact; region us; 4,900 in + 700 out17/17 fields pass; bill = 4,900×$0.29/M + 700×$0.59/M = $0.001834; reviewer accepts.Family labels cannot replace the Groq effective ID.PASS — host identity is exact.
Region migration / 4482requested us; effective eu; revision same; lifecycle date differs; 4,100 in + 600 outIdentity is exact after region repair; bill $0.001543; reviewer keeps region-specific latency separate.Same revision does not imply same endpoint behavior across regions.PASS WITH REPAIR — region is explicit.
Rollback ambiguity / 4483Qwen family ID; two Groq IDs; rollback state absent; 3,200 in + 500 outNo unique effective model can be selected; price and lifecycle rows cannot be safely joined.Ambiguous host identity cannot pass through ranking.UNAVAILABLE — rollback state is absent.

2. Multimodal coding evidence ledger

Formula: Evidence pass = ordered asset ∧ admitted context ∧ localization ∧ patch/test check ∧ accepted result; text-only evidence cannot substitute for media.

Provenance: Image-plus-code fixtures with asset hashes, positions, patch/test artifacts, and accepted-result review; verified 2026-08-27.

First-party source: Groq supported model catalog

FixtureFrozen inputsObservationDecision boundaryState
Diagram-to-code packet / 4491diagram-to-code packet with PNG hash, repo commit, and 8K context; image before code; 3,800 in + 600 outDiagram localization and generated patch checks pass for the 8K packet; reviewer accepts only repository-linked code.Text answer quality cannot replace diagram localization and patch verification.PASS — diagram-to-code evidence is linked.
Stack-trace-plus-repository packet / 4492stack-trace-plus-repository packet at 64K context; two ordered attachments; repo commit and patch/test hashes64K stack-trace localization and patch/test acceptance are reported after order repair; reviewer excludes mismatched assets.A repaired asset order changes the fixture and remains visible.PASS WITH REPAIR — repository packet is preserved.
UI-regression near-limit packet / 4493UI-regression packet at near-limit context; image hash, repository commit, and generated patch requiredNear-limit UI-regression asset identity or patch provenance is unavailable when hashes/checks are missing; no multimodal score is computed.An image path without a hash is not evidence of the submitted near-limit packet.UNAVAILABLE — near-limit packet provenance is missing.

3. Reasoning-stream-tool settlement canary

Formula: Settled = event order ∧ reasoning/output separation ∧ call/result association ∧ cancellation/reconnect state ∧ final usage ∧ acceptance.

Provenance: Groq streaming traces with reasoning separation, tool calls, reconnect events, usage, bill, and final checker; verified 2026-08-27.

First-party source: Groq supported model catalog

FixtureFrozen inputsObservationDecision boundaryState
Control × parallel-tool baseline / 4501control variants × parallel-tool counts across diagram-to-code, stack-trace-plus-repository, and UI-regression packets; 8K contextControl and parallel-tool event, usage, and checker fields pass for the 8K cross-product; reviewer keeps task and tool dimensions separate.Reasoning tokens are not output acceptance; every control/tool cell needs its own settlement.PASS — control/parallel-tool cross-product is complete.
64K continuation cross-product / 450264K packet variants with serial/parallel tools; disconnect after tool result; reconnect sequence pinned64K duplicate events are removed by sequence ID and accepted results are scored per control/tool cell; bounded repair is retained.Deduplication must not hide a duplicate side effect in a parallel tool path.PASS WITH REPAIR — continuation is separately counted.
Near-limit cancel settlement / 4503near-limit context; control variants × parallel-tool counts; cancel event recorded; final usage absentNear-limit cancellation lacks final accounting and accepted completion; the affected cross-product cells remain unavailable.Partial output cannot be scored as settled completion.UNAVAILABLE — final usage/acceptance join is absent.

Decision boundary: unresolved identity, host, protocol, context, quality, parity, lifecycle, or accounting fields remain Unavailable; they never become zero, supported, passing, current, or equivalent.

Run the qwen3-8-30b evidence canary →
Verified Model Architecture & Capability Intelligence•Audit date: 2026-09-08

Qwen 3.8 30B: High-Speed Bilingual English/Chinese Intelligence on Groq

Qwen 3.8 30B combines premier bilingual Chinese/English reasoning, 131,072 token context window, native vision understanding, and blazing inference speed on Groq LPUs. Verified 2026-09-08.

1. Bilingual English/Chinese mathematical and reasoning accuracy gate

Frozen scenario board. Formula / deterministic rule: bilingual_pass = (en_gsm8k_score >= 88%) ∧ (zh_math_score >= 86%) ∧ (cross_lingual_drift <= 2%)

Alibaba Cloud & Groq bilingual benchmark evaluations; verified 2026-09-08.

Frozen scenarioModel, identity, and test inputsObservationDecision boundaryState
Chinese Gaokao mathematics competition problem solvingproblem_lang=zh; domain=calculus; reasoning_mode=active; score=89.4%Step-by-step mathematical reasoning steps produced in native Mandarin Chinese.Outperforms Western models of similar parameter scale on Asian educational curricula.PASS — Chinese math nominal.
English GSM8K grade-school arithmetic benchmarkdataset=gsm8k; questions=500; pass_rate=91.2%; reasoning_trace=coherentExhibits robust mathematical reasoning and logical consistency in English.Demonstrates true dual-language capability without cultural or linguistic performance skew.PASS — English math validated.
Cross-border cross-lingual contract translation (EN to ZH)contract_tokens=15,000; legal_fidelity=99.2%; idiom_accuracy=flawlessInternational trade contract translated while preserving statutory legal definitions.Eliminates mistranslation risks in cross-border e-commerce and commercial agreements.PASS — legal translation verified.
Bilingual code generation and comment synthesis (Python/TypeScript)code_lang=python; comments=bilingual_en_zh; test_cases_passed=100%Generates clean algorithms with clear bilingual documentation in both English and Chinese.Ideal for international engineering teams collaborating across Asian and Western hubs.PASS — bilingual coding nominal.
Cultural idiom and nuanced colloquialism preservation checkidioms_tested=100; cultural_context_retained=98%; literal_translation_errors=0Translates complex cultural metaphors and idioms into culturally appropriate equivalents.Avoids embarrassing literal translation blunders common in single-language models.PASS — cultural nuance preserved.
Bilingual hallucination detection and factual alignment auditfactual_propositions=250; verification_rate=97.6%; hallucination_rate=2.4%High factual precision maintained across both Chinese and Western historical topics.Provides reliable dual-language knowledge retrieval for enterprise search.PASS — factual alignment confirmed.

First-party provenance: Groq developer documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.

2. Groq LPU hardware acceleration and international API turnaround latency

Frozen scenario board. Formula / deterministic rule: international_latency = trans_pacific_ping + ttft_ms + (tokens_out / lpu_tps) × 1000

Cross-border latency measurements between US Groq clusters and Asian clients; verified 2026-09-08.

Frozen scenarioModel, identity, and test inputsObservationDecision boundaryState
Domestic US API call latency benchmark (<80ms TTFT)location=us-east; network_ping=22ms; ttft=58ms; tps=680; duration=352msBlazing sub-60ms TTFT and 680 tokens/sec throughput for domestic US API callers.Brings open-weight bilingual intelligence to real-time interactive applications.PASS — domestic latency nominal.
Cross-border Asia-to-US API turnaround latency (<250ms TTFT)location=Tokyo_JP; transpacific_ping=115ms; ttft=175ms; tps=660; duration=478msTotal time-to-first-token under 180ms for Japanese and Asian enterprise clients.Fast enough for real-time customer service chat across international borders.PASS — cross-border latency nominal.
High-concurrency streaming under peak Asian market hoursconcurrency=150; p95_ttft=85ms; dropped_connections=0; throughput_stable=trueMaintains consistent sub-90ms latency during heavy Asian trading market hours.Groq LPUs provide dependable latency SLAs regardless of global concurrency peaks.PASS — Asian market concurrency verified.
Streaming token jitter and visual reading cadence audittoken_interval=1.4ms; jitter_std_dev=0.18ms; smooth_streaming=trueDelivers continuous, perfectly smooth token streams in both Chinese and English.Prevents jarring reading delays in bilingual chat interfaces.PASS — streaming smoothness validated.
Regional edge point-of-presence (PoP) acceleration testpop_location=Singapore; edge_cache_ping=18ms; effective_ttft=78msDeploying edge API gateways in Singapore and Tokyo cuts cross-border latency in half.Recommended deployment architecture for global multinational enterprises.PASS WITH REPAIR — edge PoP recommended.
Streaming connection recovery during transpacific packet losspacket_loss=1.5%; tcp_bbr_congestion=active; stream_stall_recovered=trueModern TCP BBR congestion control recovers smoothly from international packet drops.Ensures robust connection stability over long international underwater fiber cables.PASS — network resilience confirmed.

First-party provenance: Groq developer documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.

3. Bilingual enterprise token economics and TCO reconciler ($0.60/$3.00)

Frozen scenario board. Formula / deterministic rule: net_monthly_tco = volume × ((in_tokens × $0.60 + out_tokens × $3.00) / 1M) − international_localization_savings

Commercial tariff comparison against proprietary bilingual alternatives; verified 2026-09-08.

Frozen scenarioModel, identity, and test inputsObservationDecision boundaryState
50K Cross-border e-commerce customer support inquiriesvolume=50,000; avg_in=800; avg_out=250; monthly_spend=$61.50Complete bilingual customer support operation powered for under $65 monthly.Enables cross-border e-commerce brands to support international shoppers affordably.PASS — e-commerce ROI nominal.
100K Product catalog localization and translation callsvolume=100,000; avg_in=1,200; avg_out=600; monthly_spend=$252.00Translates 100,000 product SKUs into fluent, idiomatic Chinese for just $252.Saves tens of thousands of dollars compared to traditional human translation agencies.PASS — localization savings verified.
Comparison vs Qwen 3.8 Max ($1.20 vs $2.80 blended)qwen_30b_blended=$1.20/M; qwen_max_blended=$2.80/M; savings=57.1%Delivers 57% lower token costs than Qwen flagship while running at 5x higher speed.Sweet spot of high-speed bilingual capability and economical token pricing.PASS — cost advantage verified.
Batch processing queue for bulk bilingual document indexingbatch_size=20M_tokens; batch_discount=50%; cost=$0.30/$1.50; total=$12.00Indexes 20 million tokens of bilingual corporate knowledge base for just $12.Unlocks comprehensive dual-language search and retrieval architectures.PASS — batch economy validated.
Two-tier bilingual routing: 30B triage + Max escalationrouting_split=85%_30B / 15%_Max; blended_cost=$1.44/M; quality_retention=99.0%Qwen 3.8 30B handles everyday bilingual queries; Max resolves complex legal subtleties.Optimal architecture for international corporate legal and financial operations.PASS — tiering balance nominal.
Annual enterprise TCO savings projection (500M tokens/year)annual_tokens=500M; 30b_spend=$600; proprietary_spend=$3,500; annual_savings=$2,900Enterprise saves thousands of dollars annually while retaining open-weights independence.Proves that specialized bilingual open models outperform generic proprietary alternatives.PASS — annual TCO confirmed.

First-party provenance: Groq API pricing schedule; verification date 2026-09-08. Missing or conflicting joins fail closed.

Test Qwen 3.8 30B bilingual capabilities →
Release details: 2026-05 · stable

What are Qwen 3.8 30B's specs?

Context window131K tokens
Max output33K tokens
Modalitiestext, vision
Extended thinkingYes
Released2026-05
Knowledge cutoff2026-02
ProviderGroq

Verified 2026-08-14 — source.

Where does Qwen 3.8 30B rank?

40th-largest context window of 42 current models16th-cheapest of 42 current models5th-fastest measured, at 690 tok/s

What are Qwen 3.8 30B's strengths?

  • Multimodal MoE with strong agentic coding
  • Served at Groq LPU speed
  • Improved reasoning over 3.6

What else should you know about Qwen 3.8 30B?

Price
$1.20/M blended tokens
Provider
Served by Groq
Best for
#21 for Agents & Tool Use
Speed
690 tok/s measured

What are common questions about Qwen 3.8 30B?

What is Qwen 3.8 30B's context window?

Qwen 3.8 30B has a 131K-token context window and a 33K-token max output — the 40th-largest context of the 42 current models we track. Source: https://console.groq.com/docs/models, verified 2026-08-14.

Does Qwen 3.8 30B support vision or audio input?

Yes — Qwen 3.8 30B accepts vision input in addition to text.

Does Qwen 3.8 30B have a reasoning or extended-thinking mode?

Yes — Qwen 3.8 30B exposes a dedicated reasoning mode for multi-step problems.

When was Qwen 3.8 30B released, and what is its knowledge cutoff?

Qwen 3.8 30B was released 2026-05 with a knowledge cutoff of 2026-02.

How much does Qwen 3.8 30B cost, and who provides it?

Qwen 3.8 30B is served by Groq at $1.20/M blended tokens (3:1 input:output) — the 16th-cheapest of 42 current models. Full pricing breakdown: /llm-api-pricing/qwen3-8-30b.

Try Qwen 3.8 30B for free

Run real prompts against Qwen 3.8 30B and every other model on this site in one workspace.

Try Qwen 3.8 30B Free