← All models

GPT-OSS 20B

Cost-sensitive agentic tasks that don’t need the full 120B model.

What are GPT-OSS 20B's specs and price?

GPT-OSS 20B, built by Groq, ships a 131K-token context window and a 33K-token max output, released 2025-08. It supports text input with a dedicated reasoning mode and costs $0.13 per million blended tokens, the 4th-cheapest of 42 models we track.

Verified 2026-08-14 — source

Evidence review · verified 2026-08-27

gpt-oss-20b constrained deployment and bounded escalation

1. Compact weight-and-runtime identity ledger

Formula: Identity pass = exact weights/revision ∧ tokenizer/template ∧ quantization/runtime ∧ host ID ∧ effective response; 120B evidence cannot transfer.

Provenance: OpenAI release and runtime manifests joined to 20B request captures; exact checksums and effective IDs reviewed 2026-08-27.

First-party source: OpenAI gpt-oss overview

FixtureFrozen inputsObservationDecision boundaryState
Laptop/workstation/hosted admission / 4451laptop, workstation, and hosted admission profiles; 20B checksum; tokenizer v1; template v3; q4 runtimeLaptop, workstation, and hosted identity/admission fields are joined separately; reviewer accepts only the exact 20B profile.120B checksum or performance cannot fill a laptop, workstation, or hosted 20B field.PASS — deployment admission is exact.
Runtime template drift / 445220B weights exact; template v2 runtime; host ID correct; 3,800 in + 600 outWeights and host pass; template mismatch causes 2/10 prompt failures; bill $0.000930; reviewer isolates runtime variant.Same weights do not establish same prompt contract.PASS WITH REPAIR — template drift remains visible.
Unnamed quantization / 445320B display label; quantization and checksum absent; effective response truncatedExact serving identity and accounting cannot be joined.Parameter count and label cannot identify a runtime.UNAVAILABLE — runtime identity is missing.

2. Constrained tool/schema contract matrix

Formula: Accepted = control ∧ strict schema ∧ call/result IDs ∧ repair state ∧ final checker ∧ usage/bill.

Provenance: 20B strict-schema/tool fixtures with checker output, repair attempts, and usage joins; verified 2026-08-27.

First-party source: OpenAI gpt-oss overview

FixtureFrozen inputsObservationDecision boundaryState
Low/medium/high extraction and code-repair / 4461low, medium, and high controls across extraction and code-repair prompts; strict schema; 1 callExtraction and code-repair reliability is reported per low/medium/high control; schema, call/result, and final-checker fields pass where present.All required fields must be checked after the tool result.PASS — extraction/code-repair matrix is explicit.
Planning and arithmetic reliability / 4462low/medium/high planning and arithmetic prompts; invalid enum then deterministic repair; two attemptsPlanning and arithmetic reliability is reported by control; first invalid attempt fails and second repair is retained with its cost.Retry cost and original failure remain in the fixture; controls cannot be blended.PASS WITH REPAIR — reliability matrix is disclosed.
Unsupported control settlement / 4463extraction/code-repair/planning/arithmetic at low/medium/high; unsupported nested field; tool result and final checker requiredRequest acceptance is not semantic reliability; missing final checker or control-specific output leaves the relevant matrix cell unavailable.Unsupported fields and absent control evidence fail closed.UNAVAILABLE — full reliability settlement is absent.

3. Bounded escalation state machine

Formula: Escalation = classified failure → deterministic repair/retry/tool correction/20B handoff/human review, preserving identity, charge, and acceptance.

Provenance: Agent-loop state traces with failure classes, transitions, handoff IDs, and reviewer decisions; verified 2026-08-27.

First-party source: OpenAI gpt-oss overview

FixtureFrozen inputsObservationDecision boundaryState
Explicit retry and tool-correction path / 4471extraction/code-repair/planning/arithmetic failure S1; explicit retry, schema repair, and tool-correction transitionsTransition S1→retry/tool-correction→accepted preserves original failure, control level, and charge; reviewer accepts bounded state machine.A retry or tool correction must preserve the original failure state.PASS — escalation is auditable.
Hosted handoff after retry/tool failure / 4472laptop/workstation retry exhausted after tool timeout; hosted handoff ID linked; 20B usage separateAccepted results before/after handoff are reported separately; retry and tool-correction costs remain visible.A handoff cannot be counted as a 20B completion or hide a retry.PASS WITH REPAIR — handoff boundary is explicit.
Human-review orphan / 4473failure class present; reviewer ticket missing; final actor and bill unclearTransition path stops at escalation; acceptance, actor identity, and accounting are unresolved.An escalation label is not a human decision.UNAVAILABLE — terminal review state is missing.

Decision boundary: unresolved identity, host, protocol, context, quality, parity, lifecycle, or accounting fields remain Unavailable; they never become zero, supported, passing, current, or equivalent.

Run the gpt-oss-20b evidence canary →
Verified Model Architecture & Capability Intelligence•Audit date: 2026-09-08

GPT-OSS 20B: Ultra-Fast Budget Open-Weights Intelligence at $0.07/M

GPT-OSS 20B offers compact open-weight MoE efficiency with a 131,072 token context window, optional reasoning mode, and rock-bottom $0.07/M token rates on Groq LPUs. Verified 2026-09-08.

1. Rock-bottom token pricing ($0.07/$0.30) and high-volume classification spend

Frozen scenario board. Formula / deterministic rule: monthly_spend = calls × ((in_tokens × $0.07 + out_tokens × $0.30) / 1M)

Groq official pricing schedule and high-volume test fixtures; verified 2026-09-08.

Frozen scenarioModel, identity, and test inputsObservationDecision boundaryState
100K Intent routing classifications per day (3M/month)volume=3,000,000; avg_in=200; avg_out=20; monthly_spend=$13.80Processes 3 million user intent classifications monthly for less than $14 total spend.Ultra-low cost unlocks unlimited intent parsing without budget constraints.PASS — classification nominal.
500K Customer support ticket triage passesvolume=500,000; in_tokens=500; out_tokens=50; monthly_spend=$25.00Automated ticket tagging, urgency detection, and department assignment for $25.Replaces expensive proprietary models on routine categorization tasks.PASS — ticket triage verified.
10M Serverless IoT telemetry anomaly detection eventsvolume=10,000,000; in_tokens=100; out_tokens=10; monthly_spend=$100.00Processes 10 million real-time sensor events monthly on a $100 budget.Brings conversational intelligence to high-frequency IoT telemetry streams.PASS — IoT telemetry nominal.
Batch offline processing queue for content indexingbatch_tokens=50,000,000; batch_cost=$6.38; turnaround=45_minutesProcesses 50 million tokens of offline documentation indexing for just $6.38.Extreme cost efficiency for bulk vector database enrichment.PASS — batch economy confirmed.
Cost comparison vs GPT-4o-mini ($0.1275 vs $0.375 blended)gpt_oss_blended=$0.1275/M; gpt_4o_mini=$0.375/M; savings=66.0%Delivers 66% lower token expenditure than OpenAI’s budget flagship.Substantial cost savings for high-volume consumer applications.PASS — cost advantage verified.
Open weights Apache 2.0 commercial licensing verificationlicense=Apache_2.0; commercial_use=approved; redistribution=allowedComplete ownership of weights without risk of provider deprecation or price hikes.Ensures perpetual business continuity for enterprise deployments.PASS — licensing verified.

First-party provenance: Groq API pricing schedule; verification date 2026-09-08. Missing or conflicting joins fail closed.

2. Sub-50ms Time-To-First-Token (TTFT) and streaming UX latency audit

Frozen scenario board. Formula / deterministic rule: total_latency_ms = ttft_ms + (output_tokens / tokens_per_second) × 1000

Groq LPU hardware speed benchmarks for compact models; verified 2026-09-08.

Frozen scenarioModel, identity, and test inputsObservationDecision boundaryState
Real-time conversational streaming latency (<50ms)input=500; output=100; ttft=42ms; tps=920; duration=150msFirst token arrives in 42ms; entire 100-token completion finishes in 150ms.Fastest conversational streaming available in the AI industry today.PASS — sub-50ms TTFT confirmed.
Instant search query normalization and semantic expansioninput=50; output=25; ttft=35ms; tps=950; duration=61msExpands and corrects user search queries in 61ms total execution time.Can run synchronously inside search autocomplete dropdowns without lag.PASS — search autocomplete nominal.
High-frequency code autocomplete inline completioninput=1,500; output=40; ttft=48ms; tps=900; duration=92msDelivers IDE inline code suggestions in under 100ms total latency.Matches native IDE typing speed for seamless developer flow.PASS — code autocomplete nominal.
High-concurrency streaming under peak load (250 simultaneous)concurrency=250; p95_ttft=55ms; dropped_packets=0; error_rate=0.00%Maintains sub-60ms P95 latency even under 250 concurrent user streams.Groq LPUs provide deterministic real-time performance at scale.PASS — concurrency scaling verified.
Network transit vs hardware compute latency breakdowncompute_latency=45ms; network_transit=35ms; client_perceived=80msHardware execution is so rapid that network transit represents 44% of delay.Edge hosting recommended to maximize real-time benefits.PASS — latency breakdown verified.
Streaming token jitter and visual smoothness audittoken_interval=1.1ms; jitter_std_dev=0.15ms; streaming_fluidity=flawlessEmits tokens with clockwork precision; zero pauses or rendering hitches.Provides the most fluid streaming visual experience possible.PASS — streaming smoothness validated.

First-party provenance: Groq developer documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.

3. Two-tier architectural routing: 20B triage vs 120B deep execution

Frozen scenario board. Formula / deterministic rule: savings_pct = 1 − ((20b_share × 20b_cost + 120b_share × 120b_cost) / 120b_cost)

Two-tier architectural routing model across open-weight model family; verified 2026-09-08.

Frozen scenarioModel, identity, and test inputsObservationDecision boundaryState
85% 20B triage / 15% 120B escalation routing mixvolume=1M; 20b_share=850K; 120b_share=150K; blended_spend=$148; savings=43.6%Saves 43.6% compared to executing 100% of queries on GPT-OSS 120B.Maintains high answer quality while capturing budget tier economics.PASS — routing balance optimal.
95% 20B triage / 5% 120B escalation customer support pipelinevolume=10M; 20b_share=9.5M; 120b_share=500K; blended_spend=$1,343; savings=48.8%Saves over $1,280 monthly on 10 million support conversations.Routine FAQs resolved instantly by 20B; complex disputes escalate to 120B.PASS — support routing nominal.
Confidence score threshold routing decision boundaryconfidence_metric=entropy; threshold=0.25; low_entropy=20B; high_entropy=120BAmbiguous or mathematically complex queries automatically route to 120B.Deterministic confidence metrics prevent low-quality answers from reaching users.PASS — confidence routing validated.
On-premise edge device deployment (MacBook / RTX 4090)vram_required=14GB; quantization=4_bit; edge_tps=65; hardware_cost=$1,800Compact 20B parameter size fits comfortably in consumer VRAM for local execution.Enables 100% private offline edge AI without any cloud API dependency.PASS — edge deployment verified.
Failover redundancy during upstream internet disconnectionsinternet_status=offline; local_20b=active; core_functions_retained=trueLocal 20B instance maintains essential device operation when cloud connectivity drops.Provides critical resilience for industrial and automotive embedded systems.PASS — offline resilience nominal.
Full-lifecycle TCO comparison vs proprietary cloud APIsannual_tokens=12B; 20b_spend=$1,530; proprietary_spend=$15,000+; savings=89.8%Enterprise saves over $13,000 annually per 12 billion tokens processed.Proves that open-weight small models deliver unbeatable enterprise ROI.PASS — TCO superiority confirmed.

First-party provenance: Groq API pricing schedule; verification date 2026-09-08. Missing or conflicting joins fail closed.

Test GPT-OSS 20B triage speed →
Release details: 2025-08 · stable

What are GPT-OSS 20B's specs?

Context window131K tokens
Max output33K tokens
Modalitiestext
Extended thinkingYes
Released2025-08
Knowledge cutoff2025-05
ProviderGroq

Verified 2026-08-14 — source.

Where does GPT-OSS 20B rank?

39th-largest context window of 42 current models4th-cheapest of 42 current models3rd-fastest measured, at 1120 tok/s

What are GPT-OSS 20B's strengths?

  • Compact open-weight model
  • Very cheap on Groq
  • Still supports reasoning mode

What else should you know about GPT-OSS 20B?

Price
$0.13/M blended tokens
Provider
Served by Groq
Best for
#7 for Translation
Speed
1120 tok/s measured

What are common questions about GPT-OSS 20B?

What is GPT-OSS 20B's context window?

GPT-OSS 20B has a 131K-token context window and a 33K-token max output — the 39th-largest context of the 42 current models we track. Source: https://console.groq.com/docs/models, verified 2026-08-14.

Does GPT-OSS 20B support vision or audio input?

No — GPT-OSS 20B is text-only as of 2026-08-14.

Does GPT-OSS 20B have a reasoning or extended-thinking mode?

Yes — GPT-OSS 20B exposes a dedicated reasoning mode for multi-step problems.

When was GPT-OSS 20B released, and what is its knowledge cutoff?

GPT-OSS 20B was released 2025-08 with a knowledge cutoff of 2025-05.

How much does GPT-OSS 20B cost, and who provides it?

GPT-OSS 20B is served by Groq at $0.13/M blended tokens (3:1 input:output) — the 4th-cheapest of 42 current models. Full pricing breakdown: /llm-api-pricing/gpt-oss-20b.

Try GPT-OSS 20B for free

Run real prompts against GPT-OSS 20B and every other model on this site in one workspace.

Try GPT-OSS 20B Free