← Back to all pricing

GPT-OSS 20B on Groq API Pricing: Blazing Speed at Budget Rates

Explore GPT-OSS 20B on Groq API pricing ($0.07/M input, $0.30/M output), ultra-high token generation speed on Groq LPUs, and budget cost efficiency.

Full specs, context window and API limits →

How much does GPT-OSS 20B cost per million tokens?

GPT-OSS 20B on Groq costs $0.07 per million input tokens and $0.30 per million output tokens ($0.1275/M blended at 3:1). Delivers lightning-fast inference on Groq LPUs for classification and code generation. Verified 2026-09-08.

Verified 2026-09-07 — source
Input
$0.07/M
Output
$0.30/M
Blended
$0.13/M
Provider
Verified 2026-04-06 — source

How much does GPT-OSS 20B cost per 1,000 requests?

Computed from generated token pricing. Each row assumes the listed input and output tokens per request; output is adjusted by this model's measured 2.93× verbosity factor.

Request shapeInput tokensOutput tokensCost / 1,000 requests
Short10050$0.0515
Medium1,000500$0.5145
Long4,0002,000$2.0580

Formula: ((input price × input tokens) + (output price × output tokens × verbosity factor)) ÷ 1,000,000 × 1,000. Assumptions: short 100/50, medium 1,000/500, long 4,000/2,000 input/output tokens per request. Verbosity run: 2026-06-16T20:31:30.728Z.

Three model-specific pricing decisions

Groq-hosted 20B economics stay separate from open-weight specifications and the 120B page. Every bill below uses text-token units.

1. Fixed workload bills and reasoning-output sensitivity

WorkloadInput / output100K requestsExpansion treatment
Classification300 / 30$3.15Text tokens
Coding agent4,000 / 1,000$60.00Text tokens
Reasoning ×24,000 / 1,000$90.00Output doubled explicitly

2. 20B → 120B cost-per-accepted-run threshold

Fixed input / outputGPT-OSS 20BGPT-OSS 120B (Cerebras)Narrow decision boundary
Coding run · 4,000 / 1,000$60.00$215.0015% accepted-result uplift required
Agent run · 8,000 / 1,200$96.00$370.0020% accepted-result uplift required

Formula: requests × (input tokens × input $/M + output tokens × output $/M) ÷ 1,000,000. The uplift is a planning threshold, not a measured quality claim.

3. Price / TTFT / throughput frontier

Host/modelListed input / output per MSpeed evidenceCapacity boundary
GPT-OSS 20B$0.07 / $0.301120 tokens/sec; TTFT 140 ms; 5 measured samplesRate limit and concurrency unavailable
GPT-OSS 120B (Cerebras)$0.35 / $0.752450 tokens/sec; TTFT 90 ms; 5 measured samplesDo not infer queue capacity

Verified 2026-04-06. “Unavailable” means no compatible dated evidence was found; it is never treated as zero. First-party price source · Run this scenario.

All three decisions below are computed for GPT-OSS 20B; fixed inputs, formulas, dated sources, speed sample state, and unavailable mechanics are visible.

Exact-model pricing guide · verified 2026-09-07

Exact model boundary: Groq gpt-oss-20b (slug gpt-oss-20b). First-party provider pricing and API documentation remain fact owners.

LPU high-throughput token pricing and monthly spend matrix

Frozen scenario board. Formula / deterministic rule: monthly_spend = calls * ((in_tokens * 0.07 + out_tokens * 0.30) / 1M) Boundary: Owns Groq LPU token expenditure calculations for GPT-OSS 20B.

Frozen scenarioModel, identity, provider, and evidence fieldsResultState
50K low-latency conversational repliesmodel=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=50K low-latency conversational replies; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — 50K low-latency conversational replies is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
200K code generation tasksmodel=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=200K code generation tasks; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — 200K code generation tasks is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
1M automated customer triage passesmodel=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=1M automated customer triage passes; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — 1M automated customer triage passes is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
high-concurrency request surgemodel=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=high-concurrency request surge; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — high-concurrency request surge is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch offline ingestion queuemodel=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=batch offline ingestion queue; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — batch offline ingestion queue is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
unresolved billing currencymodel=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=unresolved billing currency; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — unresolved billing currency has no matched, dated bilateral observation.FAIL CLOSED — manual, probe, or source evidence required

First-party provenance: Groq Cloud official pricing; verification date 2026-09-07. Missing or conflicting joins fail closed.

Inference turnaround latency and streaming UX SLA audit

Frozen scenario board. Formula / deterministic rule: turnaround = ttft + (tokens_out / tps); Groq LPU ultra-high generation throughput Boundary: Owns turnaround time benchmarks and user experience responsiveness.

Frozen scenarioModel, identity, provider, and evidence fieldsResultState
real-time conversational streaming (<80ms)model=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=real-time conversational streaming (<80ms); task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — real-time conversational streaming (<80ms) is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
automated code autocomplete (<150ms)model=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=automated code autocomplete (<150ms); task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — automated code autocomplete (<150ms) is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
instant document summary (<300ms)model=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=instant document summary (<300ms); task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — instant document summary (<300ms) is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
long-form synthetic data generationmodel=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=long-form synthetic data generation; task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — long-form synthetic data generation is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
network contention latency buffermodel=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=network contention latency buffer; task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — network contention latency buffer is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
unmeasured speed fixturemodel=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=unmeasured speed fixture; task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — unmeasured speed fixture is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict

First-party provenance: Groq Cloud documentation; verification date 2026-09-07. Missing or conflicting joins fail closed.

Model tier escalation and multi-model routing model

Frozen scenario board. Formula / deterministic rule: blended_cost = 20b_volume * 20b_cost + 120b_volume * 120b_cost Boundary: Owns two-tier architectural routing between fast 20B triage and deep 120B generation.

Frozen scenarioModel, identity, provider, and evidence fieldsResultState
100% GPT-OSS 20B baseline ($0.07/$0.30)model=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=100% GPT-OSS 20B baseline ($0.07/$0.30); escalation mix; total monthly queries; blended spend; cost savings vs pure 120B; accuracy retention; routing recommendation; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — 100% GPT-OSS 20B baseline ($0.07/$0.30) is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
90% 20B triage / 10% 120B escalationmodel=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=90% 20B triage / 10% 120B escalation; escalation mix; total monthly queries; blended spend; cost savings vs pure 120B; accuracy retention; routing recommendation; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — 90% 20B triage / 10% 120B escalation is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
80% 20B triage / 20% 120B escalationmodel=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=80% 20B triage / 20% 120B escalation; escalation mix; total monthly queries; blended spend; cost savings vs pure 120B; accuracy retention; routing recommendation; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — 80% 20B triage / 20% 120B escalation is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
50% 20B triage / 50% 120B escalationmodel=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=50% 20B triage / 50% 120B escalation; escalation mix; total monthly queries; blended spend; cost savings vs pure 120B; accuracy retention; routing recommendation; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — 50% 20B triage / 50% 120B escalation is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
100% direct GPT-OSS 120B executionmodel=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=100% direct GPT-OSS 120B execution; escalation mix; total monthly queries; blended spend; cost savings vs pure 120B; accuracy retention; routing recommendation; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — 100% direct GPT-OSS 120B execution is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
unresolved confidence threshold triggermodel=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=unresolved confidence threshold trigger; escalation mix; total monthly queries; blended spend; cost savings vs pure 120B; accuracy retention; routing recommendation; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — unresolved confidence threshold trigger has no matched, dated bilateral observation.FAIL CLOSED — manual, probe, or source evidence required

First-party provenance: Groq Cloud official pricing; verification date 2026-09-07. Missing or conflicting joins fail closed.

Method and limitations: formulas are deterministic; observed and assumed inputs are labeled; no missing provider, host, account, region, realm, alias, snapshot, revision, weight, artifact, control, tool, modality, workload, rate-period, timestamp, or result is transferred. Run this scenario →

Verified Model Pricing Intelligence•Audit date: 2026-09-08

GPT-OSS 20B on Groq API Pricing: Blazing Speed at Budget Rates

GPT-OSS 20B on Groq costs $0.07 per million input tokens and $0.30 per million output tokens ($0.1275/M blended at 3:1). Delivers lightning-fast inference on Groq LPUs for classification and code generation. Verified 2026-09-08.

LPU high-throughput token pricing and monthly spend matrix

Frozen scenario board. Formula / deterministic rule: monthly_spend = calls * ((in_tokens * 0.07 + out_tokens * 0.30) / 1M) Boundary: Owns Groq LPU token expenditure calculations for GPT-OSS 20B.

Frozen scenarioModel, identity, provider, and evidence fieldsResultState
50K low-latency conversational repliesmodel=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=50K low-latency conversational replies; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — 50K low-latency conversational replies is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
200K code generation tasksmodel=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=200K code generation tasks; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — 200K code generation tasks is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
1M automated customer triage passesmodel=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=1M automated customer triage passes; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — 1M automated customer triage passes is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
high-concurrency request surgemodel=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=high-concurrency request surge; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — high-concurrency request surge is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch offline ingestion queuemodel=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=batch offline ingestion queue; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — batch offline ingestion queue is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
unresolved billing currencymodel=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=unresolved billing currency; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — unresolved billing currency has no matched, dated bilateral observation.FAIL CLOSED — manual, probe, or source evidence required

First-party provenance: Groq official API pricing; verification date 2026-09-08. Missing or conflicting joins fail closed.

Inference turnaround latency and streaming UX SLA audit

Frozen scenario board. Formula / deterministic rule: turnaround = ttft + (tokens_out / tps); Groq LPU ultra-high generation throughput Boundary: Owns turnaround time benchmarks and user experience responsiveness.

Frozen scenarioModel, identity, provider, and evidence fieldsResultState
real-time conversational streaming (<80ms)model=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=real-time conversational streaming (<80ms); task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — real-time conversational streaming (<80ms) is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
automated code autocomplete (<150ms)model=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=automated code autocomplete (<150ms); task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — automated code autocomplete (<150ms) is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
instant document summary (<300ms)model=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=instant document summary (<300ms); task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — instant document summary (<300ms) is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
long-form synthetic data generationmodel=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=long-form synthetic data generation; task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — long-form synthetic data generation is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
network contention latency buffermodel=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=network contention latency buffer; task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — network contention latency buffer is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
unmeasured speed fixturemodel=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=unmeasured speed fixture; task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — unmeasured speed fixture is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict

First-party provenance: Groq developer documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.

Model tier escalation and multi-model routing model

Frozen scenario board. Formula / deterministic rule: blended_cost = 20b_volume * 20b_cost + 120b_volume * 120b_cost Boundary: Owns two-tier architectural routing between fast 20B triage and deep 120B generation.

Frozen scenarioModel, identity, provider, and evidence fieldsResultState
100% GPT-OSS 20B baseline ($0.07/$0.30)model=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=100% GPT-OSS 20B baseline ($0.07/$0.30); escalation mix; total monthly queries; blended spend; cost savings vs pure 120B; accuracy retention; routing recommendation; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — 100% GPT-OSS 20B baseline ($0.07/$0.30) is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
90% 20B triage / 10% 120B escalationmodel=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=90% 20B triage / 10% 120B escalation; escalation mix; total monthly queries; blended spend; cost savings vs pure 120B; accuracy retention; routing recommendation; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — 90% 20B triage / 10% 120B escalation is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
80% 20B triage / 20% 120B escalationmodel=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=80% 20B triage / 20% 120B escalation; escalation mix; total monthly queries; blended spend; cost savings vs pure 120B; accuracy retention; routing recommendation; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — 80% 20B triage / 20% 120B escalation is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
50% 20B triage / 50% 120B escalationmodel=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=50% 20B triage / 50% 120B escalation; escalation mix; total monthly queries; blended spend; cost savings vs pure 120B; accuracy retention; routing recommendation; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — 50% 20B triage / 50% 120B escalation is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
100% direct GPT-OSS 120B executionmodel=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=100% direct GPT-OSS 120B execution; escalation mix; total monthly queries; blended spend; cost savings vs pure 120B; accuracy retention; routing recommendation; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — 100% direct GPT-OSS 120B execution is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
unresolved confidence threshold triggermodel=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=unresolved confidence threshold trigger; escalation mix; total monthly queries; blended spend; cost savings vs pure 120B; accuracy retention; routing recommendation; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — unresolved confidence threshold trigger has no matched, dated bilateral observation.FAIL CLOSED — manual, probe, or source evidence required

First-party provenance: Groq official API pricing; verification date 2026-09-08. Missing or conflicting joins fail closed.

Exact-Model Pricing Evidence•Verified 2026-04-06; revalidation required before current claims.

GPT-OSS 20B on Groq pricing evidence

Source-backed dated rate shape: dated input=$0.075000/M; output=$0.300000/M. Groq-hosted GPT-OSS 20B unit economics only; 120B and Cerebras hosting, edge hardware, and cheapest-model verdicts retain their owners. Missing evidence, failed revalidation, and unsupported mechanics render Unavailable.

Module 1 of 3: Classification and autocomplete bill ladder

Novel contribution boundary: Owns dated Groq-hosted 20B arithmetic for fixed utility shapes; classification quality and latency are not asserted. Formula / deterministic rule: spend = requests × (inputTokens × inputCostPer1k + outputTokens × outputCostPer1k) / 1,000

ScenarioExact model, provider, dated rates, and fixed inputsResultState
100K classification requestsmodel=gpt-oss-20b; provider=groq; registryRates=dated input=$0.075000/M; output=$0.300000/M; comparisonModel=gpt-oss-120b; fixedInputs=requests=100,000; inputTokens=300; outputTokens=80; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicitUnavailable — Last verified 2026-04-06, before revalidation date 2026-09-14.FAIL CLOSED — no revalidated observation or provider guarantee
1M autocomplete requestsmodel=gpt-oss-20b; provider=groq; registryRates=dated input=$0.075000/M; output=$0.300000/M; comparisonModel=gpt-oss-120b; fixedInputs=requests=1,000,000; inputTokens=600; outputTokens=120; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicitUnavailable — Last verified 2026-04-06, before revalidation date 2026-09-14.FAIL CLOSED — no revalidated observation or provider guarantee
10M compact extraction requestsmodel=gpt-oss-20b; provider=groq; registryRates=dated input=$0.075000/M; output=$0.300000/M; comparisonModel=gpt-oss-120b; fixedInputs=requests=10,000,000; inputTokens=1,000; outputTokens=200; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicitUnavailable — Last verified 2026-04-06, before revalidation date 2026-09-14.FAIL CLOSED — no revalidated observation or provider guarantee

First-party provenance: https://console.groq.com/docs/models; registry provider groq; verified 2026-04-06; freshness gate=2026-09-14. Revalidate before any current-price claim.

Module 2 of 3: 20B→120B accepted-result premium

Novel contribution boundary: Owns dated cross-model cost arithmetic with explicit retry input; accepted-result quality and routing policy are not sourced. Formula / deterministic rule: requiredAcceptedRate = 120BAttemptCost / 20BAttemptCost; observed acceptance = Unavailable

ScenarioExact model, provider, dated rates, and fixed inputsResultState
Short classification; 10% retry share; 100 retry-input tokensmodel=gpt-oss-20b; provider=groq; registryRates=dated input=$0.075000/M; output=$0.300000/M; comparisonModel=gpt-oss-120b; fixedInputs=requests=1; inputTokens=300; outputTokens=80; retryInputTokens=100; retryShare=10%; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicitUnavailable — Last verified 2026-04-06, before revalidation date 2026-09-14.FAIL CLOSED — no revalidated observation or provider guarantee
Medium extraction; 20% retry share; 250 retry-input tokensmodel=gpt-oss-20b; provider=groq; registryRates=dated input=$0.075000/M; output=$0.300000/M; comparisonModel=gpt-oss-120b; fixedInputs=requests=1; inputTokens=600; outputTokens=120; retryInputTokens=250; retryShare=20%; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicitUnavailable — Last verified 2026-04-06, before revalidation date 2026-09-14.FAIL CLOSED — no revalidated observation or provider guarantee
Long autocomplete; 30% retry share; 500 retry-input tokensmodel=gpt-oss-20b; provider=groq; registryRates=dated input=$0.075000/M; output=$0.300000/M; comparisonModel=gpt-oss-120b; fixedInputs=requests=1; inputTokens=1,000; outputTokens=200; retryInputTokens=500; retryShare=30%; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicitUnavailable — Last verified 2026-04-06, before revalidation date 2026-09-14.FAIL CLOSED — no revalidated observation or provider guarantee

First-party provenance: https://console.groq.com/docs/models; registry provider groq; verified 2026-04-06; freshness gate=2026-09-14. Revalidate before any current-price claim.

Module 3 of 3: LPU versus local deployment input ledger

Novel contribution boundary: Owns exact-model registry evidence only; VRAM, utilization, energy, local hardware, and quota are not inferred. Formula / deterministic rule: localTCO = hardware + energy + operations; hardware, utilization, and quota inputs = Unavailable

ScenarioExact model, provider, dated rates, and fixed inputsResultState
VRAM or local hardware requirementmodel=gpt-oss-20b; provider=groq; registryRates=dated input=$0.075000/M; output=$0.300000/M; comparisonModel=gpt-oss-120b; fixedInputs=evidenceField=hardware; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicitUnavailable — Last verified 2026-04-06, before revalidation date 2026-09-14.FAIL CLOSED — no revalidated observation or provider guarantee
LPU quota or rate-limit fieldmodel=gpt-oss-20b; provider=groq; registryRates=dated input=$0.075000/M; output=$0.300000/M; comparisonModel=gpt-oss-120b; fixedInputs=evidenceField=quota; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicitUnavailable — Last verified 2026-04-06, before revalidation date 2026-09-14.FAIL CLOSED — no revalidated observation or provider guarantee
Utilization or energy inputmodel=gpt-oss-20b; provider=groq; registryRates=dated input=$0.075000/M; output=$0.300000/M; comparisonModel=gpt-oss-120b; fixedInputs=evidenceField=throughput; FAIL CLOSED — no revalidated observation or provider guarantee; measurement versus assumption=explicitUnavailable — Last verified 2026-04-06, before revalidation date 2026-09-14.FAIL CLOSED — no revalidated observation or provider guarantee

First-party provenance: https://console.groq.com/docs/models; registry provider groq; verified 2026-04-06; freshness gate=2026-09-14. Revalidate before any current-price claim.

Try GPT-OSS 20B on Groq pricing analysis →

How fast is GPT-OSS 20B?

Tokens / sec
1120
TTFT
140 ms
Rank
#3 of 31
$ / M ÷ t/s
$0.0001
Measured with 5 runs on a fixed prompt — see the full methodology.

How much does GPT-OSS 20B cost at scale?

Tokens / monthEst. cost (blended 3:1)
100,000$0.01
1,000,000$0.13
10,000,000$1.31
100,000,000$13.12

How does GPT-OSS 20B compare with other models?

GPT-OSS 120B — $0.26/MLlama 4 Maverick — $0.30/MQwen 3.8 30B — $1.20/MQwen 3.6 27B — $1.20/MMuse Spark 1.3 Contributor — $0.13/MGPT-5 Nano — $0.14/MMinistral 8B — $0.15/M
See all Groq models →

What is GPT-OSS 20B best for?

#7 for Translation#8 for Structured Data Extraction#8 for Writing & Content

What are common questions about GPT-OSS 20B?

Is GPT-OSS 20B cheaper than Muse Spark 1.3 Contributor?

GPT-OSS 20B costs $0.13/M blended tokens, Muse Spark 1.3 Contributor costs $0.13/M — Muse Spark 1.3 Contributor is cheaper.

How much does 1 million tokens cost with GPT-OSS 20B?

At a 3:1 input:output ratio, 1 million blended tokens costs approximately $0.13. Pure input costs $0.07/M; pure output costs $0.30/M.

What does GPT-OSS 20B cost at high volume?

At 100 million blended tokens a month, GPT-OSS 20B costs approximately $13.12. See the cost-at-scale table below for other volumes.

Try GPT-OSS 20B for free

Run real prompts against GPT-OSS 20B and every other model on this page in one workspace.

Try GPT-OSS 20B Free