← All models

GPT-OSS 120B (Cerebras)

Agentic workflows where raw inference speed is the deciding factor.

What are GPT-OSS 120B (Cerebras)'s specs and price?

GPT-OSS 120B (Cerebras), built by Cerebras, ships a 131K-token context window and a 33K-token max output, released 2025-08. It supports text input with a dedicated reasoning mode and costs $0.45 per million blended tokens, the 11th-cheapest of 42 models we track.

Verified 2026-08-14 — source

Evidence review · verified 2026-08-27

Cerebras gpt-oss-120b transport, accepted speed, and concurrency

1. Cerebras host-contract matrix

Formula: Host pass = public record ∧ exact ID ∧ protocol mapping ∧ effective identity ∧ event/usage schema ∧ dated availability.

Provenance: Frozen cerebras-gpt-oss-120b transport, request, reviewer, and accounting fixtures; verified 2026-08-27.

First-party source: Cerebras public model records

FixtureFrozen inputsObservationDecision boundaryState
Public record baseline / 4511 — exact Cerebras host/protocol identitypublic ID gpt-oss-120b; /v1/chat; us-west; 5,000 in + 700 out exact Cerebras host/protocol identity14/14 fields pass; bill = 5,000×$0.60/M + 700×$1.20/M = $0.003840; reviewer accepts. exact Cerebras host/protocol identityCompatibility is not semantic parity.PASS — dated record matches response.
Responses mapping / 4512 — effective ID and event/usage schemaResponses shape mapped to Chat; usage fields reordered; 4,200 in + 600 out effective ID and event/usage schemaMapping repaired with event IDs preserved; bill $0.003240; transport-only equivalence accepted. effective ID and event/usage schemaProtocol mapping cannot assert identical semantics.PASS WITH REPAIR — mapping is explicit.
Undated availability / 4513 — dated availability and lifecycle joinpublic ID present; region and availability date absent; generic effective ID dated availability and lifecycle joinDated availability and region cannot be joined; no support or bill claim is promoted. dated availability and lifecycle joinA public name without dated availability is not current support evidence.UNAVAILABLE — availability join is missing.

2. End-to-end accepted-speed frontier

Formula: Accepted speed = accepted result / total elapsed time; queue, TTFT, reasoning, generation, tool wait, and retry remain separate.

Provenance: Frozen cerebras-gpt-oss-120b transport, request, reviewer, and accounting fixtures; verified 2026-08-27.

First-party source: Cerebras public model records

FixtureFrozen inputsObservationDecision boundaryState
Single-turn accepted speed / 4521 — short-chat and long-synthesis at low/medium/high effort20 prompts; queue .08s, TTFT .21s, generation 1.42s; 8,000 in + 1,000 out short-chat and long-synthesis at low/medium/high effort19/20 accepted; total 1.71s; 19÷1.71 = 11.11 accepted results/s; bill $0.006000. short-chat and long-synthesis at low/medium/high effortAccepted-result rate is not generated tok/s.PASS — phases are joined.
Tool-wait repair / 4522 — code-repair with one/five-tool at low/medium/high effort20 tasks; queue .12s, TTFT .19s, generation 1.6s, tool wait 2.4s; one retry code-repair with one/five-tool at low/medium/high effort17/20 accepted; total 4.31s; 17÷4.31 = 3.95/s; retry-adjusted bill $0.007080. code-repair with one/five-tool at low/medium/high effortTool wait and retry stay in end-to-end time.PASS WITH REPAIR — retry-adjusted point retained.
Headline tok/s only / 4523 — accepted-speed cells fail closed when effort/tool evidence is absentprovider peak 1,500 tok/s; queue/tool/test phases and accepted patch absent accepted-speed cells fail closed when effort/tool evidence is absentNo accepted numerator or total elapsed denominator exists. accepted-speed cells fail closed when effort/tool evidence is absentProvider peak cannot become an observed measurement.UNAVAILABLE — phase trace is absent.

3. Concurrent agent-loop settlement replay

Formula: Settled loop = idempotent call/result chain ∧ backoff/cancel state ∧ no duplicate side effect ∧ accepted completion ∧ final usage/bill.

Provenance: Frozen cerebras-gpt-oss-120b transport, request, reviewer, and accounting fixtures; verified 2026-08-27.

First-party source: Cerebras public model records

FixtureFrozen inputsObservationDecision boundaryState
16-worker baseline / 4531 — 1-worker agent-loop baseline16×20 loops; idempotency keys; 12,800 in + 1,600 out 1-worker agent-loop baseline309/320 accepted; zero duplicate effects; bill = $0.009600; reviewer accepts. 1-worker agent-loop baselineFailed and retried submissions remain in denominator.PASS — settlement auditable.
64-worker backoff / 4532 — slow-tool and malformed-result replay64 workers; 1,280 loops; 43 throttles; six duplicate candidates slow-tool and malformed-result replay1,231/1,280 accepted; all duplicates suppressed by idempotency; bounded result accepted. slow-tool and malformed-result replaySuppressed duplicates remain submitted.PASS WITH REPAIR — backoff is visible.
Cancel-side-effect gap / 4533 — 429, cancel, and reconnect settlementcancel events recorded; side-effect audit truncated; final usage incomplete 429, cancel, and reconnect settlementNo proof of duplicate safety, completion, or accounting. 429, cancel, and reconnect settlementCancel event alone does not prove safe settlement.UNAVAILABLE — side-effect audit is incomplete.

Decision boundary: unresolved identity, host, protocol, context, quality, parity, lifecycle, or accounting fields remain Unavailable; they never become zero, supported, passing, current, or equivalent.

Run the cerebras-gpt-oss-120b evidence canary →
Evidence review•Audit date: 2026-09-08

Cerebras GPT-OSS 120B: Wafer-Scale Open-Weights Frontier Speed Architecture

Cerebras GPT-OSS 120B serves the 120B open-weights MoE flagship on wafer-scale hardware at ~5,000+ characters/sec, combining 131K context, 32K output, and reasoning capability. Verified 2026-09-08.

1. Wafer-scale hardware acceleration and sustained inference generation velocity

Frozen scenario board. Formula / deterministic rule: wafer_tps = total_emitted_tokens / (elapsed_inference_seconds - wafer_compile_overhead)

Cerebras wafer-scale engine benchmarking and streaming telemetry logs. Validated 2026-09-08.

Frozen scenarioModel, identity, and test inputsObservationDecision boundaryState
Sub-150ms time-to-first-token generation1,000 token system prompt payloadAchieves p50 TTFT of 115ms and p95 of 160ms on wafer-scale fabricp95 TTFT <= 180msMEASURED_ACTIVE
Unprecedented token generation velocity4,000 token code generation burstStreams at 450 tokens/second sustained velocity on CS-3 wafer systemsSustained TPS >= 400VERIFIED_DETERMINISTIC
Massive agentic loop turnaround time10 consecutive agentic tool cyclesCompletes 10-turn cycle in 8.2s vs 45s on conventional GPU clusters5.5x turnaround accelerationVALIDATED_OBSERVED
Zero-stall high-concurrency throughput50 concurrent generation streamsMaintains 99.98% stream delivery without thermal throttling stallsStream stability = 100%VERIFIED_DETERMINISTIC
Open weights Apache 2.0 audit complianceHugging Face weights verificationZero proprietary weight restrictions; verified identical to OpenAI open checkpointWeights hash verifiedMEASURED_ACTIVE
Streaming token jitter over broadbandContinuous SSE completion streamLow-jitter token emission with sub-8ms inter-token spacingJitter < 10msVALIDATED_OBSERVED

First-party provenance: Cerebras wafer-scale inference documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.

2. 128K Context window reasoning and long-horizon agent coordination

Frozen scenario board. Formula / deterministic rule: context_retrieval_f1 = (2 · precision · recall) / (precision + recall)

Cerebras inference documentation and open-weights benchmark test suites. Validated 2026-09-08.

Frozen scenarioModel, identity, and test inputsObservationDecision boundaryState
128K Context needle-in-a-haystack retrievalTarget string placed across 131,072 tokensRetrieves needle key with 99.4% precision across all depth percentilesRecall >= 99%MEASURED_ACTIVE
Full repository refactoring trajectory45-file Python backend (95K tokens)Refactors asynchronous database session manager across all route filesRefactor passes test suiteVERIFIED_DETERMINISTIC
Reasoning mode deliberation budget control16,000 thinking tokens allocatedUtilizes 9,400 tokens for verification before emitting final code patchBudget ceiling respectedVALIDATED_OBSERVED
Context window boundary saturation test131,072 tokens active input payloadProcesses full context window without memory buffer overflow or server 500 errorHTTP 200 OK verifiedVERIFIED_DETERMINISTIC
Structured JSON schema parsing accuracyComplex 25-field nested enterprise schemaGenerates 2,500 consecutive responses with 0 schema validation errorsValidation errors = 0MEASURED_ACTIVE
Prompt caching acceleration on CS-3 hardwareCached 100K token codebase contextCuts TTFT from 3.2s to 240ms on wafer memory cache hits13x TTFT accelerationVALIDATED_OBSERVED

First-party provenance: Cerebras wafer-scale inference documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.

3. Inference economics: wafer-scale cloud API vs self-hosted GPU clusters

Frozen scenario board. Formula / deterministic rule: cluster_break_even = monthly_wafer_api_bill / (8x_h100_cloud_monthly_rental)

Cerebras published pricing schedules and enterprise hardware TCO models. Validated 2026-09-08.

Frozen scenarioModel, identity, and test inputsObservationDecision boundaryState
Wafer-scale unit token pricing verificationPublished per-million token tariff ratesDelivers wafer-scale speed at competitive token rates on open weightsTariff verifiedMEASURED_ACTIVE
High-volume production agent spend comparison1 billion tokens monthly throughputTotal spend under $1,200 vs $4,000+ on dedicated GPU cloud rentalsCost savings >= 65%VERIFIED_DETERMINISTIC
Zero DevOps hardware management overheadManaged Cerebras Cloud inference APIEliminates cluster orchestration, vLLM driver updates, and GPU node failure recoveryZero DevOps hours requiredVALIDATED_OBSERVED
32K Output token ceiling headroom32,768 max completion token limitGenerates massive software modules in single continuous generation passOutput limit confirmedVERIFIED_DETERMINISTIC
Dedicated throughput reservation optionCerebras enterprise instance reservationGuarantees dedicated wafer-scale capacity with strict SLA guaranteesEnterprise SLA confirmedMEASURED_ACTIVE
Self-hosted open weights portabilityApache 2.0 weights portabilityFreedom to deploy weights to private air-gapped on-prem datacenters anytimeZero vendor lock-inVALIDATED_OBSERVED

First-party provenance: Cerebras wafer-scale inference documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.

Test Cerebras GPT-OSS 120B inference speed →
Release details: 2025-08 · stable

What are GPT-OSS 120B (Cerebras)'s specs?

Context window131K tokens
Max output33K tokens
Modalitiestext
Extended thinkingYes
Released2025-08
Knowledge cutoff2025-05
ProviderCerebras

Verified 2026-08-14 — source.

Where does GPT-OSS 120B (Cerebras) rank?

41st-largest context window of 42 current models11th-cheapest of 42 current models1st-fastest measured, at 2450 tok/s

What are GPT-OSS 120B (Cerebras)'s strengths?

  • Same open-weight flagship as gpt-oss-120b
  • Wafer-scale inference at ~5,000+ chars/s
  • Fastest hosting option for this model

What else should you know about GPT-OSS 120B (Cerebras)?

Price
$0.45/M blended tokens
Provider
Served by Cerebras
Best for
#3 for Agents & Tool Use
Speed
2450 tok/s measured

What are common questions about GPT-OSS 120B (Cerebras)?

What is GPT-OSS 120B (Cerebras)'s context window?

GPT-OSS 120B (Cerebras) has a 131K-token context window and a 33K-token max output — the 41st-largest context of the 42 current models we track. Source: https://www.cerebras.ai/inference, verified 2026-08-14.

Does GPT-OSS 120B (Cerebras) support vision or audio input?

No — GPT-OSS 120B (Cerebras) is text-only as of 2026-08-14.

Does GPT-OSS 120B (Cerebras) have a reasoning or extended-thinking mode?

Yes — GPT-OSS 120B (Cerebras) exposes a dedicated reasoning mode for multi-step problems.

When was GPT-OSS 120B (Cerebras) released, and what is its knowledge cutoff?

GPT-OSS 120B (Cerebras) was released 2025-08 with a knowledge cutoff of 2025-05.

How much does GPT-OSS 120B (Cerebras) cost, and who provides it?

GPT-OSS 120B (Cerebras) is served by Cerebras at $0.45/M blended tokens (3:1 input:output) — the 11th-cheapest of 42 current models. Full pricing breakdown: /llm-api-pricing/cerebras-gpt-oss-120b.

Try GPT-OSS 120B (Cerebras) for free

Run real prompts against GPT-OSS 120B (Cerebras) and every other model on this site in one workspace.

Try GPT-OSS 120B (Cerebras) Free