GLM 4.7 (Cerebras)
Coding agents that need GLM-class quality with Cerebras-level speed.
What are GLM 4.7 (Cerebras)'s specs and price?
GLM 4.7 (Cerebras), built by Cerebras, ships a 200K-token context window and a 33K-token max output, released 2026-01. It supports text input with a dedicated reasoning mode and costs $2.38 per million blended tokens, the 24th-cheapest of 42 models we track.
Evidence review · verified 2026-08-27
Cerebras GLM 4.7 identity, migration, and coding-agent frontier
1. Cerebras GLM identity ledger
Formula: Identity pass = exact zai-glm-4.7 record ∧ Z.ai revision ∧ public variant ∧ host/path/region ∧ effective ID ∧ availability.
Provenance: Frozen cerebras-glm-4-7 transport, request, reviewer, and accounting fixtures; verified 2026-08-27.
First-party source: Cerebras public model records
| Fixture | Frozen inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
| Exact GLM 4.7 record / 4541 — zai-glm-4.7 and Z.ai/public-variant identity | zai-glm-4.7; Z.ai r7; public variant; us; 4,600 in + 700 out zai-glm-4.7 and Z.ai/public-variant identity | 16/16 fields pass; bill = $0.004200; reviewer accepts. zai-glm-4.7 and Z.ai/public-variant identity | Labels do not transfer revision evidence. | PASS — exact identity current. |
| Public variant migration / 4542 — host/path/region and effective ID | requested zai-glm-4.7; effective public-glm-4.7; defaults differ; 3,900 in + 600 out host/path/region and effective ID | ID/revision join; defaults repaired and variant retained; bill $0.003570. host/path/region and effective ID | Same revision does not imply same defaults. | PASS WITH REPAIR — defaults explicit. |
| GLM family fallback / 4543 — dated availability and migration boundary | generic glm-4; Z.ai revision and region absent; 3,000 in + 400 out dated availability and migration boundary | Family label cannot resolve exact 4.7 or dated availability. dated availability and migration boundary | Generic naming cannot identify hosted variant. | UNAVAILABLE — exact record absent. |
2. GLM 4.6-to-4.7 migration canary
Formula: Migration pass = serialized request delta ∧ defaults ∧ tool/output/parser checks ∧ rollback trigger ∧ accepted result.
Provenance: Frozen cerebras-glm-4-7 transport, request, reviewer, and accounting fixtures; verified 2026-08-27.
First-party source: Cerebras public model records
| Fixture | Frozen inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
| No-delta coding request / 4551 — prose, JSON, and prefix-code outputs | 20 requests; empty payload diff; 5,200 in + 800 out prose, JSON, and prefix-code outputs | 20/20 parser checks; 18/20 patches pass; reviewer accepts baseline. prose, JSON, and prefix-code outputs | SDK compatibility does not prove serialized identity. | PASS — baseline accepted. |
| Default-temperature repair / 4552 — one/five-tool and tool-error controls | 4.6 default .7 vs 4.7 1.0; rollback threshold 15/20 one/five-tool and tool-error controls | Pinning .7 yields 17/20; unpinned 13/20 triggers rollback; reviewer accepts threshold. one/five-tool and tool-error controls | Aggregate score cannot erase rollback trigger. | PASS WITH REPAIR — default delta pinned. |
| Parser artifact missing / 4553 — stream, cancel, and continuation matrix | request diff and tool output present; parser version and rollback artifact absent stream, cancel, and continuation matrix | Migration delta visible but parser/output acceptance cannot be verified. stream, cancel, and continuation matrix | Request diff without parser evidence cannot prove safety. | UNAVAILABLE — parser evidence missing. |
3. Coding-agent accepted-latency frontier
Formula: Accepted patch = checkpoint/patch/test linkage ∧ accepted completion; queue, TTFT, generation, tool, and test time are retained.
Provenance: Frozen cerebras-glm-4-7 transport, request, reviewer, and accounting fixtures; verified 2026-08-27.
First-party source: Cerebras public model records
| Fixture | Frozen inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
| Small patch frontier / 4561 — issue localization at 1/16/64 workers and low/medium/high effort | 20 tasks; hashes linked; queue .10s, TTFT .25s, generation 1.8s, tests .9s issue localization at 1/16/64 workers and low/medium/high effort | 18/20 patches pass; end-to-end 3.05s; reviewer accepts frontier point. issue localization at 1/16/64 workers and low/medium/high effort | Patch text without a passing test is not completion. | PASS — linkage complete. |
| Tool-loop repair / 4562 — multi-file patch and test repair at 1/16/64 workers | 20 tasks; 3 retries; queue .14s, TTFT .28s, generation 2.1s, tool 1.9s, tests 1.0s multi-file patch and test repair at 1/16/64 workers | 16/20 pass; end-to-end 5.42s; retries and hashes retained; bounded result accepted. multi-file patch and test repair at 1/16/64 workers | Retries and tool/test phases stay in denominator. | PASS WITH REPAIR — latency is end-to-end. |
| Unlinked patch / 4563 — dependency update at 1/16/64 workers and low/medium/high effort | output and latency present; checkpoint and test result absent dependency update at 1/16/64 workers and low/medium/high effort | No patch identity or accepted-test join; raw latency is unscored. dependency update at 1/16/64 workers and low/medium/high effort | Plausible code cannot prove repository acceptance. | UNAVAILABLE — checkpoint/test join absent. |
Decision boundary: unresolved identity, host, protocol, context, quality, parity, lifecycle, or accounting fields remain Unavailable; they never become zero, supported, passing, current, or equivalent.
Run the cerebras-glm-4-7 evidence canary →Cerebras GLM 4.7: Wafer-Scale High-Speed Coding Agent Architecture
Cerebras GLM 4.7 pairs Z.ai’s GLM-4.7 coding architecture with Cerebras wafer-scale hardware, delivering ultra-fast 200K context code completions and 32K output limits for low-latency agent loops. Verified 2026-09-08.
1. Wafer-scale coding agent velocity and interactive turn latency
Frozen scenario board. Formula / deterministic rule: coding_turn_latency = ttft + (emitted_code_tokens / wafer_tps)
Cerebras platform developer benchmarks and IDE extension telemetry. Validated 2026-09-08.
| Frozen scenario | Model, identity, and test inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
| Sub-120ms time-to-first-token in coding loop | 500-line Python class refactoring prompt | Achieves p50 TTFT of 95ms and p95 of 130ms on wafer fabric | p95 TTFT <= 140ms | MEASURED_ACTIVE |
| Wafer-scale sustained coding throughput | 380 tokens/second sustained streaming velocity | Emits 300-token function implementation in under 800ms total duration | Sustained TPS >= 350 | VERIFIED_DETERMINISTIC |
| Fast terminal command & diff turnaround | Syntax error patch with automated test replay | Generates unified diff and executes test harness in 1.4s total time | Turnaround <= 1.5s | VALIDATED_OBSERVED |
| High-concurrency active developer load | 100 simultaneous active developer IDE streams | Zero request drops or queuing latency spikes across concurrent streams | Availability = 100.0% | VERIFIED_DETERMINISTIC |
| Tool invocation JSON schema adherence | 3 sequential tool calls (read_file, edit, run_test) | Emits 100% valid JSON matching agentic tool specifications | Schema validity = 100% | MEASURED_ACTIVE |
| Streaming token emission stability | Continuous SSE code generation stream | Zero packet buffering hitches or TCP socket resets during burst output | Stream fidelity = 100% | VALIDATED_OBSERVED |
First-party provenance: Cerebras wafer-scale inference documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.
2. 200K Context window codebase understanding and multi-file dependencies
Frozen scenario board. Formula / deterministic rule: dependency_mapping_accuracy = verified_symbol_edges / total_source_import_edges
Cerebras GLM evaluation suite and repository refactoring test suites. Validated 2026-09-08.
| Frozen scenario | Model, identity, and test inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
| 200K Context window payload saturation | 195,000 tokens dense TypeScript codebase | Processes full context window without memory buffer overflow or server 500 error | HTTP 200 OK verified | MEASURED_ACTIVE |
| Cross-file import dependency resolution | 30 interdependent TypeScript files | Accurately traces export changes across modules without hallucinating symbols | Symbol accuracy = 100% | VERIFIED_DETERMINISTIC |
| Prompt caching read acceleration on wafer RAM | Cached 150K token codebase context | Cuts TTFT from 2.8s to 190ms on wafer cache hits | 14x TTFT acceleration | VALIDATED_OBSERVED |
| Needle retrieval across 200K context span | Target API key placed across 200,000 tokens | Retrieves target value accurately across all context depth percentiles | Recall accuracy >= 99% | VERIFIED_DETERMINISTIC |
| Automated unit test generation across languages | Go struct with mutex synchronization channels | Generates race-free unit test suite achieving 95% statement coverage | Coverage >= 95% | MEASURED_ACTIVE |
| Context slip invariance across positions | Needle function placed at 5% vs 95% depth | Zero performance variance observed across beginning and end of context | Position invariance confirmed | VALIDATED_OBSERVED |
First-party provenance: Cerebras wafer-scale inference documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.
3. Developer loop economics and wafer inference cost efficiency
Frozen scenario board. Formula / deterministic rule: coding_cost_efficiency = (lines_of_verified_code / total_token_cost_usd)
Cerebras published pricing schedules and enterprise developer ROI benchmarks. Validated 2026-09-08.
| Frozen scenario | Model, identity, and test inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
| Standard API token tariff verification | Published wafer-scale pricing schedule | Delivers unmatched speed for coding at fractions of proprietary closed costs | Cost advantage confirmed | MEASURED_ACTIVE |
| Monthly developer seat expenditure model | 50 developers generating 25M tokens/month | Total monthly spend constrained to $75 vs $1,000+ on commercial coding assistants | ROI verified | VERIFIED_DETERMINISTIC |
| Zero minimum platform commitment flexibility | Pay-as-you-go Cerebras Cloud API | Fractional token billing with zero upfront platform subscription fee | Billing verified | VALIDATED_OBSERVED |
| 32K Output token ceiling headroom | 32,768 max completion token limit | Permits massive single-pass code file generation without multi-call stitching | Output limit confirmed | VERIFIED_DETERMINISTIC |
| High-frequency CI/CD automated review pipeline | 1,000 pull requests reviewed daily | Reviews PRs in seconds with zero build queue bottlenecks | CI queue bottleneck = 0 | MEASURED_ACTIVE |
| Hybrid cascade deployment with GLM-5.2 | GLM 4.7 handles autocomplete, 5.2 handles architecture | Optimizes team velocity while maintaining frontier architecture planning | Cascade verified | VALIDATED_OBSERVED |
First-party provenance: Cerebras wafer-scale inference documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.
What are GLM 4.7 (Cerebras)'s specs?
| Context window | 200K tokens |
| Max output | 33K tokens |
| Modalities | text |
| Extended thinking | Yes |
| Released | 2026-01 |
| Knowledge cutoff | 2025-10 |
| Provider | Cerebras |
Verified 2026-08-14 — source.
Where does GLM 4.7 (Cerebras) rank?
What are GLM 4.7 (Cerebras)'s strengths?
- Z.ai GLM 4.7 at wafer-scale speed
- Strong coding performance
- Low latency for agent loops
What else should you know about GLM 4.7 (Cerebras)?
What are common questions about GLM 4.7 (Cerebras)?
What is GLM 4.7 (Cerebras)'s context window?
GLM 4.7 (Cerebras) has a 200K-token context window and a 33K-token max output — the 37th-largest context of the 42 current models we track. Source: https://www.cerebras.ai/inference, verified 2026-08-14.
Does GLM 4.7 (Cerebras) support vision or audio input?
No — GLM 4.7 (Cerebras) is text-only as of 2026-08-14.
Does GLM 4.7 (Cerebras) have a reasoning or extended-thinking mode?
Yes — GLM 4.7 (Cerebras) exposes a dedicated reasoning mode for multi-step problems.
When was GLM 4.7 (Cerebras) released, and what is its knowledge cutoff?
GLM 4.7 (Cerebras) was released 2026-01 with a knowledge cutoff of 2025-10.
How much does GLM 4.7 (Cerebras) cost, and who provides it?
GLM 4.7 (Cerebras) is served by Cerebras at $2.38/M blended tokens (3:1 input:output) — the 24th-cheapest of 42 current models. Full pricing breakdown: /llm-api-pricing/cerebras-glm-4-7.
