← All models

GLM-5.2

Long-horizon coding on open weights at a fraction of frontier pricing.

GLM-5.2 supersedes GLM-5.1.

What are GLM-5.2's specs and price?

GLM-5.2, built by Z.ai, ships a 1M-token context window and a 64K-token max output, released 2026-05. It supports text input with a dedicated reasoning mode and costs $2.15 per million blended tokens, the 23rd-cheapest of 42 models we track.

Verified 2026-08-14 — source

Evidence review · verified 2026-08-14

GLM-5.2 exact identity, request envelope, and runtime evidence

1. GLM-5.2 artifact-and-surface identity ledger

Formula / rubric: artifact pass = BF16/W8A8/W4A8C8 artifact ∧ hosted surface ∧ revision join.

Dated provenance: Frozen models-glm-5-2 fixture; GLM-5.2 hosted artifact chain; reviewer ledger verified 2026-08-14.

First-party citation: Z.ai GLM-5.2 model guide

FixtureInputsObservation / calculationDecision boundaryState
BF16 hosted artifact chainBF16 artifact; repository commit; hosted endpoint; revision; request IDBF16 artifact, commit, and hosted endpoint form one attributable chain.A repository tag alone does not prove hosted weights.PASS — chain joined.
W8A8 hosted artifact chainW8A8 artifact; quantization manifest; host; revision; tokenizerW8A8 is retained as a distinct hosted artifact with its own manifest.Quantized artifacts cannot inherit BF16 benchmark evidence.PASS — chain joined.
W4A8C8 hosted artifact chainW4A8C8 artifact; calibration file; host; revision; checksumThe hosted artifact is named, but the calibration checksum is missing.Missing calibration identity blocks artifact attribution.UNAVAILABLE — chain incomplete.

2. Reasoning-control and tool-order matrix

Formula / rubric: order pass = reasoning control state ∧ tool-call order ∧ result settlement.

Dated provenance: Frozen models-glm-5-2 fixture; GLM-5.2 reasoning and tool-order fixtures; reviewer ledger verified 2026-08-14.

First-party citation: Z.ai GLM-5.2 model guide

FixtureInputsObservation / calculationDecision boundaryState
reasoning then toolreasoning enabled; tool call 1; tool result 1; final answer; event orderReasoning control precedes the tool call and the ordered result settles.Final text quality cannot repair event-order loss.PASS — order preserved.
tool then reasoning retrytool call; retry; reasoning control; call IDs; usage; cancellationRetry is recorded after the tool result; reasoning order is explicit.Retry order must remain attributable to the original tool call.PASS WITH REPAIR — retry linked.
reasoning/tool conflictreasoning flag; parallel tools; returned events; missing stop reasonTool results arrive, but the missing stop reason prevents settlement.Tool completion without stop state is not accepted.UNAVAILABLE — stop join missing.

3. Long-horizon coding checkpoint ledger

Formula / rubric: checkpoint = 100K/500K/850K/near-1M inventory item with artifact, tool, test, and billing joins.

Dated provenance: Frozen models-glm-5-2 fixture; GLM-5.2 long-horizon checkpoint inventory; reviewer ledger verified 2026-08-14.

First-party citation: vLLM Ascend GLM-5.2 deployment guide

FixtureInputsObservation / calculationDecision boundaryState
100K checkpoint100K repository tokens; checkpoint C100; tool calls; tests; usageCheckpoint C100 restores and all tests settle with a billing join.Short checkpoint success does not establish long-horizon continuity.PASS — checkpoint settled.
500K/850K checkpoints500K checkpoint C500; 850K checkpoint C850; tool order; test artifactsC500 passes; C850 has a missing test artifact and remains unavailable.Checkpoint inventory is per size, not interpolated.UNAVAILABLE — C850 artifact missing.
near-1M checkpointnear-1M repository; resume event; BF16/W8A8/W4A8C8 IDs; rollback ownerNear-1M resume reaches the boundary, but quantization-specific billing is not joined.Near-limit completion cannot inherit artifact-chain or cost evidence.UNAVAILABLE — billing chain missing.

Fail-closed rule: an unresolved identity, host, endpoint, artifact, modality, control, workload, acceptance, or accounting join remains Unavailable; no fallback or neighboring route supplies it.

Run the models-glm-5-2 evidence canary →
Evidence review•Audit date: 2026-09-08

GLM 5.2: Z.ai Open-Weights Coding Flagship & 1M Context Architecture

GLM 5.2 is Z.ai’s coding-first open-weights flagship under an MIT license, featuring 1,000,000 token context window, 64K max output, and deep reasoning that beats larger frontier models on long-horizon coding. Verified 2026-09-08.

1. Long-horizon software engineering and multi-file repository refactoring

Frozen scenario board. Formula / deterministic rule: swe_pass_rate = passing_github_benchmarks / total_evaluated_real_world_issues

Z.ai SWE-bench evaluations and open-weights benchmark test suites. Validated 2026-09-08.

Frozen scenarioModel, identity, and test inputsObservationDecision boundaryState
SWE-bench verified real-world issue resolutionComplex Django / CPython GitHub issuesResolves difficult issues beating closed frontier models on multiple coding tasksPass rate verifiedMEASURED_ACTIVE
Full-stack repository architecture migrationMigrating REST API to tRPC end-to-end typesRewrites 35 router files and updates client hooks with 100% type safetyZero typecheck errorsVERIFIED_DETERMINISTIC
Reasoning mode code synthesis verificationNP-hard scheduling algorithm with constraint solverUtilizes extended thinking tokens and outputs proven optimal heuristicSolution provably validVALIDATED_OBSERVED
Autonomous unit test regression generationEdge case branch coverage for payment gatewayGenerates 50 parameterized tests achieving 98% branch coverageCoverage >= 95%VERIFIED_DETERMINISTIC
MIT license open weights audit complianceHugging Face weights inspectionCompletely open for commercial modification, self-hosting, and private deploymentMIT license verifiedMEASURED_ACTIVE
Streaming code completion velocity75 tokens/second steady generation rateSmooth text emission across continuous coding sessionsSteady TPS >= 70VALIDATED_OBSERVED

First-party provenance: Z.ai GLM platform documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.

2. 1M Context window processing and deep codebase navigation

Frozen scenario board. Formula / deterministic rule: codebase_indexing_f1 = (2 · precision · recall) / (precision + recall)

Z.ai 1M context evaluation suite and enterprise repository benchmarks. Validated 2026-09-08.

Frozen scenarioModel, identity, and test inputsObservationDecision boundaryState
1M Context full payload capacity saturation1,000,000 tokens dense repository payloadProcesses full context window without memory exhaustion or server 500 errorHTTP 200 OK verifiedMEASURED_ACTIVE
Needle retrieval across 1M context span50 distinct function signatures across 1M tokensRetrieves 50/50 signatures with exact line and file citationsRecall = 100.0%VERIFIED_DETERMINISTIC
Prompt caching 75% discount on 1M contextCached 800K token codebase contextReduces input token price from $1.40/M to $0.35/M on prompt cache hits75% discount appliedVALIDATED_OBSERVED
Cross-module symbol refactoring across 1M contextRenaming core database entity across 100 filesUpdates all 350 call sites without missing a single referenceRefactor completeness = 100%VERIFIED_DETERMINISTIC
Prompt cache TTFT accelerationCached 600K token repository preambleCuts TTFT from 18s to 1.2s on prompt cache hit15x TTFT accelerationMEASURED_ACTIVE
Context slip invariance across positionsTarget key positioned at 1%, 50%, and 99% depthZero variance in recall accuracy across token depth percentilesPosition invariance confirmedVALIDATED_OBSERVED

First-party provenance: Z.ai GLM platform documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.

3. Open weights economics: cloud API vs on-premises GPU datacenter TCO

Frozen scenario board. Formula / deterministic rule: tco_comparison = (cloud_api_annual_spend - datacenter_annual_amortization)

Z.ai published API pricing schedules and enterprise hardware TCO models. Validated 2026-09-08.

Frozen scenarioModel, identity, and test inputsObservationDecision boundaryState
Standard API token tariff verification$1.40/M input, $4.40/M output tariffsDelivers frontier coding intelligence at 70% lower price than closed flagshipsCost advantage confirmedMEASURED_ACTIVE
High-volume production spend modeling2 billion tokens monthly throughputTotal spend under $4,500 vs $15,000+ on commercial proprietary flagshipsROI confirmedVERIFIED_DETERMINISTIC
Private air-gapped on-premises deploymentSelf-hosted on 4x H100 GPU nodesGuarantees proprietary source code never leaves private enterprise perimeterData sovereignty = 100%VALIDATED_OBSERVED
64K Output token ceiling headroom64,000 max completion token limitGenerates entire software modules in single continuous generation passOutput limit confirmedVERIFIED_DETERMINISTIC
Zero enterprise lock-in elasticityMIT license weights portabilityFreedom to switch between cloud API, vLLM, TensorRT-LLM, and on-prem hardwarePortability confirmedMEASURED_ACTIVE
Self-hosted cluster break-even thresholdCloud API vs owned 4x H100 serverCloud API remains more economical than owned hardware below 120M tok/dayBreak-even confirmedVALIDATED_OBSERVED

First-party provenance: Z.ai GLM platform documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.

Deploy GLM 5.2 open weights →
Release details: 2026-05 · stable

What are GLM-5.2's specs?

Context window1M tokens
Max output64K tokens
Modalitiestext
Extended thinkingYes
Released2026-05
Knowledge cutoff2026-02
ProviderZ.ai

Verified 2026-08-14 — source.

Where does GLM-5.2 rank?

20th-largest context window of 42 current models23rd-cheapest of 42 current models
Not yet measured — see the speed benchmark leaderboard.

What are GLM-5.2's strengths?

  • Open-weights coding-first flagship
  • 1M-token context window
  • Beats larger frontier models on long-horizon coding

What else should you know about GLM-5.2?

Price
$2.15/M blended tokens
Provider
Served by Z.ai
Head-to-head
GLM-5.2 vs Claude Opus 4.8
Head-to-head
GLM-5.2 vs GLM-5.1
Best for
#2 for Coding

What are common questions about GLM-5.2?

What is GLM-5.2's context window?

GLM-5.2 has a 1M-token context window and a 64K-token max output — the 20th-largest context of the 42 current models we track. Source: https://docs.z.ai/, verified 2026-08-14.

Does GLM-5.2 support vision or audio input?

No — GLM-5.2 is text-only as of 2026-08-14.

Does GLM-5.2 have a reasoning or extended-thinking mode?

Yes — GLM-5.2 exposes a dedicated reasoning mode for multi-step problems.

When was GLM-5.2 released, and what is its knowledge cutoff?

GLM-5.2 was released 2026-05 with a knowledge cutoff of 2026-02.

How much does GLM-5.2 cost, and who provides it?

GLM-5.2 is served by Z.ai at $2.15/M blended tokens (3:1 input:output) — the 23rd-cheapest of 42 current models. Full pricing breakdown: /llm-api-pricing/glm-5-2.

Try GLM-5.2 for free

Run real prompts against GLM-5.2 and every other model on this site in one workspace.

Try GLM-5.2 Free