Live results · Jun 21, 2026 · Audit verified 2026-09-01

Premium AI Model Tests — The Flagships on Genuinely Hard Tasks

Our cheap-model tests answer “what's the best value LLM?” This page asks the opposite question: when you reach for the most capable models money can buy, what do the extra dollars and seconds actually buy you? We took the four high-end flagships — GPT-5.4 Pro, Claude Opus 4.6, Gemini 3.1 Pro and GLM 5.2 in Max reasoning mode — and sent each the same deliberately hard prompts through the live All AI Ask API. Every number below is real, and both tasks were verified programmatically, not just eyeballed.

2
Hard tasks
4
Flagship models
8
Live API runs
41×
Price spread

The tasks

∴
Complex Reasoning

Solving a Constraint Logic Puzzle

A five-house constraint-satisfaction puzzle with five interlocking clues and a single valid solution. The model must reason through both cases, eliminate the dead end, and report the exact arrangement and count.

Best value: Claude Opus 4.8 (100/100)
View full results →
〈/〉
Advanced Coding

Median of Two Sorted Arrays in O(log n)

A classic hard algorithm: compute the median of two sorted lists in O(log(min(m,n))) time. A merge is explicitly disallowed, so the model must implement the tricky binary-search partition correctly — including empty-list and even/odd edge cases — and return code only.

Best value: GLM 5.2 (Max) (100/100)
View full results →

Overall leaderboard

Averaged across both hard tasks. Accuracy is graded against each task's published criteria and cross-checked programmatically. Ties on accuracy break toward the cheaper model.

#ModelAvg accuracyAvg speedList priceTotal cost
🥇
Claude Opus 4.8Anthropic
10099.1 t/s$25/M$0.036505
2
GPT-5.4 ProOpenAI
10012.2 t/s$180/M$0.2565
3
GLM 5.2 (Max)Z.ai
99.559.4 t/s$4.4/M$0.02237
4
Gemini 3.1 ProGoogle
98.520.5 t/s$12/M$0.010744

Evidence guide. Every section below is computed from dated scenarios; historical outputs are not presented as new runs. Verification date: 2026-09-01.

Deterministic test receipts

These values are read from the exported test records. Prompt hashes identify the immutable test input; raw-output hashes identify each verbatim model response. Version and validation fields are recorded metadata, not a new run or a re-grade.

Test / modelPrompt SHA-256Raw output SHA-256GraderSolver / harnessValidation flags
constraint-logic-puzzle
gpt-5.4-pro
119b0025774a31bd409ae9f1f387f96eaea0c8eb3f793dc0186496fecc3df336beea172063b8f2d3e67aa379ad19d2339907b45b4525c9897ca741db8b2e46b5claude-agent-grader-v1constraint-solver-v1-120-permutationsrecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent
constraint-logic-puzzle
claude-opus-4-8
119b0025774a31bd409ae9f1f387f96eaea0c8eb3f793dc0186496fecc3df336cac6fe506285c705adbd84cb5a5d0e1560dbde69fe6ce791e76167c1ebd38fe5claude-agent-grader-v1constraint-solver-v1-120-permutationsrecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected
constraint-logic-puzzle
gemini-3.1-pro
119b0025774a31bd409ae9f1f387f96eaea0c8eb3f793dc0186496fecc3df336598a199b193a2c86c727861f860be1aafa3c6b395b594c7a2e95c1863f4051ddclaude-agent-grader-v1constraint-solver-v1-120-permutationsrecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent
constraint-logic-puzzle
glm-5.2
119b0025774a31bd409ae9f1f387f96eaea0c8eb3f793dc0186496fecc3df3362dfceaafb7a800a36299028ff1928a564e4158c4228e5d77bf4c32a676288310claude-agent-grader-v1constraint-solver-v1-120-permutationsrecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected
median-two-sorted-arrays
gpt-5.4-pro
abc8c17072e616feae256681c64aa944e31b145375338fb5269d4eb831708a57a2aefdcfc47376d961f5c43a54d8fe0f7226dd2a8e0334b1127cbd755f2f3945claude-agent-grader-v1median-harness-v1-seed-2026-5000-casesrecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, code-fence=present
median-two-sorted-arrays
claude-opus-4-8
abc8c17072e616feae256681c64aa944e31b145375338fb5269d4eb831708a572ba7b9c31f2a871ff6037dee27a26cfa4aba1d36f417e2b01ac65473b802a3a6claude-agent-grader-v1median-harness-v1-seed-2026-5000-casesrecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, code-fence=present
median-two-sorted-arrays
gemini-3.1-pro
abc8c17072e616feae256681c64aa944e31b145375338fb5269d4eb831708a57c87151f5b92f95381d9b0bf0ec89b041834f164eda008b3f0a2baa885fc37617claude-agent-grader-v1median-harness-v1-seed-2026-5000-casesrecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, code-fence=present
median-two-sorted-arrays
glm-5.2
abc8c17072e616feae256681c64aa944e31b145375338fb5269d4eb831708a571e909c8417f150d1aac0eb4b7d4c3219ed4678847a6c0d18a24005722a5b6ae3claude-agent-grader-v1median-harness-v1-seed-2026-5000-casesrecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected, code-fence=present

Premium cohort identity and eligibility receipt

Frozen fixture board. Formula / decision rule: historical cohort = exact ID + host + snapshot + available run + comparable prompt; later catalog entries do not append Boundary: The four-model 2026-06-21 cohort is immutable; a renamed or newer model cannot be silently added.

Frozen fixtureIdentity keysDeterministic ruleOutput / bounded stateValidation
four included exact IDssuite=premium; run=2026-06-21; provider/host/model IDs=exact; snapshots=joined; two tests=availableall four exact IDs and both test joins are presenthistorical inclusion=Included; cohort size=4INCLUDED — frozen cohort
renamed aliassuite=premium; run=2026-06-21; requested ID=alias; resolved ID=not proven equal; host=joinedalias rename requires exact ID equivalence evidenceinclusion=Unavailable; do not merge by display nameUNAVAILABLE — alias join
newer untested flagshipsuite=premium; run=2026-06-21; model=post-run flagship; run availability=noneno historical run means no cohort membershipexclude from historical four-model cohortEXCLUDED — untested
failed endpointsuite=premium; run=2026-06-21; exact model/host=joined; endpoint status=failed; raw hash=presentfailed endpoint has no valid comparable resultexclude from complete cohort; preserve failure receiptEXCLUDED — failed endpoint
alternate hostsuite=premium; run=2026-06-21; model display name=same; host=alternate; snapshot=Unknownhost is an identity key, not a display labelcomparability=No; do not transfer runEXCLUDED — host mismatch
unresolved model-version fixturessuite=premium; run=2026-06-21; model ID=ambiguous; snapshot=Unknown; prompt hash=joinedversion ambiguity blocks selectioninclusion=Unavailable; require version joinUNAVAILABLE — version identity

Provenance: premium-model-tests module 1; historical run 2026-06-21; audit verification 2026-09-01. First-party premium-model suite evidence (JSON). Missing or conflicting joins fail closed.

Two-test consistency and hard-failure matrix

Frozen fixture board. Formula / decision rule: complete = valid solver/harness + rubric + identity joins for both tests; hard failure blocks overall rank Boundary: One missing or invalid required test prevents an overall premium-suite rank.

Frozen fixtureIdentity keysDeterministic ruleOutput / bounded stateValidation
both-passsuite=premium; model/host=exact; logic solver=pass; median harness=pass; rubric=joined; run=2026-06-21both required tests pass their verification gatescomplete-case=Yes; cross-test consistency=PassPASS — both tests
logic-pass/code-failsuite=premium; exact test/model join; logic solver=pass; median harness=fail; raw hashes=joinedany hard failure blocks suite acceptancehard-failure=code; overall rank=UnavailableBLOCKED — code failure
code-pass/logic-failsuite=premium; exact test/model join; logic solver=fail; median harness=pass; raw hashes=joinedany hard failure blocks suite acceptancehard-failure=logic; overall rank=UnavailableBLOCKED — logic failure
instruction-leaksuite=premium; exact model/run; test output contains reasoning wrapper where code-only required; grader=joinedcode-only hard failure is preserved separately from correctnesscomplete-case=No; no overall rankBLOCKED — instruction leak
incomplete runsuite=premium; exact model/run; one test status=missing; prompt hash=present; metrics=partialmissing test evidence breaks two-test denominatorcomplete-case=No; exclusion reason=incomplete runBLOCKED — incomplete
missing-harness fixturessuite=premium; exact model/run; logic solver version=present; median harness version=missingverification version is decision-criticaloverall rank=Unavailable; require harness joinUNAVAILABLE — harness identity

Provenance: premium-model-tests module 2; historical run 2026-06-21; audit verification 2026-09-01. First-party premium-model suite evidence (JSON). Missing or conflicting joins fail closed.

Premium observed-frontier sensitivity board

Frozen fixture board. Formula / decision rule: frontier(policy) = eligible recorded metrics after declared hard gates/normalization; missing price => Unavailable Boundary: Two frozen prompts cannot establish a universal frontier-model verdict.

Frozen fixtureIdentity keysDeterministic ruleOutput / bounded stateValidation
accuracy-firstsuite=premium; policy=accuracy-first; recorded rubric scores=joined; eligible models=complete casesselect by recorded accuracy then declared tie rulewinner/tie=Unavailable until all complete scores joinUNAVAILABLE — score join
both-tests-passsuite=premium; policy=both-tests-pass; solver/harness gates=joined; eligible models=passersfilter to models passing both exact verifierseligible set=Unavailable without complete receiptUNAVAILABLE — eligibility
lowest run cost among passerssuite=premium; policy=lowest recorded run cost; cost=historical only; pass gates=joinedchoose lowest recorded cost among exact passerswinner=Unavailable if any cost field is missing; current tariff excludedUNAVAILABLE — cost field
lowest latency among passerssuite=premium; policy=lowest recorded latency; latency=joined; pass gates=joinedchoose lowest recorded latency among exact passerswinner/tie=Unavailable without joined latency vectorUNAVAILABLE — latency field
balanced normalized scoresuite=premium; policy=balanced; normalization=declared per test; complete cases=joinednormalize only recorded test metrics under declared policyrank movement=Unavailable; no cross-test imputationUNAVAILABLE — normalization
no-price-field policiessuite=premium; policy=quality/latency only; price=missing; exact run IDs=joinedquality/latency policy may proceed only if its own fields joinprice-dependent result=Unavailable; two-prompt boundary remainsBOUNDED — no price

Provenance: premium-model-tests module 3; historical run 2026-06-21; audit verification 2026-09-01. First-party premium-model suite evidence (JSON). Missing or conflicting joins fail closed.

Run this scenario →

How we tested

  • The cohort: the four high-end flagships — GPT-5.4 Pro, Claude Opus 4.6, Gemini 3.1 Pro, and GLM 5.2 run in Max reasoning mode.
  • Identical prompts: each model received the same prompt through the live API; GLM 5.2 used reasoningEffort: "max", the others their default flagship settings.
  • Real metrics: latency, token counts, and cost come straight from the API for each run. (GPT-5.4 Pro's token count is estimated from output length — the API under-reports it for that model.)
  • Verified accuracy: the logic puzzle was checked by exhaustive search and every code answer was executed against a 5,000-case correctness harness — so the grades aren't guesswork.

Pit the flagships against your own hard prompt

Send one prompt to every premium model at once and watch the speed, cost, and quality side by side.

Try the live playground free →