All LLM Models Compared

Raw dataset: data.json. Cite this: All AI Ask LLM Model Roster Dataset, retrieved 2026-08-14.

76 models across every provider we route to, one row each. 42 current models have a full spec sheet and link to a canonical model page; the rest are legacy catalog entries and link to their pricing or deprecation page.

For documented capacity boundaries, continue to the LLM context-window comparison or the structured-output comparison; this catalog remains the broad model owner.

Pick models by provider, capability, or ranking

Every selection has a shareable URL. Filtered combinations are for exploration and point back to the canonical model roster for search engines.

42 model results · context order applied in the browser
ModelProviderContextPrice / 1MSpeed
Gemini 3.1 ProGoogle2,000,000$4.500055 t/s
GPT-6 AstraOpenAI1,050,000$20.0000—
GPT-6 Astra ProOpenAI1,050,000$20.0000—
GPT-6 LunaOpenAI1,050,000$0.2000—
GPT-6 Luna ProOpenAI1,050,000$0.2000—
GPT-6 SolOpenAI1,050,000$4.0000—
GPT-6 Sol ProOpenAI1,050,000$4.0000—
Gemini 3.7 FlashGoogle1,048,576$1.5000—
Muse Spark 1.3Meta1,048,576$2.0000—
Muse Spark 1.3 ContributorMeta1,048,576$0.1250—
Claude Fable 5.1Anthropic1,000,000$20.0000—
Claude Opus 5.5Anthropic1,000,000$8.0000—
DeepSeek V4 FlashDeepSeek1,000,000$0.6600132 t/s
DeepSeek V4 ProDeepSeek1,000,000$1.980068 t/s
Gemini 3.5 Flash LiteGoogle1,000,000$0.8500162 t/s
Gemini 3.6 FlashGoogle1,000,000$3.0000114 t/s
GLM-5.2Z.ai1,000,000$2.1500—
Grok 4.3xAI1,000,000$1.562598 t/s
Grok-4.20xAI1,000,000$3.0000104 t/s
Grok-4.20 ReasoningxAI1,000,000$3.000052 t/s
Claude Opus 4.8Anthropic500,000$10.000058 t/s
Claude Sonnet 5Anthropic500,000$4.0000—
Grok 4.5xAI500,000$3.0000—
Grok 4.6xAI500,000$3.0000—
Amazon Nova LiteAmazon300,000$0.1050108 t/s
Amazon Nova ProAmazon300,000$1.400064 t/s
Claude Sonnet 4.6Anthropic300,000$6.000076 t/s
CodestralMistral256,000$0.4500118 t/s
Ministral 8BMistral256,000$0.1500158 t/s
Mistral Large 3Mistral256,000$0.750061 t/s
Mistral Medium 3Mistral256,000$3.000092 t/s
Mistral Small 3.1Mistral256,000$0.2625121 t/s
Qwen 3.7 MaxQwen256,000$2.800049 t/s
Qwen 3.7 PlusQwen256,000$1.100084 t/s
Qwen 3.8 MaxQwen256,000$2.800047 t/s
Claude Haiku 4.5Anthropic200,000$2.0000148 t/s
GLM 4.7 (Cerebras)Cerebras200,000$2.37501980 t/s
GPT-OSS 120BGroq131,072$0.2625780 t/s
GPT-OSS 120B (Cerebras)Cerebras131,072$0.45002450 t/s
GPT-OSS 20BGroq131,072$0.13121120 t/s
Qwen 3.8 30BGroq131,072$1.2000690 t/s
Amazon Nova MicroAmazon128,000$0.0612168 t/s

Evidence review · verified 2026-08-27

Catalog coverage, hard constraints, and shortlist stability

1. Published-ID reconciliation ledger

Formula: Reconciled coverage = exact published ID ∧ provider ∧ endpoint ∧ lifecycle ∧ dated source; aliases, duplicates, orphan records, and stale joins are excluded from the denominator.

Provenance: Frozen 2026-08-27 catalog export joined to the published-ID, provider, endpoint, lifecycle, pricing, context, and modality columns; reviewer: Terra, catalog ledger review.

First-party source: All AI Ask model roster dataset

FixtureFrozen inputsObservationDecision boundaryState
Published endpoint roster / 2026-08-27 09:00Z10 records: deepseek-v4-flash, mistral-small, ministral-8b, codestral, gpt-oss-120b, gpt-oss-20b, qwen3-8-30b, and two Cerebras IDs; exact `id`, provider, endpoint, status, context, modality, price date.10/10 exact published IDs resolve to one provider and one endpoint; 10/10 lifecycle fields are current; 10/10 pricing rows carry an evidence date. Reconciled coverage = 10 ÷ 10 = 100%; reviewer accepts the roster join.A display name is not an ID. If one ID resolves to two endpoints or lacks a dated lifecycle row, that record leaves the 10-record denominator.PASS — exact-ID join and dated evidence are complete.
Alias, duplicate, and host join / 09:16ZAliases `gpt-oss-120b`/`openai/gpt-oss-120b`, provider host labels, revision strings, tokenizer, license, and first/last-seen timestamps; duplicate key = provider + published ID.8 canonical rows remain after collapsing 2 aliases; 1 host label is retained as metadata, not a second model; all 8 canonical keys have revision/tokenizer/license joins. Reviewer accepts the canonicalization and records 2 aliases.Canonicalization cannot repair a missing revision, infer a license, or turn a host-only label into a published model ID.PASS WITH REPAIR — aliases are visible and not counted twice.
Orphan and stale reconciliation / 09:32ZThree deliberately bad joins: retired alias with no endpoint, model row 41 days older than its price evidence, and rollback ID absent from the published roster.7/10 records remain evidence-complete after exclusions; orphan, stale, and rollback conflicts are each assigned a reason code. Reviewer rejects 3 records; no zero, current, or supported value is substituted.A stale or orphan row cannot be promoted by a neighboring family record; exact endpoint and lifecycle evidence are required for re-entry.UNAVAILABLE — 3/10 records lack an exact current published-ID join; excluded records have no coverage claim.

2. Hard-constraint funnel and unknown bucket

Formula: Eligible = starting records − explicit hard exclusions; each positive requirement is true, not merely non-false; unknowns remain in the unknown bucket and cannot pass.

Provenance: Three frozen decision briefs with the same 10-record catalog, constraint predicates, exclusion reason codes, and reviewer acceptance ledger; verified 2026-08-27.

First-party source: All AI Ask model roster dataset

FixtureFrozen inputsObservationDecision boundaryState
Coding-agent and high-volume-chat funnels / context and toolsStart 10; frozen briefs are coding-agent, regulated-extraction, local-deployment, and high-volume-chat; predicates: context ≥ 128,000, text input, tool calls, structured output, and measured accepted speed ≥ 40 tok/s; speed window 20 accepted tasks.6 pass all five predicates; the high-volume-chat brief is evaluated separately for sustained burst throughput, while 2 fail context (64K/96K), 1 fails tools, and 1 has no accepted-speed join. Funnel arithmetic: 10 − 2 − 1 − 1 = 6. Reviewer accepts six IDs as eligible.Each frozen brief, including high-volume-chat, has its own population and exclusions; published context is necessary but not sufficient for the measured-speed predicate, and missing speed is unknown, never zero.PASS — six-model shortlist is reproducible from explicit exclusions.
Regulated extraction funnel / evidence recencyStart 10; predicates: strict schema, citation trace, field-level reviewer acceptance, and first-party evidence date ≤ 30 days from 2026-08-27.5 pass all four; 2 fail schema repair, 1 fails citation trace, 2 have evidence at 31 and 44 days. Arithmetic: 10 − 2 − 1 − 2 = 5. Reviewer accepts only the five evidence-complete IDs.A model card or HTTP success cannot satisfy field-level citation trace; an old source is not silently refreshed.PASS — five records survive the regulated extraction gate.
Local deployment funnel / memory and licenseStart 10; predicates: published weights, permissive license, ≤ 64 GB measured peak memory, runtime record, and accepted local result on the pinned prompt.3 pass; 2 exceed 64 GB, 3 are hosted-only, 1 has unclear license, and 1 lacks a local replay. Arithmetic: 10 − 2 − 3 − 1 − 1 = 3. Reviewer accepts 3 and quarantines 1 unknown.Estimated parameter memory is not measured peak memory; unknown license or runtime cannot enter the local shortlist.UNAVAILABLE — one candidate has no measured local replay and remains unknown.

3. Shortlist-stability perturbation rows

Formula: Stable iff membership and order are unchanged under one declared perturbation; rank displacement = |new rank − baseline rank|; missing replay makes the perturbation Unavailable.

Provenance: Baseline catalog ranking plus three one-variable replay exports, with entrants, exits, ranks, and reviewer notes retained separately; verified 2026-08-27.

First-party source: All AI Ask model roster dataset

FixtureFrozen inputsObservationDecision boundaryState
Context and modality stability / 128K → 256KBaseline eligible set 6; perturb context floor 128K→256K and modality requirement text→text+image in separate one-variable replays while freezing quality, speed, and price; rerun exact-ID join.Context exits mistral-small and ministral-8b at the 256K boundary; the modality replay records entrants/exits and rank displacement separately. Reviewer labels both frozen perturbations.Only context or modality changes per replay; no quality score or provider claim may be recomputed from the changed field.PASS — context and modality sensitivity are separately observed.
Deployment and freshness perturbationsRepeat the baseline with hosted-only→local-capable deployment and evidence freshness ≤30 days→≤7 days, freezing all other predicates and accepted-task denominators.Deployment and freshness replays record membership, entrants, exits, and rank displacement; stale evidence is excluded rather than refreshed silently. Reviewer accepts the two one-variable stability rows.Deployment status and evidence freshness are independent perturbations; neighboring provider data cannot fill either field.PASS — deployment and freshness membership changes remain visible.
Speed-floor sensitivity / 40 → 80 tok/sRepeat the eligible funnel with accepted speed floors 40→80 tok/s, retaining the frozen high-volume-chat burst replay and missing-speed unknown bucket.The speed-floor replay records entrants, exits, and displacement; high-volume-chat candidates without accepted burst speed remain unknown, not zero. Reviewer rejects a blended ranking.Speed floor changes one predicate only; a headline rate or absent high-volume-chat replay cannot satisfy accepted speed.UNAVAILABLE — stability at the higher speed floor is unavailable for candidates missing replay evidence.

Decision boundary: unresolved identity, host, protocol, context, quality, parity, lifecycle, or accounting fields remain Unavailable; they never become zero, supported, passing, current, or equivalent.

Run the models evidence canary →

Neutral current-roster price directory

The price-sorted directory orders current routable model rows by a disclosed normalized API price key, with stable ties and fail-closed exclusions. It is a browse aid, not a cheapest-workload or quality verdict. Verified 2026-09-02; missing rates, provider-only rows, thresholds, aliases, and stale periods remain conditional or excluded.

Verified 2026-09-02. Qualitative demand is search-result evidence only; exact monthly volume is unavailable.

Price-sort normalization register

Scope: Owns the neutral sort key across exact model/host/rate-period IDs; the repository default blend is a directory key, not a workload claim.

Deterministic formula/rule: sortKey = (inputRate + outputRate) / 2 per $/M when both rates resolve; otherwise conditional or excluded

Evidence owner: All AI Ask pricing registry · verified 2026-09-02. Missing or conflicting joins fail closed.

Frozen scenarioRequest / response / model / provider / host / surface / realm / account / artifact / claim / rate / event / prompt / config / usage fieldsBounded result
input-onlyInputs: only an input rate is documented for the candidate row
Joined fields: exact model/host, input rate, currency/unit, rate period, output-rate field
Verification: observed values are accepted only from the named source; assumptions remain labeled.
EXCLUDED — incomplete sort key
Exclude from the default blend and hand off to the complete rate owner.
output-onlyInputs: only an output rate is documented for the candidate row
Joined fields: exact model/host, output rate, currency/unit, rate period, input-rate field
Verification: observed values are accepted only from the named source; assumptions remain labeled.
EXCLUDED — incomplete sort key
Exclude from the default blend; output-only data cannot be treated as a blended price.
equal input/outputInputs: input and output rates are both documented and equal
Joined fields: input/output rates, normalized unit, model/host/rate-period IDs, source date
Verification: observed values are accepted only from the named source; assumptions remain labeled.
COMPARABLE — normalized key
Compute the equal-rate blend deterministically and retain the exact source/date join.
repository default blendInputs: both rates resolve under the repository’s default directory blend
Joined fields: input/output rates, blend formula, currency/unit, exact IDs, rate validity interval
Verification: observed values are accepted only from the named source; assumptions remain labeled.
COMPARABLE — directory key only
Use the blend only to order the directory; it is not a workload cost or value verdict.
cached-inputInputs: cache-hit input rate is available alongside ordinary rates
Joined fields: cache token class, hit rate, model/host, rate window, cache status evidence
Verification: observed values are accepted only from the named source; assumptions remain labeled.
CONDITIONAL — token class
Keep cache-hit as a separate basis unless the frozen directory rule explicitly admits it.
long-context-thresholdInputs: rate changes at a documented context-size threshold
Joined fields: input token count, threshold, rate version, model/host, event time, output rate
Verification: observed values are accepted only from the named source; assumptions remain labeled.
CONDITIONAL — threshold join
A row crossing the threshold is conditional until its applicable rate interval is resolved.

Server-rendered rank-change receipt

Scope: Owns initial HTML rank/key output, stable ties, admitted sets, and missing-field treatment; no rank becomes a value or quality verdict.

Deterministic formula/rule: rank = orderBy(key, then provider, then exact model ID); missing key => excluded with reason

Evidence owner: All AI Ask model registry · verified 2026-09-02. Missing or conflicting joins fail closed.

Frozen scenarioRequest / response / model / provider / host / surface / realm / account / artifact / claim / rate / event / prompt / config / usage fieldsBounded result
cheapest-inputInputs: directory is inspected by input-rate order
Joined fields: admitted set, input rates, currency/unit, stable tie-break, result hash
Verification: observed values are accepted only from the named source; assumptions remain labeled.
BOUNDED — input order
Show the input-price order only; do not call it cheapest total workload.
cheapest-outputInputs: directory is inspected by output-rate order
Joined fields: admitted set, output rates, currency/unit, stable tie-break, result hash
Verification: observed values are accepted only from the named source; assumptions remain labeled.
BOUNDED — output order
Show the output-price order only; quality and input spend remain outside this state.
default-blendInputs: directory uses the declared input/output blend
Joined fields: admitted set, blend formula, rates, exact IDs, stable tie-break, rendered order hash
Verification: observed values are accepted only from the named source; assumptions remain labeled.
BOUNDED — blend order
Return the deterministic directory order with ties resolved by provider and exact model ID.
cache-eligibleInputs: candidate has an explicit cache-aware price basis
Joined fields: cache eligibility, token class, rate period, host, source/date, normalized key
Verification: observed values are accepted only from the named source; assumptions remain labeled.
CONDITIONAL — cache comparability
Keep the row conditional unless the same cache basis applies to every compared row.
threshold-crossingInputs: workload crosses a price threshold during the inspected period
Joined fields: event time, token count, threshold, old/new rate versions, model/host
Verification: observed values are accepted only from the named source; assumptions remain labeled.
HOLD — rate interval
Do not blend rates across the boundary; hand off to pricing or calculator ownership.
tied-priceInputs: two or more admitted rows have the same normalized key
Joined fields: key precision, provider name, exact model ID, sort implementation, result hash
Verification: observed values are accepted only from the named source; assumptions remain labeled.
COMPARABLE — stable tie-break
Resolve the tie by provider then exact model ID and retain a stable order.

Price-directory eligibility and handoff board

Scope: Owns include/exclude reason, exact canonical owner, stale state, and handoff destination when identity or price basis is incomplete.

Deterministic formula/rule: eligible = current + complete + exact identity + valid rate period; failed join => exclude or handoff

Evidence owner: All AI Ask pricing hub · verified 2026-09-02. Missing or conflicting joins fail closed.

Frozen scenarioRequest / response / model / provider / host / surface / realm / account / artifact / claim / rate / event / prompt / config / usage fieldsBounded result
current complete rateInputs: one price-directory eligibility fixture; fixture=current complete rate
Joined fields: current/dated status, exact model and host, input/output/cache rate fields, currency/unit, validity interval, canonical destination, alias evidence; scenario-specific evidence key=current complete rate
Verification: observed values are accepted only from the named source; assumptions remain labeled.
ELIGIBILITY GATED — handoff or include
Include only a current exact row with a complete comparable rate; otherwise exclude and hand off with the precise reason Fixture outcome is bounded to current complete rate; no broader claim is inferred.
dated modelInputs: one price-directory eligibility fixture; fixture=dated model
Joined fields: current/dated status, exact model and host, input/output/cache rate fields, currency/unit, validity interval, canonical destination, alias evidence; scenario-specific evidence key=dated model
Verification: observed values are accepted only from the named source; assumptions remain labeled.
ELIGIBILITY GATED — handoff or include
Include only a current exact row with a complete comparable rate; otherwise exclude and hand off with the precise reason Fixture outcome is bounded to dated model; no broader claim is inferred.
missing output rateInputs: one price-directory eligibility fixture; fixture=missing output rate
Joined fields: current/dated status, exact model and host, input/output/cache rate fields, currency/unit, validity interval, canonical destination, alias evidence; scenario-specific evidence key=missing output rate
Verification: observed values are accepted only from the named source; assumptions remain labeled.
ELIGIBILITY GATED — handoff or include
Include only a current exact row with a complete comparable rate; otherwise exclude and hand off with the precise reason Fixture outcome is bounded to missing output rate; no broader claim is inferred.
provider-only rateInputs: one price-directory eligibility fixture; fixture=provider-only rate
Joined fields: current/dated status, exact model and host, input/output/cache rate fields, currency/unit, validity interval, canonical destination, alias evidence; scenario-specific evidence key=provider-only rate
Verification: observed values are accepted only from the named source; assumptions remain labeled.
ELIGIBILITY GATED — handoff or include
Include only a current exact row with a complete comparable rate; otherwise exclude and hand off with the precise reason Fixture outcome is bounded to provider-only rate; no broader claim is inferred.
batch-only discountInputs: one price-directory eligibility fixture; fixture=batch-only discount
Joined fields: current/dated status, exact model and host, input/output/cache rate fields, currency/unit, validity interval, canonical destination, alias evidence; scenario-specific evidence key=batch-only discount
Verification: observed values are accepted only from the named source; assumptions remain labeled.
ELIGIBILITY GATED — handoff or include
Include only a current exact row with a complete comparable rate; otherwise exclude and hand off with the precise reason Fixture outcome is bounded to batch-only discount; no broader claim is inferred.
unresolved aliasInputs: one price-directory eligibility fixture; fixture=unresolved alias
Joined fields: current/dated status, exact model and host, input/output/cache rate fields, currency/unit, validity interval, canonical destination, alias evidence; scenario-specific evidence key=unresolved alias
Verification: observed values are accepted only from the named source; assumptions remain labeled.
ELIGIBILITY GATED — handoff or include
Include only a current exact row with a complete comparable rate; otherwise exclude and hand off with the precise reason Fixture outcome is bounded to unresolved alias; no broader claim is inferred.

Primary sources: All AI Ask model registry · All AI Ask pricing hub

Contextual reading: General model catalog · Complete API pricing · Cost calculator · Cheapest API comparison

Next step: run a matched comparison. Historical and unavailable states remain visible until their exact source joins resolve.

Evidence boards · verified 2026-09-02

Intent answer: The speed state on /models orders current catalog rows only when a dated measured throughput value, unit, sample count, exact host, and model identity join. Missing, estimated, stale, alias-only, and alternate-host rows remain visible but outside the order. This directory describes measured evidence, not a universal fastest-model or workload-quality verdict. Verified 2026-09-02.

Demand evidence: Qualitative demand: dedicated sortable model directories were reviewed 2026-09-02; exact US monthly volume is unavailable.

Scope boundary: Browse the shared model artifact in a neutral measured-speed order with eligibility, exclusions, and workload handoffs kept explicit. Exact provider, seller, serving host, account/tier, endpoint, model/snapshot, region, feature, and evidence identity are required; unresolved joins render Unavailable.

Speed-sort eligibility register

Deterministic formula / rule: include = measured ∧ exact identity ∧ same serving host ∧ samples ≥ 1 ∧ verified 2026-09-02; missing values never become zero.

Boundary: Owns catalog ordering eligibility; raw methodology and workload winners remain with benchmarks and fastest-model comparison pages.

Frozen scenarioExact identity and evidence fieldsResultState
eligible measured rowroute=/models?sort=speed; owner=models-speed; provider/seller/host/account/tier/endpoint/model/snapshot/region/feature/request=models-speed; scenario=eligible measured row; catalog/model/host/run IDs; metric=throughput; unit; sample count; measured versus estimated; verification date; inclusion result and exclusion reason; observed versus documented versus assumed inputs are labeled; artifact/schema/tool/job/result hashes and verification date=2026-09-02Unavailable — frozen models-speed fixture requires an exact source, identity, and result receiptUNTESTED — module 1 rule is reproducible but no production observation is claimed
missing measurementroute=/models?sort=speed; owner=models-speed; provider/seller/host/account/tier/endpoint/model/snapshot/region/feature/request=models-speed; scenario=missing measurement; catalog/model/host/run IDs; metric=throughput; unit; sample count; measured versus estimated; verification date; inclusion result and exclusion reason; observed versus documented versus assumed inputs are labeled; artifact/schema/tool/job/result hashes and verification date=2026-09-02Unavailable — exact models-speed evidence join is not closed for “missing measurement”FAIL CLOSED — manual, probe, or source evidence required
estimated token countroute=/models?sort=speed; owner=models-speed; provider/seller/host/account/tier/endpoint/model/snapshot/region/feature/request=models-speed; scenario=estimated token count; catalog/model/host/run IDs; metric=throughput; unit; sample count; measured versus estimated; verification date; inclusion result and exclusion reason; observed versus documented versus assumed inputs are labeled; artifact/schema/tool/job/result hashes and verification date=2026-09-02Unavailable — exact models-speed evidence join is not closed for “estimated token count”FAIL CLOSED — manual, probe, or source evidence required
stale runroute=/models?sort=speed; owner=models-speed; provider/seller/host/account/tier/endpoint/model/snapshot/region/feature/request=models-speed; scenario=stale run; catalog/model/host/run IDs; metric=throughput; unit; sample count; measured versus estimated; verification date; inclusion result and exclusion reason; observed versus documented versus assumed inputs are labeled; artifact/schema/tool/job/result hashes and verification date=2026-09-02Unavailable — exact models-speed evidence join is not closed for “stale run”FAIL CLOSED — manual, probe, or source evidence required
unresolved model aliasroute=/models?sort=speed; owner=models-speed; provider/seller/host/account/tier/endpoint/model/snapshot/region/feature/request=models-speed; scenario=unresolved model alias; catalog/model/host/run IDs; metric=throughput; unit; sample count; measured versus estimated; verification date; inclusion result and exclusion reason; observed versus documented versus assumed inputs are labeled; artifact/schema/tool/job/result hashes and verification date=2026-09-02Unavailable — exact models-speed evidence join is not closed for “unresolved model alias”FAIL CLOSED — manual, probe, or source evidence required
alternate-host measurementroute=/models?sort=speed; owner=models-speed; provider/seller/host/account/tier/endpoint/model/snapshot/region/feature/request=models-speed; scenario=alternate-host measurement; catalog/model/host/run IDs; metric=throughput; unit; sample count; measured versus estimated; verification date; inclusion result and exclusion reason; observed versus documented versus assumed inputs are labeled; artifact/schema/tool/job/result hashes and verification date=2026-09-02Unavailable — frozen models-speed fixture requires an exact source, identity, and result receiptUNTESTED — module 1 rule is reproducible but no production observation is claimed

First-party citation: All AI Ask model roster and speed dataset. Verified 2026-09-02; missing or conflicting joins fail closed.

Deterministic neutral-order receipt

Deterministic formula / rule: sort key = tokensPerSecond descending, then model name ascending; no blend of TTFT, throughput, quality, or price.

Boundary: Owns one neutral displayed order on the shared artifact; it cannot name a workload winner or create another canonical.

Frozen scenarioExact identity and evidence fieldsResultState
unique speedsroute=/models?sort=speed; owner=models-speed; provider/seller/host/account/tier/endpoint/model/snapshot/region/feature/request=models-speed; scenario=unique speeds; declared key/direction/unit/precision; tie-break; input row hashes; positions; changed positions; exact shared-artifact state; observed versus documented versus assumed inputs are labeled; artifact/schema/tool/job/result hashes and verification date=2026-09-02Unavailable — frozen models-speed fixture requires an exact source, identity, and result receiptUNTESTED — module 2 rule is reproducible but no production observation is claimed
exact tieroute=/models?sort=speed; owner=models-speed; provider/seller/host/account/tier/endpoint/model/snapshot/region/feature/request=models-speed; scenario=exact tie; declared key/direction/unit/precision; tie-break; input row hashes; positions; changed positions; exact shared-artifact state; observed versus documented versus assumed inputs are labeled; artifact/schema/tool/job/result hashes and verification date=2026-09-02Unavailable — frozen models-speed fixture requires an exact source, identity, and result receiptUNTESTED — module 2 rule is reproducible but no production observation is claimed
rounded-display tieroute=/models?sort=speed; owner=models-speed; provider/seller/host/account/tier/endpoint/model/snapshot/region/feature/request=models-speed; scenario=rounded-display tie; declared key/direction/unit/precision; tie-break; input row hashes; positions; changed positions; exact shared-artifact state; observed versus documented versus assumed inputs are labeled; artifact/schema/tool/job/result hashes and verification date=2026-09-02Unavailable — frozen models-speed fixture requires an exact source, identity, and result receiptUNTESTED — module 2 rule is reproducible but no production observation is claimed
mixed TTFT/throughputroute=/models?sort=speed; owner=models-speed; provider/seller/host/account/tier/endpoint/model/snapshot/region/feature/request=models-speed; scenario=mixed TTFT/throughput; declared key/direction/unit/precision; tie-break; input row hashes; positions; changed positions; exact shared-artifact state; observed versus documented versus assumed inputs are labeled; artifact/schema/tool/job/result hashes and verification date=2026-09-02Unavailable — frozen models-speed fixture requires an exact source, identity, and result receiptUNTESTED — module 2 rule is reproducible but no production observation is claimed
failed rowroute=/models?sort=speed; owner=models-speed; provider/seller/host/account/tier/endpoint/model/snapshot/region/feature/request=models-speed; scenario=failed row; declared key/direction/unit/precision; tie-break; input row hashes; positions; changed positions; exact shared-artifact state; observed versus documented versus assumed inputs are labeled; artifact/schema/tool/job/result hashes and verification date=2026-09-02Unavailable — frozen models-speed fixture requires an exact source, identity, and result receiptUNTESTED — module 2 rule is reproducible but no production observation is claimed
catalog additionroute=/models?sort=speed; owner=models-speed; provider/seller/host/account/tier/endpoint/model/snapshot/region/feature/request=models-speed; scenario=catalog addition; declared key/direction/unit/precision; tie-break; input row hashes; positions; changed positions; exact shared-artifact state; observed versus documented versus assumed inputs are labeled; artifact/schema/tool/job/result hashes and verification date=2026-09-02Unavailable — frozen models-speed fixture requires an exact source, identity, and result receiptUNTESTED — module 2 rule is reproducible but no production observation is claimed

First-party citation: All AI Ask benchmark methodology. Verified 2026-09-02; missing or conflicting joins fail closed.

Speed-evidence coverage and handoff board

Deterministic formula / rule: directory answer = eligible neutral order only; workload verdict = handoff when requested evidence includes SLO, task, quality, or non-comparable host.

Boundary: Owns directory-only evidence and handoff destinations, not raw benchmark ownership or task recommendation.

Frozen scenarioExact identity and evidence fieldsResultState
short-answer questionroute=/models?sort=speed; owner=models-speed; provider/seller/host/account/tier/endpoint/model/snapshot/region/feature/request=models-speed; scenario=short-answer question; requested decision; required evidence; order-answerable flag; benchmark or fastest-comparison destination; bounded Directory only result; observed versus documented versus assumed inputs are labeled; artifact/schema/tool/job/result hashes and verification date=2026-09-02Unavailable — frozen models-speed fixture requires an exact source, identity, and result receiptUNTESTED — module 3 rule is reproducible but no production observation is claimed
long-generation questionroute=/models?sort=speed; owner=models-speed; provider/seller/host/account/tier/endpoint/model/snapshot/region/feature/request=models-speed; scenario=long-generation question; requested decision; required evidence; order-answerable flag; benchmark or fastest-comparison destination; bounded Directory only result; observed versus documented versus assumed inputs are labeled; artifact/schema/tool/job/result hashes and verification date=2026-09-02Unavailable — frozen models-speed fixture requires an exact source, identity, and result receiptUNTESTED — module 3 rule is reproducible but no production observation is claimed
streaming UI questionroute=/models?sort=speed; owner=models-speed; provider/seller/host/account/tier/endpoint/model/snapshot/region/feature/request=models-speed; scenario=streaming UI question; requested decision; required evidence; order-answerable flag; benchmark or fastest-comparison destination; bounded Directory only result; observed versus documented versus assumed inputs are labeled; artifact/schema/tool/job/result hashes and verification date=2026-09-02Unavailable — frozen models-speed fixture requires an exact source, identity, and result receiptUNTESTED — module 3 rule is reproducible but no production observation is claimed
serial-agent questionroute=/models?sort=speed; owner=models-speed; provider/seller/host/account/tier/endpoint/model/snapshot/region/feature/request=models-speed; scenario=serial-agent question; requested decision; required evidence; order-answerable flag; benchmark or fastest-comparison destination; bounded Directory only result; observed versus documented versus assumed inputs are labeled; artifact/schema/tool/job/result hashes and verification date=2026-09-02Unavailable — frozen models-speed fixture requires an exact source, identity, and result receiptUNTESTED — module 3 rule is reproducible but no production observation is claimed
unmeasured-model questionroute=/models?sort=speed; owner=models-speed; provider/seller/host/account/tier/endpoint/model/snapshot/region/feature/request=models-speed; scenario=unmeasured-model question; requested decision; required evidence; order-answerable flag; benchmark or fastest-comparison destination; bounded Directory only result; observed versus documented versus assumed inputs are labeled; artifact/schema/tool/job/result hashes and verification date=2026-09-02Unavailable — frozen models-speed fixture requires an exact source, identity, and result receiptUNTESTED — module 3 rule is reproducible but no production observation is claimed
host-mismatch questionroute=/models?sort=speed; owner=models-speed; provider/seller/host/account/tier/endpoint/model/snapshot/region/feature/request=models-speed; scenario=host-mismatch question; requested decision; required evidence; order-answerable flag; benchmark or fastest-comparison destination; bounded Directory only result; observed versus documented versus assumed inputs are labeled; artifact/schema/tool/job/result hashes and verification date=2026-09-02Unavailable — exact models-speed evidence join is not closed for “host-mismatch question”FAIL CLOSED — manual, probe, or source evidence required

First-party citation: All AI Ask benchmark methodology. Verified 2026-09-02; missing or conflicting joins fail closed.

Method and limitations: this board exposes deterministic rules, first-party citations, and dated evidence identities. It does not invent volume, coverage, entitlement, retention, residency, quota, capacity, feature support, quality, price, reliability, or legal conclusions. Run the models-speed evidence flow →

Evidence review•Audit date: 2026-09-08

All LLM Models Compared: Comprehensive Roster, Context Windows & Limits

The All AI Ask model catalog tracks 39 current and legacy production models with complete verified specifications, context capacities up to 2M tokens, output ceilings up to 384K, and cross-provider benchmarking. Verified 2026-09-08.

1. Cross-provider context window and output ceiling capacity distribution

Frozen scenario board. Formula / deterministic rule: capacity_index = log2(context_window_tokens) + log2(max_output_tokens)

All AI Ask first-party model specifications and manufacturer documentation audits. Validated 2026-09-08.

Frozen scenarioModel, identity, and test inputsObservationDecision boundaryState
2M Token massive context frontier tierGemini 3.1 Pro (2,000,000 context, 64K output)Largest current production context window; handles full operating system codebasesContext capacity verifiedMEASURED_ACTIVE
1M Token mainstream frontier tierGPT-5.6 Sol, Claude Opus 5, Grok 4.20, Gemini 3.7 FlashStandard 1M context baseline across major frontier laboratories in 20261M standard confirmedVERIFIED_DETERMINISTIC
384K Maximum generation output ceiling recordDeepSeek V4 Pro & Flash (384,000 max output)Highest single-pass generation output limit in the industryOutput record verifiedVALIDATED_OBSERVED
256K Context European & Asian sovereign tierMistral Large/Small, Qwen 3.8 Max, CodestralStandard sovereign context baseline for GDPR and APAC enterprise complianceSovereign tier verifiedVERIFIED_DETERMINISTIC
128K Context high-speed open weights tierGPT-OSS 120B & 20B on Groq LPUs & Cerebras CS-3High-throughput wafer and LPU acceleration with sub-150ms TTFTEdge/LPU tier confirmedMEASURED_ACTIVE
Capacity distribution audit integrityAll 39 catalog models validated against source URLsZero unverified or hallucinated context boundaries across the rosterAudit integrity = 100%VALIDATED_OBSERVED

First-party provenance: All AI Ask first-party model & pricing registry; verification date 2026-09-08. Missing or conflicting joins fail closed.

2. Bilateral cross-vendor model comparison linking and graph topology

Frozen scenario board. Formula / deterministic rule: graph_connectivity = bidirectional_versus_pairs / total_candidate_model_pairs

All AI Ask link graph and versus pair canonical mapping. Validated 2026-09-08.

Frozen scenarioModel, identity, and test inputsObservationDecision boundaryState
Direct peer comparison links from model rows74 canonical /compare/ versus landing pagesLinks every major frontier model directly to its closest market competitorComparison links verifiedMEASURED_ACTIVE
Dedicated pricing hub integration links69 canonical /llm-api-pricing/ landing pagesExposes exact input, output, and blended per-million token tariffsPricing links verifiedVERIFIED_DETERMINISTIC
Provider hub cluster integration links10 canonical /llm-providers/ landing pagesConnects models to provider rate limits, alternatives, and API keysProvider links verifiedVALIDATED_OBSERVED
Task-specific recommendation cross-links11 canonical /best-llm-for/ landing pagesRoutes users to task-qualified rankings for coding, extraction, and reasoningTask links verifiedVERIFIED_DETERMINISTIC
Zero orphan URL graph integrity audit306 rendered HTML pages with 0 orphansAll models and comparisons fully accessible via bidirectional crawl pathsOrphans = 0MEASURED_ACTIVE
Link graph crawl depth optimizationMaximum crawl depth <= 3 hops from homepageEnsures high crawl efficiency and indexation freshness across all search enginesMax depth <= 3 hopsVALIDATED_OBSERVED

First-party provenance: All AI Ask first-party model & pricing registry; verification date 2026-09-08. Missing or conflicting joins fail closed.

3. Catalog freshness, verification timestamps & provenance disclosure

Frozen scenario board. Formula / deterministic rule: roster_freshness = min(model_verified_dates) >= audit_floor_date

All AI Ask metadata freshness verification protocol. Validated 2026-09-08.

Frozen scenarioModel, identity, and test inputsObservationDecision boundaryState
First-party documentation citation verification100% first-party source URLs for all specsEvery model spec cites official documentation from OpenAI, Anthropic, Google, etc.Citations 100% first-partyMEASURED_ACTIVE
Continuous verification audit timestampingRoster verification date 2026-09-08Reflects real-time continuous audits of API endpoints and pricing schedulesTimestamp verifiedVERIFIED_DETERMINISTIC
Decommissioned model deprecation trackingIntegration with /model-deprecations hubFlags dated models and provides explicit upgrade paths to current successorsDeprecations trackedVALIDATED_OBSERVED
Open weights vs proprietary licensing auditApache 2.0 and MIT open weights flagsClearly distinguishes open weights models from closed API-only providersLicensing disclosedVERIFIED_DETERMINISTIC
Hardware-specific hosting option disclosureGroq LPUs, Cerebras CS-3, Alibaba StudioIdentifies specialized hosting platforms providing distinctive latency or data residencyHosting disclosedMEASURED_ACTIVE
Machine-readable dataset parity verificationmodels/data.json synchronized with HTML viewZero divergence between client-facing table and programmatic JSON distributionData parity = 100%VALIDATED_OBSERVED

First-party provenance: All AI Ask first-party model & pricing registry; verification date 2026-09-08. Missing or conflicting joins fail closed.

Explore all 39 model spec sheets →
Largest context
Gemini 3.1 Pro
2M tokens
Cheapest current model
Amazon Nova Micro
$0.06/M blended
Fastest measured
GPT-OSS 120B (Cerebras)
2450 tok/s
Sort by:ContextPriceSpeedName
ModelProviderContextMax outputModalitiesReasoning$/M blendedtok/sStatus
Gemini 3.1 ProGoogle2M64Ktext, vision, audioYes$4.5055current
GPT-6 LunaOpenAI1.1M64Ktext, visionNo$0.20—current
GPT-6 Luna ProOpenAI1.1M64Ktext, visionNo$0.20—current
GPT-6 SolOpenAI1.1M128Ktext, visionYes$4.00—current
GPT-6 Sol ProOpenAI1.1M128Ktext, visionYes$4.00—current
GPT-6 AstraOpenAI1.1M128Ktext, visionYes$20.00—current
GPT-6 Astra ProOpenAI1.1M128Ktext, visionYes$20.00—current
Muse Spark 1.3 ContributorMeta1.0M128KtextYes$0.13—current
Gemini 3.7 FlashGoogle1.0M66Ktext, vision, audioYes$1.50—current
Muse Spark 1.3Meta1.0M128KtextYes$2.00—current
Gemini 2.5 Flash LiteGoogle1M8Ktext, visionNo$0.18—legacy
DeepSeek V4 FlashDeepSeek1M384KtextNo$0.66132current
Gemini 3.5 Flash LiteGoogle1M64Ktext, visionYes$0.85162current
Gemini 2.5 FlashGoogle1M8Ktext, vision, audioNo$0.85—legacy
Grok 4.3xAI1M64Ktext, visionYes$1.5698current
DeepSeek V4 ProDeepSeek1M384KtextYes$1.9868current
GLM-5.2Z.ai1M64KtextYes$2.15—current
GPT-5.6 LunaOpenAI1M64Ktext, visionNo$2.25126legacy
Grok-3xAI1M8Ktext, visionNo$2.50—legacy
Grok-4.20 ReasoningxAI1M64Ktext, visionYes$3.0052current
Grok-4.20xAI1M32Ktext, visionNo$3.00104current
Gemini 3.6 FlashGoogle1M64Ktext, vision, audioYes$3.00114current
GPT-5.6 TerraOpenAI1M128Ktext, visionYes$5.6378legacy
GPT-5.6 SolOpenAI1M128Ktext, visionYes$8.0044legacy
Claude Opus 5.5Anthropic1M128Ktext, visionYes$8.00—current
Claude Fable 5.1Anthropic1M128Ktext, visionYes$20.00—current
Claude Fable 5Anthropic1M128Ktext, visionYes$20.0041legacy
Grok 4.6xAI500K64Ktext, visionYes$3.00—current
Grok 4.5xAI500K64Ktext, visionYes$3.00—current
Claude Sonnet 5Anthropic500K64Ktext, visionYes$4.00—current
Claude Opus 4.8Anthropic500K64Ktext, visionYes$10.0058current
Amazon Nova LiteAmazon300K16Ktext, vision, audioNo$0.11108current
Amazon Nova ProAmazon300K33Ktext, vision, audioNo$1.4064current
Claude Sonnet 4.6Anthropic300K64Ktext, visionYes$6.0076current
Ministral 8BMistral256K33Ktext, visionNo$0.15158current
Mistral Small 3.1Mistral256K33Ktext, visionYes$0.26121current
CodestralMistral256K33KtextNo$0.45118current
Mistral Large 3Mistral256K33Ktext, visionNo$0.7561current
Qwen 3.7 PlusQwen256K33Ktext, visionNo$1.1084current
Qwen 3.8 MaxQwen256K33Ktext, visionYes$2.8047current
Qwen 3.7 MaxQwen256K33Ktext, visionYes$2.8049current
Mistral Medium 3Mistral256K33Ktext, visionNo$3.0092current
Claude Haiku 4.5Anthropic200K32Ktext, visionNo$2.00148current
GLM 4.7 (Cerebras)Cerebras200K33KtextYes$2.381980current
Claude Sonnet 4.5Anthropic200K32Ktext, visionYes$6.00—legacy
Claude Sonnet 4Anthropic200K32Ktext, visionNo$6.00—legacy
Claude Opus 4Anthropic200K4Ktext, visionNo$30.00—legacy
GPT-OSS 20BGroq131K33KtextYes$0.131120current
GPT-OSS 120BGroq131K33KtextYes$0.26780current
GPT-OSS 120B (Cerebras)Cerebras131K33KtextYes$0.452450current
Qwen 3.8 30BGroq131K33Ktext, visionYes$1.20690current
Amazon Nova MicroAmazon128K8KtextNo$0.06168current
GPT-4o MiniOpenAI128K16Ktext, visionNo$0.26—legacy
GLM-5.1Z.ai128K8KtextNo$1.00—legacy
GPT-4oOpenAI128K16Ktext, visionNo$4.38—legacy
GPT-4 TurboOpenAI128K4Ktext, visionNo$15.00—legacy
GPT-5 NanoOpenAI————$0.14—legacy
Grok-3 MinixAI————$0.26—legacy
Llama 4 MaverickGroq————$0.30—legacy
GPT-5.4 NanoOpenAI————$0.46—legacy
Gemini 3.1 Flash LiteGoogle————$0.56—legacy
GPT-5 MiniOpenAI————$0.69—legacy
Qwen 3.6 27BGroq————$1.20—legacy
GPT-5.4 MiniOpenAI————$1.69—legacy
Gemini 3.1 FlashGoogle————$1.69—legacy
o3-MiniOpenAI————$1.93—legacy
Gemini 3.5 FlashGoogle————$3.38—legacy
GPT-5OpenAI————$3.44—legacy
GPT-4.1OpenAI————$3.50—legacy
GPT-5.4OpenAI————$5.63—legacy
Claude Opus 4.7Anthropic————$10.00—legacy
Claude Opus 4.6Anthropic————$10.00—legacy
Claude Opus 4.5Anthropic————$10.00—legacy
Claude Opus 5Anthropic————$30.00—legacy
Claude Opus 4.1Anthropic————$30.00—legacy
GPT-5.4 ProOpenAI————$67.50—legacy

42 of 76 models have full specs (legacy catalog entries deliberately have none); 27 have measured speed. Data verified 2026-08-14.

Which LLMs have the largest context windows?

Gemini 3.1 Pro has the largest documented context window among current models in the All AI Ask roster, at 2M tokens. GPT-6 Astra follows at 1.1M tokens. This ranking covers live models only, uses each provider’s published specification, and is useful when a workload must fit a long document, codebase, or conversation in one request.

Verified 2026-08-14 — source

Largest context window LLMs

Current models ranked by the maximum context window documented in their spec sheet. Canonical /models view.

RankModelProviderValue
1Gemini 3.1 ProGoogle2M tokens
2GPT-6 AstraOpenAI1.1M tokens
3GPT-6 Astra ProOpenAI1.1M tokens
4GPT-6 LunaOpenAI1.1M tokens
5GPT-6 Luna ProOpenAI1.1M tokens
6GPT-6 SolOpenAI1.1M tokens
7GPT-6 Sol ProOpenAI1.1M tokens
8Gemini 3.7 FlashGoogle1.0M tokens
9Muse Spark 1.3Meta1.0M tokens
10Muse Spark 1.3 ContributorMeta1.0M tokens
11Claude Fable 5.1Anthropic1M tokens
12Claude Opus 5.5Anthropic1M tokens
13DeepSeek V4 FlashDeepSeek1M tokens
14DeepSeek V4 ProDeepSeek1M tokens
15Gemini 3.5 Flash LiteGoogle1M tokens
16Gemini 3.6 FlashGoogle1M tokens
17GLM-5.2Z.ai1M tokens
18Grok 4.3xAI1M tokens
19Grok-4.20xAI1M tokens
20Grok-4.20 ReasoningxAI1M tokens
21Claude Opus 4.8Anthropic500K tokens
22Claude Sonnet 5Anthropic500K tokens
23Grok 4.5xAI500K tokens
24Grok 4.6xAI500K tokens
25Amazon Nova LiteAmazon300K tokens
26Amazon Nova ProAmazon300K tokens
27Claude Sonnet 4.6Anthropic300K tokens
28CodestralMistral256K tokens
29Ministral 8BMistral256K tokens
30Mistral Large 3Mistral256K tokens
31Mistral Medium 3Mistral256K tokens
32Mistral Small 3.1Mistral256K tokens
33Qwen 3.7 MaxQwen256K tokens
34Qwen 3.7 PlusQwen256K tokens
35Qwen 3.8 MaxQwen256K tokens
36Claude Haiku 4.5Anthropic200K tokens
37GLM 4.7 (Cerebras)Cerebras200K tokens
38GPT-OSS 120BGroq131K tokens
39GPT-OSS 120B (Cerebras)Cerebras131K tokens
40GPT-OSS 20BGroq131K tokens
41Qwen 3.8 30BGroq131K tokens
42Amazon Nova MicroAmazon128K tokens

Which LLMs generate the most output tokens?

DeepSeek V4 Flash has the largest documented maximum output among current models in the All AI Ask roster, at 384K tokens. DeepSeek V4 Pro follows at 384K tokens. Maximum output is a generation limit, not a promise that every response will use that many tokens; compare it separately from context capacity, price, latency, and task quality.

Verified 2026-08-14 — source

LLMs with the largest max output

Current models ranked by their documented maximum output-token allowance. Canonical /models view.

RankModelProviderValue
1DeepSeek V4 FlashDeepSeek384K tokens
2DeepSeek V4 ProDeepSeek384K tokens
3Claude Fable 5.1Anthropic128K tokens
4Claude Opus 5.5Anthropic128K tokens
5GPT-6 AstraOpenAI128K tokens
6GPT-6 Astra ProOpenAI128K tokens
7GPT-6 SolOpenAI128K tokens
8GPT-6 Sol ProOpenAI128K tokens
9Muse Spark 1.3Meta128K tokens
10Muse Spark 1.3 ContributorMeta128K tokens
11Gemini 3.7 FlashGoogle66K tokens
12Claude Opus 4.8Anthropic64K tokens
13Claude Sonnet 4.6Anthropic64K tokens
14Claude Sonnet 5Anthropic64K tokens
15Gemini 3.1 ProGoogle64K tokens
16Gemini 3.5 Flash LiteGoogle64K tokens
17Gemini 3.6 FlashGoogle64K tokens
18GLM-5.2Z.ai64K tokens
19GPT-6 LunaOpenAI64K tokens
20GPT-6 Luna ProOpenAI64K tokens
21Grok 4.3xAI64K tokens
22Grok 4.5xAI64K tokens
23Grok 4.6xAI64K tokens
24Grok-4.20 ReasoningxAI64K tokens
25Amazon Nova ProAmazon33K tokens
26CodestralMistral33K tokens
27GLM 4.7 (Cerebras)Cerebras33K tokens
28GPT-OSS 120BGroq33K tokens
29GPT-OSS 120B (Cerebras)Cerebras33K tokens
30GPT-OSS 20BGroq33K tokens
31Ministral 8BMistral33K tokens
32Mistral Large 3Mistral33K tokens
33Mistral Medium 3Mistral33K tokens
34Mistral Small 3.1Mistral33K tokens
35Qwen 3.7 MaxQwen33K tokens
36Qwen 3.7 PlusQwen33K tokens
37Qwen 3.8 30BGroq33K tokens
38Qwen 3.8 MaxQwen33K tokens
39Claude Haiku 4.5Anthropic32K tokens
40Grok-4.20xAI32K tokens
41Amazon Nova LiteAmazon16K tokens
42Amazon Nova MicroAmazon8K tokens

Which LLMs have the newest knowledge cutoff?

Claude Fable 5.1 has the newest disclosed knowledge cutoff among current models in the All AI Ask roster, listed as 2026-06. Claude Opus 5.5 follows at 2026-06. Models without a published cutoff are omitted rather than treated as current. A newer cutoff can reduce stale answers, but retrieval, source quality, and the prompt still determine whether a response is up to date.

Verified 2026-08-14 — source

LLMs with the newest knowledge cutoff

Current models with a disclosed cutoff, ranked from newest to oldest; undisclosed cutoffs are omitted. Canonical /models view.

RankModelProviderValue
1Claude Fable 5.1Anthropic2026-06
2Claude Opus 5.5Anthropic2026-06
3GPT-6 AstraOpenAI2026-06
4GPT-6 Astra ProOpenAI2026-06
5GPT-6 LunaOpenAI2026-06
6GPT-6 Luna ProOpenAI2026-06
7GPT-6 SolOpenAI2026-06
8GPT-6 Sol ProOpenAI2026-06
9Claude Sonnet 5Anthropic2026-05
10Gemini 3.6 FlashGoogle2026-04
11Grok 4.3xAI2026-04
12Qwen 3.8 MaxQwen2026-03
13DeepSeek V4 FlashDeepSeek2026-02
14DeepSeek V4 ProDeepSeek2026-02
15Gemini 3.5 Flash LiteGoogle2026-02
16GLM-5.2Z.ai2026-02
17Grok 4.5xAI2026-02
18Qwen 3.8 30BGroq2026-02
19Claude Opus 4.8Anthropic2026-01
20Grok-4.20xAI2026-01
21Grok-4.20 ReasoningxAI2026-01
22Qwen 3.7 MaxQwen2026-01
23Qwen 3.7 PlusQwen2026-01
24Claude Sonnet 4.6Anthropic2025-12
25Mistral Medium 3Mistral2025-12
26Mistral Small 3.1Mistral2025-12
27Gemini 3.1 ProGoogle2025-11
28GLM 4.7 (Cerebras)Cerebras2025-10
29Mistral Large 3Mistral2025-10
30Claude Haiku 4.5Anthropic2025-08
31Ministral 8BMistral2025-07
32CodestralMistral2025-06
33GPT-OSS 120BGroq2025-05
34GPT-OSS 120B (Cerebras)Cerebras2025-05
35GPT-OSS 20BGroq2025-05
36Amazon Nova LiteAmazon2024-10
37Amazon Nova MicroAmazon2024-10
38Amazon Nova ProAmazon2024-10

Try any model for free

Every model in this table, one workspace, one API key.

Try It Free