LLM Cost Calculator — Real Monthly Cost, Not Just Rate

Raw dataset: data.json. Verbosity measured 2026-06-21. Cite this: All AI Ask LLM Verbosity Index and Effective Cost Dataset, retrieved 2026-06-21.

A $/M rate is not your bill. Set your call shape below and see every priced model ranked by effective monthly cost — adjusted for how many output tokens each model actually spends on a job of this size, not its list price alone.

ModelProviderList monthlyEffective monthlyVerbosityRank Δ
Amazon Nova MicrobudgetAmazon$26.60$24.250.76×—
Amazon Nova LitebudgetAmazon$45.60$44.090.91×—
GPT-5 NanobudgetlegacyOpenAI$52.00$52.00——
Gemini 2.5 Flash LitebudgetlegacyGoogle$76.00$76.00—▲2
Ministral 8BbudgetMistral$82.50$81.770.93×▲2
GPT-6 LunabudgetOpenAI$83.00$83.00—▲2
GPT-6 Luna ProbudgetOpenAI$83.00$83.00—▲2
GPT-OSS 20BbudgetGroq$57.00$97.532.93×▼4
Mistral Small 3.1budgetMistral$114$1080.85×▲4
GPT-4o MinibudgetlegacyOpenAI$114$114——
Grok-3 MinibudgetlegacyxAI$114$114——
Llama 4 MaverickbudgetlegacyGroq$138$138—▲2
GPT-OSS 120BbudgetGroq$114$1481.82×▼1
GPT-5.4 NanobudgetlegacyOpenAI$184$1640.78×▲1
CodestralbudgetMistral$207$1940.79×▲1
Show all 76 models
Gemini 3.1 Flash LitebudgetlegacyGoogle$225$2110.87×▲2
Muse Spark 1.3 ContributorbudgetMeta$62.00$22912.95×▼12
GPT-5 MinibudgetlegacyOpenAI$260$260—▲1
GPT-OSS 120B (Cerebras)budgetCerebras$221$2902.32×▼2
Gemini 3.5 Flash LitebudgetGoogle$319$319—▲1
Gemini 2.5 FlashbudgetlegacyGoogle$319$319—▲1
Mistral Large 3budgetMistral$345$345—▲1
GLM-5.1midlegacyZ.ai$442$442—▲1
DeepSeek V4 FlashbudgetDeepSeek$304$4512.59×▼4
Qwen 3.8 30BmidGroq$498$498——
Qwen 3.6 27BmidlegacyGroq$498$498——
Qwen 3.7 PlusmidQwen$524$524——
Amazon Nova PromidAmazon$608$608——
Gemini 3.7 FlashmidGoogle$623$623——
GPT-5.4 MinimidlegacyOpenAI$675$675——
Gemini 3.1 FlashmidlegacyGoogle$675$675——
Claude Haiku 4.5midAnthropic$830$830—▲1
o3-MinimidlegacyOpenAI$836$836—▲1
Grok 4.3midxAI$775$8471.41×▼2
Muse Spark 1.3midMeta$898$898——
GPT-5.6 LunamidlegacyOpenAI$900$900——
Mistral Medium 3midMistral$1,245$1,1610.84×▲6
Qwen 3.8 MaxmidQwen$1,216$1,216—▲1
Qwen 3.7 MaxmidQwen$1,216$1,216—▲1
Grok-3midlegacyxAI$1,240$1,240—▲1
Gemini 3.6 FlashmidGoogle$1,245$1,245—▲1
GPT-5midlegacyOpenAI$1,300$1,300—▲3
Gemini 3.5 FlashmidlegacyGoogle$1,350$1,350—▲3
Grok-4.20 ReasoningmidxAI$1,380$1,380—▲3
Grok-4.20midxAI$1,380$1,380—▲3
Grok 4.6midxAI$1,380$1,380—▲3
Grok 4.5midxAI$1,380$1,380—▲3
Gemini 3.1 PromidGoogle$1,800$1,4890.63×▲7
GPT-4.1midlegacyOpenAI$1,520$1,520—▲2
DeepSeek V4 PromidDeepSeek$911$1,5483.30×▼13
GPT-6 SolmidOpenAI$1,660$1,660—▲1
GPT-6 Sol PromidOpenAI$1,660$1,660—▲1
Claude Sonnet 5midAnthropic$1,660$1,660—▲1
GLM-5.2midZ.ai$980$1,8643.87×▼16
GPT-4omidlegacyOpenAI$1,900$1,900—▲1
GPT-5.6 TerramidlegacyOpenAI$2,250$2,250—▲1
GPT-5.4midlegacyOpenAI$2,250$2,250—▲1
Claude Sonnet 4.6midAnthropic$2,490$2,490—▲1
Claude Sonnet 4.5midlegacyAnthropic$2,490$2,490—▲1
Claude Sonnet 4midlegacyAnthropic$2,490$2,490—▲1
GLM 4.7 (Cerebras)midCerebras$1,273$2,5307.53×▼17
GPT-5.6 SolmidlegacyOpenAI$3,320$3,320——
Claude Opus 5.5midAnthropic$3,320$3,320——
Claude Opus 4.8midAnthropic$4,150$4,0800.96×—
Claude Opus 4.7midlegacyAnthropic$4,150$4,150——
Claude Opus 4.6midlegacyAnthropic$4,150$4,150——
Claude Opus 4.5midlegacyAnthropic$4,150$4,150——
GPT-4 TurbofrontierlegacyOpenAI$6,900$6,900——
GPT-6 AstrafrontierOpenAI$8,300$8,300——
GPT-6 Astra ProfrontierOpenAI$8,300$8,300——
Claude Fable 5.1frontierAnthropic$8,300$8,300——
Claude Fable 5frontierlegacyAnthropic$8,300$8,300——
Claude Opus 5frontierlegacyAnthropic$12,450$12,450——
Claude Opus 4.1frontierlegacyAnthropic$12,450$12,450——
Claude Opus 4frontierlegacyAnthropic$12,450$12,450——
GPT-5.4 ProfrontierlegacyOpenAI$27,000$27,5041.04×—

List price lied to you

These models move the most once verbosity is priced in — a model that talks more costs more, regardless of its list rate.

  • GLM 4.7 (Cerebras) is priced #44 by list rate but #61 once its 7.53× verbosity is billed — $1,273 list vs $2,530 effective.
  • GLM-5.2 is priced #38 by list rate but #54 once its 3.87× verbosity is billed — $980 list vs $1,864 effective.
  • DeepSeek V4 Pro is priced #37 by list rate but #50 once its 3.30× verbosity is billed — $911 list vs $1,548 effective.
  • Muse Spark 1.3 Contributor is priced #5 by list rate but #17 once its 12.95× verbosity is billed — $62.00 list vs $229 effective.
  • Gemini 3.1 Pro is priced #55 by list rate but #48 once its 0.63× verbosity is billed — $1,800 list vs $1,489 effective.
  • Mistral Medium 3 is priced #43 by list rate but #37 once its 0.84× verbosity is billed — $1,245 list vs $1,161 effective.

Decision evidence · verified 2026-08-27

Whole-workload cost routing and price replay

Every field is tied to a frozen input and a dated provenance record. Unsupported facts fail closed as Unavailable; they are not treated as zero, free, equivalent, current, fastest, cheapest, private, or passing.

Workload-component router

Formula / rubric: Total scope = base token bill + applicable retrieval + embeddings + storage + tools + media + review + retries + parallelism; route to specialist when component exists.

Provenance: Frozen defaults cover chatbot, RAG, coding agent, extraction, summarization, and tool loops. Unsupported fields fail closed as Unavailable.

Frozen fixture / runVisible inputsField-level resultDecision boundaryState / reproducible bill
chatbot
batch39-llm-cost-calculator-m1-r1
30K calls; 1,200 input/220 output; no tools; 10% cachebase token estimate $279.00/month.Use global calculator; no non-token component detected.ROUTED — base-token path.
RAG
batch39-llm-cost-calculator-m1-r2
30K calls; embeddings + vector reads + rerank + generationbase token path closes; full RAG path routes to /llm-cost-calculator/rag-question-answering.Do not reuse chatbot estimate as end-to-end RAG cost.ROUTED — specialist calculator.
coding agent
batch39-llm-cost-calculator-m1-r3
5–30 turns; tools; tests; compaction; parallel workerstoken estimate closes; repository fan-out and repair route to /llm-cost-calculator/coding-agent.Generic token output is a lower-bound component, not task cost.ROUTED — specialist calculator.

Module citations: All AI Ask pricing registry. All AI Ask evidence registry (verified 2026-08-27).

Five-axis break-even matrix

Formula / rubric: winner threshold = first axis value where totalCostA ≤ totalCostB, holding other declared inputs fixed; unavailable mechanic blocks threshold.

Provenance: Axes: output expansion, cache hit, batch share, retry rate, and monthly calls; dated units must match. Unsupported fields fail closed as Unavailable.

Frozen fixture / runVisible inputsField-level resultDecision boundaryState / reproducible bill
output expansion
batch39-llm-cost-calculator-m2-r1
220 → 440 output tokens; input 1,200; 30K callsmodel A crosses model B at 318 output tokens under frozen rates.Threshold is valid only for these model IDs and rates.CALCULATED — axis closed.
cache/batch axes
batch39-llm-cost-calculator-m2-r2
cache 0/50/100%; batch 0/50/100%; same token shapecache-read multiplier for one candidate is Unavailable — not returned in compatible registry unitsNo break-even matrix cell is emitted for that candidate.Unavailable — cache-read multiplier is not returned in compatible registry units
retry/call volume
batch39-llm-cost-calculator-m2-r3
retry 0/5/20%; calls 1K/30K/1M; output 220call-volume scaling is linear; retry debit for candidate C is Unavailable — not returnedRetain a threshold only where retry billing is observed.Unavailable — retry debit is not returned

Module citations: OpenAI API pricing. All AI Ask evidence registry (verified 2026-08-27).

Current versus prior price-change budget ledger

Formula / rubric: Impact = replay(frozen workload, current registry) − replay(frozen workload, prior registry), with catalog/coverage changes separated from rate deltas.

Provenance: Current and immediately prior verified registry snapshots; no live pricing call is made. Unsupported fields fail closed as Unavailable.

Frozen fixture / runVisible inputsField-level resultDecision boundaryState / reproducible bill
small workload
batch39-llm-cost-calculator-m3-r1
1K calls; 2K input/400 output; current $5/$15; prior $5/$15rate delta $0.00; model catalog unchanged.Attribute only rate movement to budget impact.NO CHANGE — replay closes.
production workload
batch39-llm-cost-calculator-m3-r2
30K calls; 12K input/1.2K output; prior output coverage absentcurrent replay exists; prior verbosity-adjusted output is Unavailable — not covered by prior snapshotDo not label coverage change as price change.Unavailable — prior verbosity-adjusted output is not covered
scale workload
batch39-llm-cost-calculator-m3-r3
1M calls; batch 40%; cache 60%; provider TTL mechanic changedcatalog and cache mechanics changed together; attributable delta is Unavailable — not separablePublish no single price-change percentage.Unavailable — catalog and cache mechanic deltas are not separable

Module citations: All AI Ask model pricing registry. All AI Ask evidence registry (verified 2026-08-27).

Calculate a shareable whole-workload estimate →

The formula, published

outputTokensBilled = outputTokens × (verbosityIndex ?? 1)
inputCost  = inputPerM  × inputTokens        / 1e6
outputCost = outputPerM × outputTokensBilled / 1e6
effective  = (inputCost + outputCost) × callsPerMonth

with caching: inputCost × (1 − cacheableInputPct × 0.9)
with batching: (inputCost + outputCost) × (1 − batchDiscountPct/100)

up to 90% cache-read saving is the largest sourced read discount in the provider table; provider-specific estimates use that provider's multiplier and account for cache writes. The free-form estimator assumes 30% of input is cache-eligible; pick a workload preset below for a shape-specific assumption instead.

Coverage: 20 of 76 priced models have a measured verbosity index (2+ graded runs). The rest render with a verbosity of — and are shown at unadjusted list price.

Cost by workload shape

LLM Chatbot
growing conversation history in, short reply out
RAG Question Answering
large retrieved context in, short answer out
Coding Agent
large file context in, large diff out
Document Extraction
medium document in, tiny structured JSON out
Long-Document Summarization
very large document in, medium summary out
Content Generation
tiny brief in, long piece out
Classification at Volume
tiny input in, single-label output out
Agentic Tool Loop
many small round-trips per completed task