How Much Does an LLM Chatbot Cost per Month?

At production volume (200,000 calls/month), the cheapest effective option is Amazon Nova Micro at $24.25/month. The most expensive frontier option, GPT-5.4 Pro, runs $27,504/month — A support or product chatbot re-sends a growing conversation history on every turn, so the input side of the bill grows even though each individual reply stays short.

How much does llm chatbot cost per month?

At production volume (200,000 calls/month), the cheapest effective option for llm chatbot is Amazon Nova Micro at $24.25 per month, verbosity-adjusted rather than list price. The most expensive frontier model, GPT-5.4 Pro, runs $27,504 per month for the same workload.

Verified 2026-08-08
Cheapest ≠ best. This page ranks by effective cost only. See our graded results for this task on Best LLM For →

Token shape

Shapegrowing conversation history in, short reply out
Input / output tokens per call2K in / 0K out
Cacheable input35%
Batch-eligibleNo

Input is a system prompt plus roughly 8-10 turns of history by the middle of a conversation; output is a typical single-paragraph reply. Only the system prompt is stable enough to cache.

What drives this workload's cost?

The main token-volume driver here is input tokens: each call sends 2K input tokens and requests up to 0K output tokens, at 200,000 calls per month in the default volume. 35% of input is modeled as cache-eligible, so repeated prefixes can reduce the input charge. It is not marked batch-eligible because the profile assumes a synchronous response.

Volume

Side project
20,000 calls/mo
$2.42/mo cheapest
Production
200,000 calls/mo
$24.25/mo cheapest
Scale
2,000,000 calls/mo
$242/mo cheapest

Ranked cost — Production volume

ModelProviderList monthlyEffective monthlyVerbosityRank Δ
Amazon Nova MicrobudgetAmazon$26.60$24.250.76×—
Amazon Nova LitebudgetAmazon$45.60$44.090.91×—
GPT-5 NanobudgetlegacyOpenAI$52.00$44.51——
Gemini 2.5 Flash LitebudgetlegacyGoogle$76.00$60.93—▲2
Ministral 8BbudgetMistral$82.50$81.770.93×▲2
GPT-6 LunabudgetOpenAI$83.00$68.02—▲2
GPT-6 Luna ProbudgetOpenAI$83.00$68.02—▲2
GPT-OSS 20BbudgetGroq$57.00$97.532.93×▼4
Mistral Small 3.1budgetMistral$114$1080.85×▲4
GPT-4o MinibudgetlegacyOpenAI$114$91.53——
Grok-3 MinibudgetlegacyxAI$114$114——
Llama 4 MaverickbudgetlegacyGroq$138$138—▲2
GPT-OSS 120BbudgetGroq$114$1481.82×▼1
GPT-5.4 NanobudgetlegacyOpenAI$184$1340.78×▲1
CodestralbudgetMistral$207$1940.79×▲1
Show all 76 models
Gemini 3.1 Flash LitebudgetlegacyGoogle$225$1740.87×▲2
Muse Spark 1.3 ContributorbudgetMeta$62.00$22912.95×▼12
GPT-5 MinibudgetlegacyOpenAI$260$223—▲1
GPT-OSS 120B (Cerebras)budgetCerebras$221$2902.32×▼2
Gemini 3.5 Flash LitebudgetGoogle$319$274—▲1
Gemini 2.5 FlashbudgetlegacyGoogle$319$274—▲1
Mistral Large 3budgetMistral$345$345—▲1
GLM-5.1midlegacyZ.ai$442$442—▲1
DeepSeek V4 FlashbudgetDeepSeek$304$4512.59×▼4
Qwen 3.8 30BmidGroq$498$498——
Qwen 3.6 27BmidlegacyGroq$498$498——
Qwen 3.7 PlusmidQwen$524$524——
Amazon Nova PromidAmazon$608$608——
Gemini 3.7 FlashmidGoogle$623$510——
GPT-5.4 MinimidlegacyOpenAI$675$563——
Gemini 3.1 FlashmidlegacyGoogle$675$562——
Claude Haiku 4.5midAnthropic$830$687—▲1
o3-MinimidlegacyOpenAI$836$671—▲1
Grok 4.3midxAI$775$8471.41×▼2
Muse Spark 1.3midMeta$898$898——
GPT-5.6 LunamidlegacyOpenAI$900$750——
Mistral Medium 3midMistral$1,245$1,1610.84×▲6
Qwen 3.8 MaxmidQwen$1,216$1,216—▲1
Qwen 3.7 MaxmidQwen$1,216$1,216—▲1
Grok-3midlegacyxAI$1,240$1,240—▲1
Gemini 3.6 FlashmidGoogle$1,245$1,019—▲1
GPT-5midlegacyOpenAI$1,300$1,113—▲3
Gemini 3.5 FlashmidlegacyGoogle$1,350$1,124—▲3
Grok-4.20 ReasoningmidxAI$1,380$1,380—▲3
Grok-4.20midxAI$1,380$1,380—▲3
Grok 4.6midxAI$1,380$1,380—▲3
Grok 4.5midxAI$1,380$1,380—▲3
Gemini 3.1 PromidGoogle$1,800$1,1880.63×▲7
GPT-4.1midlegacyOpenAI$1,520$1,220—▲2
DeepSeek V4 PromidDeepSeek$911$1,5483.30×▼13
GPT-6 SolmidOpenAI$1,660$1,360—▲1
GPT-6 Sol PromidOpenAI$1,660$1,360—▲1
Claude Sonnet 5midAnthropic$1,660$1,374—▲1
GLM-5.2midZ.ai$980$1,8643.87×▼16
GPT-4omidlegacyOpenAI$1,900$1,525—▲1
GPT-5.6 TerramidlegacyOpenAI$2,250$1,875—▲1
GPT-5.4midlegacyOpenAI$2,250$1,875—▲1
Claude Sonnet 4.6midAnthropic$2,490$2,061—▲1
Claude Sonnet 4.5midlegacyAnthropic$2,490$2,061—▲1
Claude Sonnet 4midlegacyAnthropic$2,490$2,061—▲1
GLM 4.7 (Cerebras)midCerebras$1,273$2,5307.53×▼17
GPT-5.6 SolmidlegacyOpenAI$3,320$2,721——
Claude Opus 5.5midAnthropic$3,320$2,749——
Claude Opus 4.8midAnthropic$4,150$3,3660.96×—
Claude Opus 4.7midlegacyAnthropic$4,150$3,436——
Claude Opus 4.6midlegacyAnthropic$4,150$3,436——
Claude Opus 4.5midlegacyAnthropic$4,150$3,436——
GPT-4 TurbofrontierlegacyOpenAI$6,900$5,402——
GPT-6 AstrafrontierOpenAI$8,300$6,802——
GPT-6 Astra ProfrontierOpenAI$8,300$6,802——
Claude Fable 5.1frontierAnthropic$8,300$6,871——
Claude Fable 5frontierlegacyAnthropic$8,300$6,871——
Claude Opus 5frontierlegacyAnthropic$12,450$10,307——
Claude Opus 4.1frontierlegacyAnthropic$12,450$10,307——
Claude Opus 4frontierlegacyAnthropic$12,450$10,307——
GPT-5.4 ProfrontierlegacyOpenAI$27,000$23,0101.04×—

Decision evidence · verified 2026-08-27

Chatbot conversation cost evidence

Frozen inputs, formulas, provenance, and decision boundaries are visible in the initial HTML. Unsupported evidence fails closed as Unavailable.

Conversation-cohort history ledger

Formula / scoring rule: Conversation bill = Σ(turn input history × input rate + turn output × output rate); history policy changes the denominator.

Provenance: Frozen 30K monthly conversations, 8 turns, 1,200 initial input, 220 output/turn; unresolved chats are retained.

Frozen fixture / runVisible inputsField-level resultDecision boundaryState
full history8 turns; 1,200 initial; 220 output/turn; history retainedApprox input 9,680; output 1,760; bill = (9,680×5 + 1,760×15)/1M = $0.0748/conversation.Do not multiply the first turn by eight.CALCULATED — growing history.
summary after turn 4turn 4 summary 600 tokens; turns 5–8 use summary + recent 2 turnsEstimated input 6,480; output 1,760; bill = $0.0588/conversation.Summary quality and summary-call cost must be added when measured.CALCULATED — summary path.
unresolved cohort30K conversations; 12% unresolved; 2 extra turns; resolution log absentExtra turns are counted; resolved denominator Unavailable — conversation outcome log is absentDo not report cost per resolved conversation as cost per chat.Unavailable — conversation outcome log is absent

Module citation: All AI Ask evidence registry (verified 2026-08-27).

Turn-policy crossover surface

Formula / scoring rule: Policy cost = model turns + summary/compaction calls + retries; crossover is first policy with lower cost at the same resolution floor.

Provenance: Matched 4/8/12-turn scenarios; compaction is modeled as a separate call, not free state.

Frozen fixture / runVisible inputsField-level resultDecision boundaryState
4-turn cap4 turns; 1,200 input; 220 output; stop reason=capBill subtotal = $0.0408; unresolved rate Unavailable — observed resolution is absentA lower cap cannot claim equivalent service.Unavailable — observed resolution is absent
8-turn + summary8 turns; summary call at turn 4; 600 summary tokensModel bill calculable; summary acceptance Unavailable — not graded in fixtureSummary overhead remains in numerator.Unavailable — not graded in fixture
12-turn overflow12 turns; 3 retries; duplicate user messages 2%Duplicate and retry tokens counted; exact bill Unavailable — retry usage export is absentDo not hide overflow in average turn count.Unavailable — retry usage export is absent

Module citation: All AI Ask chatbot workload registry.

Resolved-conversation TCO tree

Formula / scoring rule: TCO/resolved = (model + moderation + retrieval + human escalation) / resolved conversations; unsupported components remain Unavailable.

Provenance: 30K conversation baseline; escalation minutes/rate are user-supplied; chatbot quality is out of scope.

Frozen fixture / runVisible inputsField-level resultDecision boundaryState
model subtotal30K × $0.0748 full-history conversationModel subtotal = 30,000 × $0.0748 = $2,244.This is token cost, not resolved-conversation TCO.CALCULATED — token component.
retrieval / moderationretrieval calls 20%; moderation endpoint and vector feeAdditional fees Unavailable — provider/tool tariffs are not suppliedDo not treat ancillary calls as included for free.Unavailable — provider/tool tariffs are not supplied
human escalation6% escalated; 4 minutes; $60/hour; resolution denominator absentScenario labor = 1,800 × 4/60 × $60 = $7,200.Publish per-resolved only after outcome counts close.Unavailable — resolution denominator absent

Module citation: All AI Ask chatbot calculator.

Calculate cost per resolved conversation →

Prompt caching uses the provider's sourced read multiplier and amortises cache writes over the documented TTL; the largest published read discount is 90%. See the full formula and coverage disclosure →

Related

Alternatives to Amazon Nova Micro →Amazon cost calculator →RAG Question Answering cost →Coding Agent cost →

FAQ

Why does conversation history dominate the cost?
Most chat APIs are stateless — the full history is re-sent on every turn, so a long conversation costs more per reply even though the model only generates one short answer at a time.
Does prompt caching help chatbot costs?
Only for the stable part of the prompt (system instructions). The conversation history itself changes every turn, so it is rarely cache-eligible.