How Much Does LLM Classification at Volume Cost per Month?

At production volume (5,000,000 calls/month), the cheapest effective option is Amazon Nova Micro at $98.14/month. The most expensive frontier option, GPT-5.4 Pro, runs $93,720/month — Classification is the smallest per-call shape in this cluster, but volume is the entire story — a tiny per-call cost still adds up at millions of calls a month.

How much does classification at volume cost per month?

At production volume (5,000,000 calls/month), the cheapest effective option for classification at volume is Amazon Nova Micro at $98.14 per month, verbosity-adjusted rather than list price. The most expensive frontier model, GPT-5.4 Pro, runs $93,720 per month for the same workload.

Verified 2026-08-08
Cheapest ≠ best. This page ranks by effective cost only.

Token shape

Shapetiny input in, single-label output out
Input / output tokens per call1K in / 0K out
Cacheable input40%
Batch-eligibleYes

Input is one short text plus a stable instruction/label-set prompt; output is a single label or short code. The instruction portion is highly cacheable, and this is the clearest batch-processing candidate in the cluster.

What drives this workload's cost?

The main token-volume driver here is input tokens: each call sends 1K input tokens and requests up to 0K output tokens, at 5,000,000 calls per month in the default volume. 40% of input is modeled as cache-eligible, so repeated prefixes can reduce the input charge. It is also batch-eligible, which can reduce the modeled total when asynchronous processing is acceptable.

Volume

Side project
500,000 calls/mo
$9.81/mo cheapest
Production
5,000,000 calls/mo
$98.14/mo cheapest
Scale
50,000,000 calls/mo
$981/mo cheapest

Ranked cost — Production volume (caching + batch applied where available)

ModelProviderList monthlyEffective monthlyVerbosityRank Δ
Amazon Nova MicrobudgetAmazon$101$49.070.76×—
GPT-5 NanobudgetlegacyOpenAI$165$120——
Amazon Nova LitebudgetAmazon$174$85.920.91×—
GPT-OSS 20BbudgetGroq$218$1382.93×—
Gemini 2.5 Flash LitebudgetlegacyGoogle$290$200—▲1
GPT-6 LunabudgetOpenAI$300$210—▲1
GPT-6 Luna ProbudgetOpenAI$300$210—▲1
Ministral 8BbudgetMistral$390$1940.93×▲1
Mistral Small 3.1budgetMistral$435$2130.85×▲4
GPT-4o MinibudgetlegacyOpenAI$435$300——
Grok-3 MinibudgetlegacyxAI$435$435——
GPT-OSS 120BbudgetGroq$435$2421.82×—
Muse Spark 1.3 ContributorbudgetMeta$270$50912.95×▼8
Llama 4 MaverickbudgetlegacyGroq$560$280——
GPT-5.4 NanobudgetlegacyOpenAI$625$4180.78×—
Show all 76 models
Gemini 3.1 Flash LitebudgetlegacyGoogle$775$5310.87×—
CodestralbudgetMistral$840$4110.79×▲1
GPT-5 MinibudgetlegacyOpenAI$825$600—▼1
Gemini 3.5 Flash LitebudgetGoogle$1,000$730—▲1
Gemini 2.5 FlashbudgetlegacyGoogle$1,000$730—▲1
GPT-OSS 120B (Cerebras)budgetCerebras$950$1,0492.32×▼2
Mistral Large 3budgetMistral$1,400$700—▲1
DeepSeek V4 FlashbudgetDeepSeek$1,232$1,4422.59×▼1
GLM-5.1midlegacyZ.ai$1,720$1,720——
Qwen 3.8 30BmidGroq$1,800$900——
Qwen 3.6 27BmidlegacyGroq$1,800$900——
Qwen 3.7 PlusmidQwen$2,200$2,200——
Gemini 3.7 FlashmidGoogle$2,250$1,575——
Amazon Nova PromidAmazon$2,320$1,160——
GPT-5.4 MinimidlegacyOpenAI$2,325$1,650——
Gemini 3.1 FlashmidlegacyGoogle$2,325$1,650——
Claude Haiku 4.5midAnthropic$3,000$2,102——
GPT-5.6 LunamidlegacyOpenAI$3,100$2,200——
o3-MinimidlegacyOpenAI$3,190$2,200——
Grok 4.3midxAI$3,375$3,4781.41×—
Muse Spark 1.3midMeta$3,550$3,550——
GPT-5midlegacyOpenAI$4,125$3,000—▲2
Mistral Medium 3midMistral$4,500$2,1900.84×▲3
Gemini 3.6 FlashmidGoogle$4,500$3,150—▲1
DeepSeek V4 PromidDeepSeek$3,696$4,6073.30×▼3
Qwen 3.8 MaxmidQwen$4,640$4,640—▲1
Qwen 3.7 MaxmidQwen$4,640$4,640—▲1
Gemini 3.5 FlashmidlegacyGoogle$4,650$3,300—▲1
GLM-5.2midZ.ai$3,940$5,2033.87×▼6
Grok-3midlegacyxAI$5,400$5,400——
Grok-4.20 ReasoningmidxAI$5,600$5,600——
Grok-4.20midxAI$5,600$5,600——
Grok 4.6midxAI$5,600$5,600——
Grok 4.5midxAI$5,600$5,600——
Gemini 3.1 PromidGoogle$6,200$3,9560.63×▲5
GPT-4.1midlegacyOpenAI$5,800$4,001—▼1
GPT-6 SolmidOpenAI$6,000$4,201——
GPT-6 Sol PromidOpenAI$6,000$4,201——
Claude Sonnet 5midAnthropic$6,000$4,204——
GPT-4omidlegacyOpenAI$7,250$5,001—▲1
GLM 4.7 (Cerebras)midCerebras$5,900$7,6967.53×▼5
GPT-5.6 TerramidlegacyOpenAI$7,750$5,501——
GPT-5.4midlegacyOpenAI$7,750$5,501——
Claude Sonnet 4.6midAnthropic$9,000$6,306——
Claude Sonnet 4.5midlegacyAnthropic$9,000$6,306——
Claude Sonnet 4midlegacyAnthropic$9,000$6,306——
GPT-5.6 SolmidlegacyOpenAI$12,000$8,401——
Claude Opus 5.5midAnthropic$12,000$8,408——
Claude Opus 4.8midAnthropic$15,000$10,4100.96×—
Claude Opus 4.7midlegacyAnthropic$15,000$10,510——
Claude Opus 4.6midlegacyAnthropic$15,000$10,510——
Claude Opus 4.5midlegacyAnthropic$15,000$10,510——
GPT-4 TurbofrontierlegacyOpenAI$28,000$19,003——
GPT-6 AstrafrontierOpenAI$30,000$21,003——
GPT-6 Astra ProfrontierOpenAI$30,000$21,003——
Claude Fable 5.1frontierAnthropic$30,000$21,020——
Claude Fable 5frontierlegacyAnthropic$30,000$21,020——
Claude Opus 5frontierlegacyAnthropic$45,000$31,530——
Claude Opus 4.1frontierlegacyAnthropic$45,000$31,530——
Claude Opus 4frontierlegacyAnthropic$45,000$31,530——
GPT-5.4 ProfrontierlegacyOpenAI$93,000$66,7301.04×—

Decision evidence · verified 2026-08-27

Classification-at-volume cost and capacity evidence

Frozen inputs, formulas, provenance, and decision boundaries are visible in the initial HTML. Unsupported evidence fails closed as Unavailable.

Label-taxonomy prompt-growth cube

Formula / scoring rule: Bill/record = (input tokens × input rate + output tokens × output rate)/1M; cached prefix is separated from uncached input.

Provenance: Frozen 500-input / 20-output record; 2/10/50/200 classes and 0/1/5 examples per class.

Frozen fixture / runVisible inputsField-level resultDecision boundaryState
2 classes / 0 examplesprompt 620; output 20; single-label JSONPer 1M records: 620M×$5 + 20M×$15 = $3,400.Tokenizer count is fixture-specific; no quality transfer.CALCULATED — exact fixture.
50 classes / 5 examplesprompt 8,420; output 34; multi-label JSONPer 1M records: 8,420M×$5 + 34M×$15 = $42,610.Schema growth must be included before choosing cache.CALCULATED — exact fixture.
200 classes / 5 examplesprompt 31,200; tokenizer result not stored; cache boundary requestedUnavailable — tokenizer result and cache-read rate are absentEstimated tokens cannot become an exact bill.Unavailable — tokenizer result and cache-read rate are absent

Module citation: OpenAI structured outputs documentation.

Confidence-and-review settlement tree

Formula / scoring rule: Cost/accepted label = total attempted path cost / accepted labels; abstentions and invalid repairs stay in numerator.

Provenance: Five-million-record scenario; confidence bands and routing probabilities are user-supplied, not accuracy observations.

Frozen fixture / runVisible inputsField-level resultDecision boundaryState
high confidence3.5M records; cheap model $0.003/record; no reviewAttempted cost $10,500; accepted labels Unavailable — calibrated acceptance is not observedConfidence is not accuracy.Unavailable — calibrated acceptance is not observed
abstain / retry1M records; retry $0.004; premium escalation $0.02; 8% invalid repairPath costs can be summed; accepted denominator Unavailable — not suppliedDo not invent pass probabilities.Unavailable — not supplied
human adjudication500K records; 6% review; 3 min at $60/hourScenario review labor = 30,000 × 0.05 × $60 = $90,000.Review cost is a scenario input, not provider billing.CALCULATED — user-supplied.

Module citation: All AI Ask extraction evidence registry.

Capacity and deadline planner

Formula / scoring rule: Completion capacity = min(RPM, TPM/input tokens) × active shards × window; missing quota fails closed.

Provenance: 5M and 50M record scenarios; provider quotas and completion guarantees are not inferred.

Frozen fixture / runVisible inputsField-level resultDecision boundaryState
5M / online5M × 500 input tokens; 2,500M tokens; 10 shards; RPM Unavailable — provider-specific quotaUnavailable — online completion deadline cannot be closedDo not display infinite throughput.Unavailable — provider-specific quota
50M / batch50M JSONL rows; 20 shards; submission window 24hBatch eligibility and completion SLA Unavailable — not established for selected providerA discount does not prove deadline compliance.Unavailable — not established for selected provider
partial failure100K submitted; 2,400 invalid; 700 duplicates; retry subset 1,100Unique accepted input = 96,900 before label acceptance.Deduplicate before cost-per-accepted calculation.CALCULATED — settlement counts closed.

Module citation: OpenAI rate limits.

Calculate cost per accepted label →

Prompt caching uses the provider's sourced read multiplier and amortises cache writes over the documented TTL; the largest published read discount is 90%. See the full formula and coverage disclosure →

Related

Alternatives to Amazon Nova Micro →Amazon cost calculator →LLM Chatbot cost →RAG Question Answering cost →

FAQ

Why does a tiny per-call cost matter here?
At tens of millions of calls a month, a fraction-of-a-cent difference per call compounds into a large monthly delta between models.
Should this always run through a batch API?
Whenever the classification does not need a synchronous response, yes — batch discounts apply directly to a workload this uniform.