How Much Does an Agentic Tool-Use Loop Cost per Month?

At production volume (25,000 calls/month), the cheapest effective option is Amazon Nova Micro at $27.38/month. The most expensive frontier option, GPT-5.4 Pro, runs $29,232/month — An agent task is not one call — it is a chain of several small round-trips (plan, call a tool, read the result, decide the next step), and the per-task cost is the sum of the whole chain.

How much does agentic tool loop cost per month?

At production volume (25,000 calls/month), the cheapest effective option for agentic tool loop is Amazon Nova Micro at $27.38 per month, verbosity-adjusted rather than list price. The most expensive frontier model, GPT-5.4 Pro, runs $29,232 per month for the same workload.

Verified 2026-08-08
Cheapest ≠ best. This page ranks by effective cost only. See our graded results for this task on Best LLM For →

Token shape

Shapemany small round-trips per completed task
Input / output tokens per call3K in / 0K out × 8 turns
Cacheable input70%
Batch-eligibleNo

"callsPerMonth" here counts completed tasks, each made of 8 model round-trips (plan → tool call → observe, repeated). Tool schemas and growing scratchpad context make most of each turn's input cacheable.

What drives this workload's cost?

The main token-volume driver here is input tokens: each call sends 3K input tokens and requests up to 0K output tokens, at 25,000 calls per month in the default volume. 70% of input is modeled as cache-eligible, so repeated prefixes can reduce the input charge. It is not marked batch-eligible because the profile assumes a synchronous response.

Volume

Side project
2,500 calls/mo
$2.74/mo cheapest
Production
25,000 calls/mo
$27.38/mo cheapest
Scale
250,000 calls/mo
$274/mo cheapest

Ranked cost — Production volume

ModelProviderList monthlyEffective monthlyVerbosityRank Δ
Amazon Nova MicrobudgetAmazon$29.40$27.380.76×—
Amazon Nova LitebudgetAmazon$50.40$49.100.91×—
GPT-5 NanobudgetlegacyOpenAI$54.00$36.49——
Gemini 2.5 Flash LitebudgetlegacyGoogle$84.00$47.29—▲2
GPT-6 LunabudgetOpenAI$90.00$54.98—▲2
GPT-6 Luna ProbudgetOpenAI$90.00$54.98—▲2
GPT-OSS 20BbudgetGroq$63.00$97.742.93×▼3
Ministral 8BbudgetMistral$99.00$98.370.93×▲1
Mistral Small 3.1budgetMistral$126$1210.85×▲4
GPT-4o MinibudgetlegacyOpenAI$126$73.47——
Grok-3 MinibudgetlegacyxAI$126$126——
GPT-OSS 120BbudgetGroq$126$1561.82×—
Llama 4 MaverickbudgetlegacyGroq$156$156—▲1
GPT-5.4 NanobudgetlegacyOpenAI$195$1080.78×▲1
Muse Spark 1.3 ContributorbudgetMeta$72.00$21512.95×▼10
Show all 76 models
CodestralbudgetMistral$234$2230.79×—
Gemini 3.1 Flash LitebudgetlegacyGoogle$240$1370.87×—
GPT-5 MinibudgetlegacyOpenAI$270$182—▲1
GPT-OSS 120B (Cerebras)budgetCerebras$255$3142.32×▼1
Gemini 3.5 Flash LitebudgetGoogle$330$220——
Gemini 2.5 FlashbudgetlegacyGoogle$330$220——
Mistral Large 3budgetMistral$390$390—▲1
DeepSeek V4 FlashbudgetDeepSeek$343$4692.59×▼1
GLM-5.1midlegacyZ.ai$492$492——
Qwen 3.8 30BmidGroq$540$540——
Qwen 3.6 27BmidlegacyGroq$540$540——
Qwen 3.7 PlusmidQwen$600$600——
Amazon Nova PromidAmazon$672$672——
Gemini 3.7 FlashmidGoogle$675$400——
GPT-5.4 MinimidlegacyOpenAI$720$457——
Gemini 3.1 FlashmidlegacyGoogle$720$445——
Claude Haiku 4.5midAnthropic$900$689—▲1
o3-MinimidlegacyOpenAI$924$539—▲1
GPT-5.6 LunamidlegacyOpenAI$960$610—▲1
Grok 4.3midxAI$900$9621.41×▼3
Muse Spark 1.3midMeta$1,005$1,005——
Mistral Medium 3midMistral$1,350$1,2780.84×▲5
Qwen 3.8 MaxmidQwen$1,344$1,344—▲1
Qwen 3.7 MaxmidQwen$1,344$1,344—▲1
Gemini 3.6 FlashmidGoogle$1,350$799—▲1
GPT-5midlegacyOpenAI$1,350$912—▲2
Grok-3midlegacyxAI$1,440$1,440—▲2
Gemini 3.5 FlashmidlegacyGoogle$1,440$889—▲2
Grok-4.20 ReasoningmidxAI$1,560$1,560—▲3
Grok-4.20midxAI$1,560$1,560—▲3
Grok 4.6midxAI$1,560$1,560—▲3
Grok 4.5midxAI$1,560$1,560—▲3
DeepSeek V4 PromidDeepSeek$1,030$1,5763.30×▼11
Gemini 3.1 PromidGoogle$1,920$9190.63×▲6
GPT-4.1midlegacyOpenAI$1,680$980—▲1
GPT-6 SolmidOpenAI$1,800$1,100—▲1
GPT-6 Sol PromidOpenAI$1,800$1,100—▲1
Claude Sonnet 5midAnthropic$1,800$1,378—▲1
GLM-5.2midZ.ai$1,104$1,8623.87×▼16
GPT-4omidlegacyOpenAI$2,100$1,225—▲1
GPT-5.6 TerramidlegacyOpenAI$2,400$1,525—▲1
GPT-5.4midlegacyOpenAI$2,400$1,525—▲1
GLM 4.7 (Cerebras)midCerebras$1,515$2,5927.53×▼12
Claude Sonnet 4.6midAnthropic$2,700$2,067——
Claude Sonnet 4.5midlegacyAnthropic$2,700$2,067——
Claude Sonnet 4midlegacyAnthropic$2,700$2,067——
GPT-5.6 SolmidlegacyOpenAI$3,600$2,199——
Claude Opus 5.5midAnthropic$3,600$2,756——
Claude Opus 4.8midAnthropic$4,500$3,3850.96×—
Claude Opus 4.7midlegacyAnthropic$4,500$3,445——
Claude Opus 4.6midlegacyAnthropic$4,500$3,445——
Claude Opus 4.5midlegacyAnthropic$4,500$3,445——
GPT-4 TurbofrontierlegacyOpenAI$7,800$4,298——
GPT-6 AstrafrontierOpenAI$9,000$5,498——
GPT-6 Astra ProfrontierOpenAI$9,000$5,498——
Claude Fable 5.1frontierAnthropic$9,000$6,889——
Claude Fable 5frontierlegacyAnthropic$9,000$6,889——
Claude Opus 5frontierlegacyAnthropic$13,500$10,334——
Claude Opus 4.1frontierlegacyAnthropic$13,500$10,334——
Claude Opus 4frontierlegacyAnthropic$13,500$10,334——
GPT-5.4 ProfrontierlegacyOpenAI$28,800$18,7271.04×—

Decision evidence · verified 2026-08-27

Agentic tool-loop cost evidence

Frozen inputs, formulas, provenance, and decision boundaries are visible in the initial HTML. Unsupported evidence fails closed as Unavailable.

Recurrence-based context ledger

Formula / scoring rule: Cₙ = system + userₙ + Σ(previous assistant/tool) + tool schema + scratchpad; total bill = Σ(inputₙ×rate + outputₙ×rate).

Provenance: Frozen 8-step loop: 3K initial input, 300 output/step, 1.2K tool result/step; cache hit is explicit.

Frozen fixture / runVisible inputsField-level resultDecision boundaryState
step 1system 600 + user 3,000 + tools 800; output 300Input 4,400; output 300; bill = (4,400×5 + 300×15)/1M = $0.0265.Constant-per-step shortcut is invalid after history grows.CALCULATED — recurrence row.
step 4prior history 6,000; current tool result 1,200; schema 800; output 300Input 11,000; cache-hit prefix Unavailable — provider returned cache usage is absentDo not apply a cache discount from prompt similarity.Unavailable — provider returned cache usage is absent
step 8history 14,400; tool result 1,200; output 300; stop=successInput 17,000; cumulative input 85,600; cumulative output 2,400.Cost is task-shaped, not 8 × first-step cost.CALCULATED — recurrence row.

Module citation: All AI Ask evidence registry (verified 2026-08-27).

Heterogeneous tool settlement tree

Formula / scoring rule: Tool-loop cost = model bill + Σ(tool fee + retry fee + compensating action); unsupported fee remains Unavailable.

Provenance: Frozen search, browser, code, database, and write-action calls; write action is side-effect gated.

Frozen fixture / runVisible inputsField-level resultDecision boundaryState
search / browser2 searches; 1 browser fetch; 3,100 payload tokens; 1 timeoutModel payload bill calculated; external search fee Unavailable — user tool tariff is not suppliedDo not treat tool fee as zero.Unavailable — user tool tariff is not supplied
code / database1 sandbox run; 2 DB reads; 1 retry; 4,800 payload tokensRetry is counted once; database egress Unavailable — not priced in fixtureExternal costs remain separate from model tokens.Unavailable — not priced in fixture
write action1 side-effecting write; idempotency key w-40; timeout before acknowledgementCompensating action required; duplicate-effect risk Unavailable — not observableNo successful completion credit until acknowledgement closes.Unavailable — not observable

Module citation: OpenAI function calling documentation.

Completed-task budget controller

Formula / scoring rule: Cost/completed = total path cost / completed tasks; stopped and escalated tasks remain in numerator.

Provenance: 5/10/20-step caps; planner/executor routing and completion rates are declared scenarios only.

Frozen fixture / runVisible inputsField-level resultDecision boundaryState
5-step cap25K tasks; 5 steps; completion scenario 82%; budget $0.08Stopped 4,500; completed 20,500; cost per completed Unavailable — observed completion rate is absentScenario completion is not reliability evidence.Unavailable — observed completion rate is absent
10-step capbranch factor 1.3; retry ceiling 2; budget $0.12Escalations and over-budget count Unavailable — not measured on a live runDo not publish a completion winner from assumptions.Unavailable — not measured on a live run
planner/executorcheap planner $0.002/step; premium executor $0.02/step; 70/30 splitIllustrative blended step = .7×.002 + .3×.02 = $0.0074.Routing cost excludes tool fees and acceptance.CALCULATED — scenario only.

Module citation: All AI Ask agent workload registry.

Calculate cost per completed agent task →

Prompt caching uses the provider's sourced read multiplier and amortises cache writes over the documented TTL; the largest published read discount is 90%. See the full formula and coverage disclosure →

Related

Alternatives to Amazon Nova Micro →Amazon cost calculator →LLM Chatbot cost →RAG Question Answering cost →

FAQ

Why 8 turns per task?
A representative agent task (research, then act, then verify) typically chains several tool calls before returning a final answer — this is a stated modelling assumption, not a measured average.
Does verbosity compound across turns?
Yes — a chatty model pays its verbosity penalty on every one of the 8 round-trips, not once, so the effect on total task cost is larger here than on a single-call workload.